compara-deep-learning/README.md
2019-08-12 09:58:45 +05:30

4.2 KiB

New Readme

The aim of this project is to use Deep Neural-Nets to predict homology type between the give pair of genes. This file provides the instructions to replicate the results of this project from scratch.

Requirements:

  1. A machine with atleast 8GB of RAM (although 16-32GB is recommended. It depends on the no. of homology databases that you are willing to use in the preparation of the dataset), a graphic card for training the deep neural nets. A single GPU machine would suffice. The model can be trained on CPU as well but will be a lot faster if trained on a GPU.
  2. A stable Internet Connection.
  3. A native/virtual python environment. Install the dependencies from requirements.txt using :
    pip install -r requirements.txt

Step 1: Data Preparation:

The model uses a synteny matrix and some other factors derived from the species tree to make predictions.
Download The Required Files: So as to prepare data we will need the following files:
1.All the gtf files to get the start and end locations of the genes and find their neighboring genes. This is used to create the synteny matrix which helps to see the conserved synteny among the genes.
2. All the cds files in FAST-A format. They are required but are not mandatory, if the files are not provided the sequences are directly accessed from the REST API but the process can be slow :( . It's better to have all the cds files.
3. Homology Databases of Your Choice. All the databases have the same name so its better to change the name of the files with their respective speicies names. One format that works best is species_name.tsv.gz.
4. All the pep files in FAST-A format. They are required to get the protein sequences to run the pfam scan on.

To download the gtf, cdsand pep files ftpg.py can be used. This scripts writes the links of all the required gtf and cds files to gtf-link.txt and seq_link.txt. You can use your own script to download the files or just enter y when prompted for permission to download the files. It will automatically download all the files and store them in designated folder. You can manually download each files by pasting link from the files in the browser.

Create Genome Maps:
The purpose is to create maps of all the genes present in the gtf files with respect to their chromosomes, a map of all the genes belonging to the same chromosome in the given species, a map of all the genes in the given species, a map of all the species whose data has been successfully read.
To create genome maps run this command:
python create_genome_maps.py:
Note: Genome Maps can be downloaded from this link.

Select the Records from each homology database:
This step will select the specified no. of records from each of the homology databases on the basis of distant species,GOC score,homology type etc.
To select the data run:
python select_data.py number_of_records_to_be_selected_from_each_file.

Find the Neighbor Genees of the Selected Records:
This step will find the neighbor genes of all the selected records from the homology databases and write it to the processed directory.
Run:
python neighbor_gene.py

Create Synteny Matrix and Write Protein Sequences:
This step will create the synteny matrices and write the protein sequences on a FAST-A file named protein_seq_positive.fa.You will have to run the hmmer-scan on this file to get the PFAm domains.
To create the synteny matrices:
python prepare_synteny_matrix.py number_of_threads.
Note:The script uses multiprocessing to prepare synteny matrices. Since each thread has it own copy of all the resources its better to run with higher number of threads on a high-RAM machine.

Process The Negative Dataset:
Negative samples are a non-homologous pair of genes. You can get the negative samples from here.Download one of your choice :).
Process the negative set by using:
python process_negative.py negative_database_file_name number_of_threads<br>