compara-deep-learning/README.md
2019-06-24 10:42:21 +05:30

9.4 KiB
Raw Blame History

New Readme

The aim of this project is to use Deep Neural-Nets to predict homology type between the give pair of genes. This file provides the instructions to replicate the results of this project from scratch.

Requirements:

  1. A machine with atleast 8GB of RAM (although 16-32GB is recommended. It depends on the no. of homology databases that you are willing to use in the preparation of the dataset), a graphic card for training the deep neural nets. A single GPU machine would suffice. The model can be trained on CPU as well but will be a lot faster if trained on a GPU.
  2. A stable Internet Connection.
  3. A native/virtual python environment. Install the dependencies from requirements.txt using :
    pip install -r requirements.txt

Step 1: Data Preparation:

The model uses a synteny matrix and some other factors derived from the species tree to make predictions.
Download The Required Files: So as to prepare data we will need the following files:
1.All the gtf files to get the start and end locations of the genes and find their neighboring genes. This is used to create the synteny matrix which helps to see the conserved synteny among the genes.
2. All the cds files in FAST-A format. They are required but are not mandatory, if the files are not provided the sequences are directly accessed from the REST API but the process can be slow :( . It's better to have all the cds files.
3. Homology Databases of Your Choice. All the databases have the same name so its better to change the name of the files with their respective speicies names. One format that works best is species_name.tsv.gz.

To download the gtf and cds files ftpg.py can be used. This scripts writes the links of all the required gtf and cds files to gtf-link.txt and seq_link.txt. You can use your own script to download the files or just enter y when prompted for permission to download the files. It will automatically download all the files and store them in designated folder. You can manually download each files by pasting link from the files in the browser.

The designated folders to store the files are as follows:
data all the gtf files.
data_homology all the homology databases that you want to use.
geneseqTo store the cds sequence files in fasta format.
You can assign different folder names if you want but it's better to stick to them.

Create Genome Maps:
The purpose is to create maps of all the genes present in the gtf files with respect to their chromosomes, a map of all the genes belonging to the same chromosome in the given species, a map of all the genes in the given species, a map of all the species whose data has been successfully read.
To create genome maps run this command:
python create_genome_maps.py -d path -r where:
path: link to the directory where all the gtf files exist. If you have used ftpg.py then the path is data.
Note: Genome Maps can be downloaded from this link.

Create/Update Neighbor Genes File:
This project uses the measure of conserved synteny to predict the homology type. Therefore, to predict the homology type we need to find the neighboring genes of the given homologous pair of genes.
This file finds the neighboring genes of all the rows in the given databases and writes it to a file called processed/neighbor_genes.json .
To find the neighboring genes of the homologous genes in the databases run this command:
python update_neighbor_genes.py -d path -r -test to test the file for 5 samples or you can directly run:
python update_neighbor_genes.py -d path -r -run.
It will update/create the neighbor_genes.json file in the processed directory.

Prepare Synteny Matrices:
The aim is to get all the alignments of the homologous neighboring genes with respect to one another. This step will create the synteny matrix for each record in the datset where the i,jth entry of the matrix is the alignment of ith gene with jth gene. We use both local and global alignments. Also to provide extra features each ith gene is reversed and again the alignment is taken again with the jth gene. This has two purposes:
1.It makes sure that the + strand gene sequence is always aligned with + strand gene sequence and vice-versa.
2.It provides an extra depth to the synteny matrix which helps in better recognition of features.

To prepare synteny matrices run this command:
python prepare_synteny_matrix.py nos where:
nos is an integer specifying the number of rows that you want to sample from each file to prepare the dataset. This file assumes that:
neighbor_genes.json is present in the processed sub-directory.
homology databases are present in the data_homology sub-directory
. all the cds sequence files are present in the geneseq directory, but this is not mandatory if you want to directly access all the sequences over the REST-API. It accesses all the not found gene sequences from the REST-API after reading all the files.
All the synteny matrices are created and stored in processed/synteny_matrices directory in .npy format.

Extract Other Features:
This will extract other features from the species trees. Run this command:
python prepare_other_features.py
. After the completion of this process you should get a binary file named dataset in the directory.

TILL NOW WE HAVE PREPARED ONLY THE DATASET CONTAINING POSITIVE SAMPLES.
I know, right!!
          (I know right)

TO PREPARE NEGATIVE DATASET:

This dataset will contain the gene pairs which are not homologous to each other.
Sample the Negative Dataset:
Negative dataset takes one gene from the homology database and another gene from the gtf files which is not present in any of the homology databases read. This ensures that the two genes selected belong to mutually exclusive sets and hence are not homologous to each other.
To sample the negative dataset run this command:
python prepare_negative_dataset.py nos random_seed where
nos:number of rows to sample from the dataset.
random_seed:a random number to seed the random number generator.

Update Neigbor Genes
(Again!!!, ¯\(ツ)/¯). To update the neighbor genes with the new sampled dataset use this command:
python update_neighbor_genes_ndf.py
(Seriously!,That's just it)

Prepare Synteny Matices
To prepare the synteny matrices use this command:
python prepare_synteny_matrix_negative.py

Extract Other Features:
To extract other features run this command:
python prepare_other_features_negative.py

NOW WE HAVE BOTH: A DATASET CONTAINING POSITIVE SAMPLES AND A DATASET CONTAINING NEGATIVE SAMPLES
I know, right!!

STEP 2: TRAIN THE MODEL:

This will come later.

STEP 3: TESTING THE MODEL:

Download the model from here and extract it to the code directory.
To test the model run:
python test_model.py
It asks for a gene id and a second gene id and predicts the homology type.

SUMMARY:

If you have directly scrolled to this part(as you were checking how long the page is, we all have been there), Congratulations!! you have come to the right part.
The simple step by step summaray is:

  1. Run ftpg.py and enter y when it asks to download for file. Believe me, you don't want to do it manually, there are 199 of them.
  2. Download the homology databases of your choice to the data_homology directory and rename them with this format species_name.tsv.gz.
  3. Run create_genome_maps.py -d data -r to create genome maps.
  4. Run update_neighbor_genes.py -d data_homology -r -run to create/update the neighboring genes map file.
  5. Run prepare_synteny_matrix.py nos to sample nos records from the dataset and create their synteny matrices.
  6. Run prepare_other_features.py to finalize the positive dataset.
  7. Run prepare_negative_dataset.py nos random_seed to sample the negative dataset.
  8. Run update_neighbor_genes_ndf.py to update the neighbor genes file.
  9. Run prepare_synteny_matrix_negative.py to prepare the synteny matrices.
  10. Run prepare_other_features_negative.py to finalize the negative dataset.

To test the model:
Run test_model.py with the genome_maps.

Some Tips:
1.Don't change the directory names as they might create additional problems in the execution.
2.Preparing the dataset might take some time especially the part where synteny matrices are created, so it's better to run multiple concurrent jobs.
3.Here is the link to the genome maps,neighbor_genes.json,negative dataset with 1 million entries.