From da442a91b126bed5d75976be39e959b8d2dfa02d Mon Sep 17 00:00:00 2001 From: HarshitGupta11 <50410275+HarshitGupta11@users.noreply.github.com> Date: Sun, 23 Jun 2019 00:49:47 +0530 Subject: [PATCH] Update README.md --- README.md | 32 +++++++++++++++++++++++++++++--- 1 file changed, 29 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 62a3ae8..91b2160 100644 --- a/README.md +++ b/README.md @@ -39,13 +39,13 @@ To find the neighboring genes of the homologous genes in the databases run this `python update_neighbor_genes.py -d path -r -test` to test the file for 5 samples or you can directly run:
`python update_neighbor_genes.py -d path -r -run`.
It will update/create the `neighbor_genes.json` file in the `processed` directory.
-
**Prepare Synteny Matrices:** -The aim is to get all the alignments of the genes with respect to one another. This step will create the synteny matrix for each record in the datset where the *i,jth* entry of the matrix is the alignment of *ith* gene with *jth* gene. We use both local and global alignments. Also to provide extra features each *ith* gene is reversed and again the alignment is taken again with the *jth* gene. This solves two purposes:
+
**Prepare Synteny Matrices:**
+The aim is to get all the alignments of the homologous neighboring genes with respect to one another. This step will create the synteny matrix for each record in the datset where the *i,jth* entry of the matrix is the alignment of *ith* gene with *jth* gene. We use both local and global alignments. Also to provide extra features each *ith* gene is reversed and again the alignment is taken again with the *jth* gene. This has two purposes:
1.It makes sure that the `+` strand gene sequence is always aligned with `+` strand gene sequence and vice-versa.
2.It provides an extra depth to the synteny matrix which helps in better recognition of features.
To prepare synteny matrices run this command:
-`python prepare_synteny_matrices.py nos` where:
+`python prepare_synteny_matrix.py nos` where:
`nos` is an integer specifying the number of rows that you want to sample from each file to prepare the dataset. This file assumes that:
`neighbor_genes.json` is present in the `processed` sub-directory.
@@ -89,3 +89,29 @@ To extract other features run this command:
**NOW WE HAVE BOTH: A DATASET CONTAINING POSITIVE SAMPLES AND A DATASET CONTAINING NEGATIVE SAMPLES**
I know, right!! +## STEP 2: TRAIN THE MODEL: +This will come later.
+ +## STEP 3: TESTING THE MODEL: +Download the model from [here](https://drive.google.com/open?id=1_TmsH8uLrKAmZKbCOqRo33by0n7jFvty) and extract it to the code directory.
+**To test the model** run:
+`python test_model.py`
+It asks for a gene id and a second gene id and predicts the homology type.
+ +## SUMMARY: +I know you are probably thinking (and angry,bored,frustrated) "*Why did I have to read this long guide when I could have directly come here*" or you have directly scrolled to this(as you were checking how long the page is, *we all have been there*), So congratulations!! you have come to the right point.
+The simple step by step summaray is:
+1. Run `ftpg.py` and enter `y` when it asks to download for file. Believe me, you don't want to do it manually there are **199** of them.
+2. Download the homology databases of your choice and rename them with this format `species_name.tsv.gz`.
+3. Run `create_genome_maps.py -d data -r` to create genome maps.
+4. Run `update_neighbor_genes.py -d data_homology -r -run` to create/update the neighboring genes map file.
+5. Run `prepare_synteny_matrix.py nos` to sample `nos` records from the dataset and create their synteny matrices.
+6. Run `prepare_other_features.py` to finalize the positive dataset.
+7. Run `prepare_negative_dataset.py nos random_seed` to sample the negative dataset.
+8. Run `update_neighbor_genes.py` to update the neighbor genes file.
+9. Run `prepare_synteny_matrix_negative.py` to prepare the synteny matrices.
+10. Run `prepare_other_features_negative.py` to finalize the negative dataset.
+ +**To test the model:**
+Run `test_model.py` with the genome_maps.
+