From b68588ee5d2dd767ada835f6ac042b7ec29b9eb0 Mon Sep 17 00:00:00 2001 From: HarshitGupta11 <50410275+HarshitGupta11@users.noreply.github.com> Date: Sun, 23 Jun 2019 00:16:54 +0530 Subject: [PATCH] Update README.md --- README.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/README.md b/README.md index 6387099..62a3ae8 100644 --- a/README.md +++ b/README.md @@ -24,12 +24,14 @@ The designated folders to store the files are as follows:
`data_homology` all the homology databases that you want to use.
`geneseq`To store the `cds` sequence files in `fasta` format.
You can assign different folder names if you want but it's better to stick to them.
+ **Create Genome Maps:**
The purpose is to create maps of all the genes present in the `gtf` files with respect to their chromosomes, a map of all the genes belonging to the same chromosome in the given species, a map of all the genes in the given species, a map of all the species whose data has been successfully read.
To create genome maps run this command:
`python create_genome_maps.py -d path -r` where:
`path`: link to the directory where all the `gtf` files exist. If you have used `ftpg.py` then the path is `data`.
Note: Genome Maps can be downloaded from this [link](https://drive.google.com/open?id=1GjV6dT-Hpf2LWQ-vSpekqqQ7RF_tH8So).
+ **Create/Update Neighbor Genes File:**
This project uses the measure of conserved synteny to predict the homology type. Therefore, to predict the homology type we need to find the neighboring genes of the given homologous pair of genes.
This file finds the neighboring genes of all the rows in the given databases and writes it to a file called `processed/neighbor_genes.json` .
@@ -73,5 +75,17 @@ To sample the negative dataset run this command:
**Update Neigbor Genes**
(Again!!!, ¯\\_(ツ)_/¯). To update the neighbor genes with the new sampled dataset use this command:
`python update_neighbor_genes.py`
+*(Seriously!,That's just it)* +
+**Prepare Synteny Matices**
+To prepare the synteny matrices use this command:
+`python prepare_synteny_matrix_negative.py`
+ +**Extract Other Features:**
+To extract other features run this command:
+`python prepare_other_features_negative.py`
+ +**NOW WE HAVE BOTH: A DATASET CONTAINING POSITIVE SAMPLES AND A DATASET CONTAINING NEGATIVE SAMPLES**
+I know, right!!