From a12d6e2ee57a624546e0da1a37f629d8db2cdb90 Mon Sep 17 00:00:00 2001
From: HarshitGupta11 <50410275+HarshitGupta11@users.noreply.github.com>
Date: Sat, 22 Jun 2019 23:28:05 +0530
Subject: [PATCH] Update README.md
---
README.md | 17 ++++++++++++++++-
1 file changed, 16 insertions(+), 1 deletion(-)
diff --git a/README.md b/README.md
index 9df9cd1..6387099 100644
--- a/README.md
+++ b/README.md
@@ -52,7 +52,7 @@ all the `cds` sequence files are present in the `geneseq` directory, but this is
All the synteny matrices are created and stored in `processed/synteny_matrices` directory in `.npy` format.
**Extract Other Features:**
This will extract other features from the species trees.
-Run this command:
+Run this command:
`python prepare_other_features.py`
.
After the completion of this process you should get a binary file named `dataset` in the directory.
@@ -60,3 +60,18 @@ After the completion of this process you should get a binary file named `dataset
*(I know right)*
+
+## TO PREPARE NEGATIVE DATASET:
+This dataset will contain the gene pairs which are not homologous to each other.
+**Sample the Negative Dataset:**
+Negative dataset takes one gene from the homology database and another gene from the `gtf` files which is not present in any of the homology databases read. This ensures that the two genes selected belong to mutually exclusive sets and hence are not homologous to each other.
+To sample the negative dataset run this command:
+`python prepare_negative_dataset.py nos random_seed` where
+`nos`:number of rows to sample from the dataset.
+`random_seed`:a random number to seed the random number generator.
+
+**Update Neigbor Genes**
+(Again!!!, ¯\\_(ツ)_/¯). To update the neighbor genes with the new sampled dataset use this command:
+`python update_neighbor_genes.py`
+
+