From 3d27debf4a90ac939c8adce1e9e9e62c5b1af439 Mon Sep 17 00:00:00 2001 From: HarshitGupta11 <50410275+HarshitGupta11@users.noreply.github.com> Date: Mon, 12 Aug 2019 10:38:50 +0530 Subject: [PATCH] Update README.md --- README.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/README.md b/README.md index 9821a6d..99bc36b 100644 --- a/README.md +++ b/README.md @@ -47,5 +47,22 @@ Negative samples are a non-homologous pair of genes. You can get the negative sa Process the negative set by using:
`python process_negative.py negative_database_file_name number_of_threads` +**Run the HMMER scan on the Protein Sequences:**
+The idea is that homologous genes will have overlapping domains. Run the hmmer scan on the `protein_seq_positve.fa` and `pro_seq_negative.fa` with the `-domtblout` option. +**Parse the PFAm domain files:**
+This file will parse the hmmer scan database. You will have to parse both the positive samples and the negative sample database.
+To parse, run:
+`python pfam_parser.py domtblout_file_name_positive domtblout_file_name_negative`. +**Create PFAM matrices:**
+This step will create the pfam matrices. This might take some time...
+Run:
+`python pfam_matrix.py name_of_negative_database_you_earlier_processed`
+ +**Finalize the Dataset:**
+This step combines everything and finalizes the dataset by reading the processed factors and extracting some basic features from the species tree. +Run:
+`python finalize_dataset.py name_of_negative_database_you_earlier_processed`
+ +IF YOU DID EVERYTHING RIGHT YOU SHOULD SEE A FILE NAMED `dataset` IN THE SAME DIRECTORY.