From 90ba01e8b69eae5d282db2f6716a5b74466d2cc6 Mon Sep 17 00:00:00 2001
From: HarshitGupta11 <50410275+HarshitGupta11@users.noreply.github.com>
Date: Mon, 12 Aug 2019 11:57:49 +0530
Subject: [PATCH] Update README.md
---
README.md | 23 +++++++++++++++++++++++
1 file changed, 23 insertions(+)
diff --git a/README.md b/README.md
index bae92a0..1dc52d3 100644
--- a/README.md
+++ b/README.md
@@ -71,3 +71,26 @@ IF YOU DID EVERYTHING RIGHT YOU SHOULD SEE A FILE NAMED `dataset` IN THE SAME DI
This is where it gets interesting. You are gonna train your own model architecture or you can use the one given in the `model.py` script. You can directly change the model as well in the `model.py`.
To train the model, run:
`python train.py model_name negative_start_composition negative_end_composition epochs learning_rate learning_rate_decay no_of_samples batch_size`
+
+## Predictions:
+To make predictions you need to have the prediction files in a pre-defined format like this [file](ftp://ftp.ebi.ac.uk/pub/databases/ensembl/mateus/gsoc_2019/balanced_random_mix_ortholog_paralog_negative.txt.gz). All the fields have to tab seperated and in the same order.
+
+**Get the prediction files:**
+Create a new directory with any name of your choice in the code directory and paste all the files on which you want to make the predictions inside it.
+Run:
+`python pfam_folder_pred.py directory_name`
+This will read all the files on the directory and write all the protein sequences on a FAST-A file named `prediction_directory_name.fa`.
+Run hmmer scan on that file.
+
+**Parse the PFAM file**
+This file parses the PFAM file and creates some maps.
+Run:
+`python pfam_db_parser.py domtblout_file_name`
+
+**Now you have all the resources required to Make predictions on the required file.**
+Copy the file you want to make predictions on from the directory that you created earlier to the code directory and Run:
+`python prediction_pfam.py file_name_with_extension model_name number_of_threads start end name domtblout_file_name`.
+where:
+`start end`:the locations from which you want to make predictions in the file.
+`name`:Since you can run multiple predcitions at the same time this serves as a unique identifier for the temporary files being created.
+`domtblout_file_name`:It is the name of the file that you get after running the hmmer scan on the fast-a files.