Daily report
‘22/12/23
Cleaning the files
Because the species were mixed in the files, redid ^^Coding_challenge_4_TENT4A_PAPD7.py^^ and ^^Coding_challenge_4_TENT4B_PAPD5.py^^ on my files and got ^^NEW_(order)_(gene).fas^^
Ran ^^Coding_challenge_5.py^^ and replaced the Sequences that were too short
[ ] Edited ^^Coding_challenge_5.py^^ using Biopython pairwise
Getting new species
- Take the NCBI accession number from ^^NEW_(order)_(gene).fas^^ and used blast to get csv and fasta file

- List the species that weren’t in the fasta file
‘22/12/26
- Make ^^Coding_challenge_5_draft.py^^ compare with the longest sequence using dict key values
d1={record.id:record.seq for record in SeqIO.parse(file, "fasta")}
d2={record.id:len(record.seq) for record in SeqIO.parse(file, "fasta")}
d2=sorted(d2.items(), key=lambda x:x[1],reverse=True)
#df_d1=pd.DataFrame.from_dict(d1) ←this gave an error (ValueError: All arrays must be of the same length)
df_d2 = pd.DataFrame(d2)
```
## ‘22/12/27
> ## Finish adding species
* Finished adding species acquired from BLAST to each fasta file
### Converting a text file into a list by splitting the text on the occurrence of ‘.’.
python
opening the file in read mode
my_file = open("file1.txt", "r")
reading the file
data = my_file.read()
replacing end of line('/n') with ' ' and
splitting the text it further when '.' is seen.
data_into_list = data.replace('\n', ' ').split(".")
printing the data
print(data_into_list) my_file.close() ```
- Left only the Longest sequence in the file. ^^NEW_12_27_(order)_(gene).fas^^ has only one sequence per species.