【发布时间】:2021-02-28 13:38:31
【问题描述】:
我编写了一个脚本来计算生物体中不同基因的值。我想得到一个 csv 文件,其中包含“物种 ID”、“基因座标签”列,然后是值本身。
我是 Python 新手,我不知道该怎么做?文件名的一个例子是:
NC_000117.1+CT_001.txt
+ 前后的每一位都是不同的。作为参考,前面的位是物种 ID,+ 后面的位是基因座标记。我有大约 6000 个这样的文件,而且它们非常小,所以处理时间应该不会太长!
这是我目前的代码:
import os
import io
import glob
import pandas as pd
import csv
work_dir = "User/..."
for path in glob.glob(os.path.join(work_dir, "*.txt")):
with io.open(path, mode="r", encoding="utf-8") as dir:
genome = dir.read()
GC3_list = list()
locus_list = list()
#calculates the total base number of each txt file
total_base = len(genome)
# print("Total no of bases:" + str(total_base))
#counts number of g and c bases at every 3rd position
g = genome[::3].count('g')
c = genome[::3].count('c')
# print("Number of G bases:"+str(g))
# print("Number of C bases:"+str(c))
#calculates final GC3 content in %
GC3 = ((g+c)/total_base) * 100
#print("GC3 content in % :" + str(GC3))
GC3_list.append(GC3)
#loci names taken from file
loci = os.path.basename(path).replace('.txt', "")
locus_list.append(loci)
df = pd.DataFrame({'Locus tag':locus_list,
'GC3 %': GC3_list})
df.to_excel('output')
【问题讨论】: