【问题标题】:Python: How to find coordinates of short sequences in a FASTA file?Python:如何在 FASTA 文件中查找短序列的坐标?
【发布时间】:2015-05-21 02:55:00
【问题描述】:

我有一个短序列列表,我想在与包含原始序列的 fasta 文件进行比较后获取其坐标,或者换句话说,获取它的床文件。

快速文件:

>PGH2
CGTAGCGGCTGAGTGCGCGGATAGCGCGTA

短序列fasta文件:

>PGH2
CGGCTGAGT

有什么方法可以获取它的坐标吗?床具帮不上什么忙。

期望的输出:

PGH2  6 14

【问题讨论】:

    标签: python python-2.7 bioinformatics biopython fasta


    【解决方案1】:

    使用BioPyton

    from Bio import SeqIO
    
    for long_sequence_record in SeqIO.parse(open('long_sequences.fasta'), 'fasta'):
        long_sequence = str(long_sequence_record.seq)
    
        for short_sequence_record in SeqIO.parse(open('short_sequences.fasta'), 'fasta'):
            short_sequence = str(short_sequence_record.seq)
    
            if short_sequence in long_sequence:
                start = long_sequence.index(short_sequence) + 1
                stop = start + len(short_sequence) - 1
                print short_sequence_record.id, start, stop
    

    【讨论】:

      【解决方案2】:
      str1 = "CGTAGCGGCTGAGTGCGCGGATAGCGCGTA"
      str2 = "CGGCTGAGT"
      index = str1.index(str2)
      print index
      

      输出:index = 5, 得到 6,14 使用 index+1, index+len(str2)

      【讨论】:

      • 谢谢。供您参考,str1 和 str2 来自两个不同的文件,它们是带有序列 ID 的 fasta 文件。
      猜你喜欢
      • 2012-10-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-01-08
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-04-24
      相关资源
      最近更新 更多