【问题标题】:Longest Common Sequence -Syntax error最长公共序列 - 语法错误
【发布时间】:2014-09-29 10:08:12
【问题描述】:

我正在尝试查找两个 DNA 序列的 LCS。我正在输出矩阵形式以及包含最长公共序列的字符串。但是,当我在代码中同时返回矩阵和列表时,会出现以下错误:IndexError: string index out of range

如果我要删除涉及变量 temp 和 higestcount 的编码,我的代码将很好地输出我的矩阵。我正在尝试对矩阵使用类似的编码来生成我的列表。有没有办法避免这个错误?根据序列 AGCTGGTCAG 和 TACGCTGGTGGCAT,最长的公共序列应该是 GCTGGT。

def lcs(x,y):
    c = len(x)
    d = len(y)
    plot = []
    temp = ''
    highestcount = ''

    for i in range(c):
        plot.append([])
        temp.join('')
        for j in range(d):
            if x[i] == y[j]:
                plot[i].append(plot[i-1][j-1] + 1)
                temp.join(temp[i-1][j-1])
            else:
                plot[i].append(0)
                temp = ''
                if temp > highestcount:
                    highestcount = temp

    return plot, temp

x = "AGCTGGTCAG"
y = "TACGCTGGTGGCAT"
test = compute_lcs(x,y)

print test

【问题讨论】:

    标签: python string sequences lcs


    【解决方案1】:

    在temp.join(temp[i-1][j-1])的第一次迭代中,temp作为变量是一个空字符串,''

    字符串中没有可以通过索引调用的字符,因此 temp[any_number] 将抛出 index out of range 异常。

    【讨论】:

    • 但是我从字符串中什么都没有开始,因为我希望程序在 for 循环中进行时填充最长的序列(当然,用前面的一个替换最长的序列一个已保存)
    • Python 将首先执行最里面的操作,该操作是 temp[i-1][j-1]。在该操作时, temp == '' 没有可以通过将索引号传递给列表来找到的字符串字符(调用 any_string[any_integer] 将 any_string 存储为字符列表,即 ['H',' e','l','l','o'])。我没有遗传学背景,所以我不确定您要加入现有临时字符串的内容,但在空字符串中按索引查找字符将永远不会起作用。您能否提供一个您希望看到的输出示例?
    【解决方案2】:

    据我所知,join() 通过另一个字符串连接一个字符串数组。例如,"-".join(["a", "b", "c"]) 将返回 a-b-c。

    此外,您首先将temp 定义为字符串,但稍后使用双索引引用它,几乎就像它是一个数组一样。据我所知,您可以通过单个索引调用来引用字符串中的字符。例如,a = "foobar"、a[3] 返回b。

    我将您的代码更改为以下内容。初始化数组以开始以避免索引问题。

    def lcs(x,y):
        c = len(x)
        d = len(y)
        plot = [[0 for j in range(d+1)] for i in range(c+1)]
        temp = [['' for j in range(d+1)] for i in range(c+1)]
        highestcount = 0
        longestWord = ''
    
        for i in range(c):
            for j in range(d):
                if x[i] == y[j]:
                    plot[i+1][j+1] = plot[i][j] + 1
                    temp[i+1][j+1] = ''.join([temp[i][j],x[i]])
                else:
                    plot[i+1][j+1] = 0
                    temp[i+1][j+1] = ''
                    if plot[i][j] > highestcount:
                        highestcount = plot[i][j]
                        longestWord = temp[i][j]
    
        return plot, temp, highestcount, longestWord
    
    x = "AGCTGGTCAG"
    y = "TACGCTGGTGGCAT"
    test = lcs(x,y)
    print test
    

    【讨论】:

      【解决方案3】:

      在我看来,您正在浏览一个不必要的复杂屏幕,这会导致混乱,包括其他人提到的空字符串。

      例如,这仍然很冗长,但我认为更容易理解(并返回预期的答案):

      def lcs(seq1, seq2):
          matches = []
          for i in range(len(seq1)):
              j = 1
              while seq1[i:j] in seq2:
                  j+=1 
                  if j > len(seq1):
                      break
              matches.append( (len(seq1[i:j-1]), seq1[i:j-1]) )
          return max(matches)
      
      seq1 = 'AGCTGGTCAG'
      seq2 = 'TACGCTGGTGGCAT'
      lcs(seq1, seq2)
      

      返回

      (6, 'GCTGGT')
      

      【讨论】:

      • 这段代码很好,但我无法输出我的矩阵。还有另一种方法可以将 i 仅用于一个序列,将 j 用于另一个序列吗?
      • 你能举例说明你希望你的输出是什么样子吗?
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-03-01
      • 2014-03-31
      • 2011-02-25
      • 2013-02-13
      • 1970-01-01
      相关资源
      最近更新 更多