【问题标题】:Nested for loops - getting the right order嵌套 for 循环 - 获得正确的顺序
【发布时间】:2015-04-20 09:45:49
【问题描述】:

在进行嵌套 for 循环时,我无法掌握如何正确地输出正确的顺序。

我有一个整数列表:

[7, 9, 12]

还有一个包含多行 DNA 序列数据的 .txt。

>Ind1 AACTCAGCTCACG
>Ind2 GTCATCGCTACGA 
>Ind3 CTTCAAACTGACT

我正在尝试创建一个嵌套的 for 循环,它采用第一个整数 (7),遍历文本行并在每行的位置 7 处打印字符。然后取下一个整数,并在每行的位置 9 处打印每个字符。

with (Input) as getletter:
    for line in getletter:
        if line [0] == ">":

            for pos in position:
                snp = line[pos]
                print line[pos], str(pos)

当我运行上面的代码时,我得到了我想要的数据,但是顺序错误,像这样:

A  7
T  9
G  12
T  7
A  9
G  12
T  7
C  9
A  12

我想要的是这个:

A  7
T  7
T  7
T  9
A  9
C  9
G  12
G  12
A  12

我怀疑可以通过更改代码的缩进来解决问题,但我无法正确解决。

--------编辑--------

我试图交换两个循环,但我显然没有得到更大的图景,这给了我与上述相同(错误)的结果。

with (Input) as getsnps:
    for line in getsnps:
        if line[0] == ">":
            hit = line
        for pos in position:
                print hit[pos], pos

【问题讨论】:

  • 您的打印语句不会产生第一个输出;它将打印A 7、T 9 等,因此交换了字母和位置。
  • 我删除了python-3.x 标签;您在这里使用的是 Python 2,正如 print 是一个声明所证明的那样。
  • 如何从列表[7, 9, 12] 到31 和119 的位置?
  • 您可以交换循环,但这意味着您每次都必须从头到尾重新读取文件。我们在这里谈论多少行?
  • 如果文件足够小,你做哪一个并不重要。如果它足够大......那么你做什么取决于你有多少 RAM(以及你是否在 64 位机器上),你的驱动器有多快和/或你有多少磁盘缓存,......但只要做任何一个首先对你有意义,然后如果它太慢,至少你有一些需要优化的东西,它正在工作并且你理解。 :)

标签: python for-loop nested indentation


【解决方案1】:

尝试回答:

with (Input) as getletter:
    lines=[x.strip() for x in getLetter.readlines() if x.startswith('>') ]
    for pos in position:
        for line in lines:
            snp = line[pos]
            print ("%s\t%s" % (pos,snp))

文件被读取并缓存到数组中(行,丢弃不以>开头的文件) 然后我们遍历位置,然后遍历行并打印预期结果。

请注意,您应该检查您的偏移量是否大于您的线。

没有列表理解的替代方案(将使用更多内存,特别是如果您有很多无用的行(即不以 '>' 开头)

with (Input) as getletter:       
    lines=getLetter.readlines()
    for pos in position:
        for line in lines:
            if line.startswith('>'):
                 snp = line[pos]
                 print ("%s\t%s" % (pos,snp))

另一种存储方式(假设位置小而输入大)

with (Input) as getletter:
    storage=dict()
    for p in positions:
        storage[p]=[]
    for line in getLetter:
        for p in positions:
            storage[p]+=[line[pos]]
for (k,v) in storage.iteritems():
    print ("%s -> %s" % (k, ",".join(v))

如果位置包含大于行大小的值,使用line[p] 将触发异常(IndexError)。您可以捕获它或测试它

try:
    a=line[pos]
except IndexError:
    a='X'

if pos>len(line):
   a='X'
else:
   a=line[pos]

【讨论】:

  • 假设positions 相当小,您也可以将字符存储在这些位置而不是所有行。
  • @Bruce,你能告诉我如何做同样的事情,但没有列表理解(我认为就是这样),因为我还不熟悉这些(对此仍然很陌生)?另外,您能否详细说明偏移量大于线的含义?
  • @hjalte :做了替代方案(和@tobias_k 提出的那个)
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2022-08-05
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多