【问题标题】:I need to speed up a nested loop in python that makes trillions of loops我需要在 python 中加速一个嵌套循环,它会产生数万亿个循环
【发布时间】:2020-05-25 00:02:13
【问题描述】:

filtradas 大约有 500000 个元素,_13 大约有 2000000 个元素,我使用 cython 达到的最短时间是 4 小时,但我需要在不到一个小时的时间内完成,我该怎么做?

两个列表都有带 1、2 或 X 的字符串元素,X 可以更改为 3

for i in filtradas:
        for x in _13:
            aciertos=0
            if i[0]==x[0]:
                aciertos+=1
            if i[1]==x[1]:
                aciertos+=1         
            if i[2]==x[2]:
                aciertos+=1
            if i[3]==x[3]:
                aciertos+=1
            if i[4]==x[4]:
                aciertos+=1
            if i[5]==x[5]:
                aciertos+=1         
            if i[6]==x[6]:
                aciertos+=1
            if i[7]==x[7]:
                aciertos+=1
            if i[8]==x[8]:
                aciertos+=1
            if i[9]==x[9]:
                aciertos+=1         
            if i[10]==x[10]:
                aciertos+=1
            if i[11]==x[11]:
                aciertos+=1
            if i[12]==x[12]:
                aciertos+=1
            if i[13]==x[13]:
                aciertos+=1
            if aciertos>=nroaciertos:
                filtradas13.append(i)
                break
    return filtradas13

【问题讨论】:

  • 可以有重复吗?如果没有,您可以使用 set 对象。如果有重复,首先对两个数组进行排序,然后遍历两个列表可能会快得多。
  • 你在linux上吗?多处理可能有助于分叉系统。
  • @tdelaney 感谢您的评论。我在 Windows 上使用 fx 8350
  • @BillLynch 感谢您的评论。我已经用 set 消除了两个列表中的重复项
  • 您能否详细说明您的意思是X 可以替换为3?这是否意味着 X 可以替换为空字符串或字符串'3'?

标签: python python-3.x list loops nested-loops


【解决方案1】:

numpy 值得一试。在此示例中,我怀疑转换为 numpy 数组的时间不值得比较,但如果您的数据首先导入 numpy,它可能会更快。

import numpy as np

_13_array = [np.array(list(x))for x in _13]
for i in filtradas:
    filtradas_array = np.array(list(i))
    for x in _13_array:
        aciertos = (x==filtradas_array).sum()
        if aciertos>=nroaciertos:
            filtradas13.append(i)
            break

【讨论】:

  • aciertos 始终为 0,我不是 numpy 专家,但我认为它不适用于字符串,列表中的元素类似于“11211X12111X1211”
  • 对,它必须类似于 np.array(tuple(x)) 和 np.array(tuple(i)) 才能将字符串转换为列。
  • @diegow98 - 我不确定格式。我已经更新了。
  • 它和其他使用cython的代码一样快,我明天用cython编译它并检查速度,谢谢。但我仍然需要检查它是否工作正常,因为我不明白 aciertos=(x==filtradas_array).sum() 的部分哈哈哈
  • numpy 将数据表示为低级 C 或 Fortran 数组,然后将 python 操作应用于整个数组,通常会释放 python GIL。所以,(x==filtradas_array) 变成了两个数组元素比较的低级布尔数组,实际上是 1 和 0。 .sum() 然后对该数组中的值求和。所以它与你在项目匹配时所做的+= 相同。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2013-05-03
  • 1970-01-01
  • 2021-04-27
  • 2020-08-19
  • 2020-04-18
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多