【问题标题】:Python - Performance when iterating over dictionary keysPython - 迭代字典键时的性能
【发布时间】:2018-10-26 02:45:35
【问题描述】:

我有一个相对较大的文本文件(大约 7m 行),我想在它上面运行一个特定的逻辑,我将在下面尝试解释:

A1KEY1
A2KEY1
B1KEY2
C1KEY3
D1KEY3
E1KEY4

我想统计按键出现的频率,然后将频率为 1 的输出到一个文本文件中,频率为 2 的输出到另一个文本文件中,频率高于 2 的输出到另一个文本文件中。

这是我目前拥有的代码,但它在字典上的迭代速度非常缓慢,并且越进展越慢。

def filetoliststrip(file):
    file_in = str(file)
    lines = list(open(file_in, 'r'))
    content = [x.strip() for x in lines] 
    return content


dict_in = dict()    
seen = []


fileinlist = filetoliststrip(file_in)
out_file = open(file_ot, 'w')
out_file2 = open(file_ot2, 'w')
out_file3 = open(file_ot3, 'w')

counter = 0

for line in fileinlist:
    counter += 1
    keyf = line[10:69]
    print("Loading line " + str(counter) + " : " + str(line))
if keyf not in dict_in.keys():
    dict_in[keyf] = []
    dict_in[keyf].append(1)
    dict_in[keyf].append(line)
else:
    dict_in[keyf][0] += 1
    dict_in[keyf].append(line)


for j in dict_in.keys():
    print("Processing key: " + str(j))
    #print(dict_in[j])
    if dict_in[j][0] < 2:
        out_file.write(str(dict_in[j][1]))
    elif dict_in[j][0] == 2:
        for line_in in dict_in[j][1:]:
            out_file2.write(str(line_in) + "\n")
    elif dict_in[j][0] > 2:
        for line_in in dict_in[j][1:]:
            out_file3.write(str(line_in) + "\n")


out_file.close()
out_file2.close()
out_file3.close()

我在具有 8GB 内存的 Windows PC i7 上运行它,这应该不会花费数小时来执行。这是我将文件读入列表的方式有问题吗?我应该使用不同的方法吗?提前致谢。

【问题讨论】:

  • 在导出频率时是否需要保持出现的顺序?
  • 顺便说一句,filetoliststrip 函数效率有点低,而且浪费内存。您将整个文件读入一个列表,然后创建一个新的剥离行列表。这些列表都不是必需的。只需遍历文件行并在阅读时剥离它们。
  • 我在这里必须同意@PM 2Ring。在我看来,这是瓶颈。而且您有大文件,这意味着所有内容都保存在内存中。另外我建议更好地命名你的函数。 thisismymythicalmethod vs this_is_my_mythical_method
  • 顺便说一句,您for 循环中的缩进似乎不正确。请修复它。

标签: python loops dictionary frequency large-files


【解决方案1】:

第一个函数:

def filetoliststrip(file):
    file_in = str(file)
    lines = list(open(file_in, 'r'))
    content = [x.strip() for x in lines] 
    return content

此处生成的原始行列表仅用于剥离。这将需要大约两倍的内存,同样重要的是,需要多次传递不适合缓存的数据。我们也不需要重复str。所以我们可以稍微简化一下:

def filetoliststrip(filename):
    return [line.strip() for line in open(filename, 'r')]

这仍然会产生一个列表。如果我们只读取一次数据,而不是存储每一行​​,请将[] 替换为() 以将其转换为生成器表达式;在这种情况下,由于行实际上在内存中保持不变,直到程序结束,我们只会为列表节省空间(在您的情况下仍然至少 30MB)。

然后我们有主解析循环(我调整了我认为应该的缩进):

counter = 0

for line in fileinlist:
    counter += 1
    keyf = line[10:69]
    print("Loading line " + str(counter) + " : " + str(line))
    if keyf not in dict_in.keys():
        dict_in[keyf] = []
        dict_in[keyf].append(1)
        dict_in[keyf].append(line)
    else:
        dict_in[keyf][0] += 1
        dict_in[keyf].append(line)

这里有几个次优的事情。

首先,计数器可以是enumerate(如果没有可迭代对象,则有range 或itertools.count)。改变这一点将有助于提高清晰度并降低出错的风险。

for counter, line in enumerate(fileinlist, 1):

其次,在一次操作中形成一个字符串比从位添加它更有效:

    print("Loading line {} : {}".format(counter, line))

第三,不需要为字典成员检查提取键。在 Python 2 中,它构建了一个新列表,这意味着复制键中保存的所有引用,并且每次迭代都会变慢。在 Python 3 中,这仍然意味着不必要地构建一个键视图对象。如果需要检查,只需使用keyf not in dict_in。

第四,确实不需要支票。在查找失败时捕获异常几乎与 if 检查一样快,而在 if 检查之后重复查找几乎肯定会更慢。就此而言,请停止重复查找:

    try:
        dictvalue = dict_in[keyf]
        dictvalue[0] += 1
        dictvalue.append(line)
    except KeyError:
        dict_in[keyf] = [1, line]

这是一种常见的模式,但是,我们有 两个 标准库实现它:Counter 和 defaultdict。我们可以在这里使用两者,但是当您只需要计数时,计数器更实用。

from collections import defaultdict
def newentry():
    return [0]
dict_in = defaultdict(newentry)

for counter, line in enumerate(fileinlist, 1):
    keyf = line[10:69]
    print("Loading line {} : {}".format(counter, line))
    dictvalue = dict_in[keyf]
    dictvalue[0] += 1
    dictvalue.append(line)

使用defaultdict 让我们不必担心条目是否存在。

我们现在到达输出阶段。再次,我们有不必要的查找,所以让我们将它们减少到一次迭代:

for key, value in dict_in.iteritems():  # just items() in Python 3
    print("Processing key: " + key)
    #print(value)
    count, lines = value[0], value[1:]
    if count < 2:
        out_file.write(lines[0])
    elif count == 2:
        for line_in in lines:
            out_file2.write(line_in + "\n")
    elif count > 2:
        for line_in in lines:
            out_file3.write(line_in + "\n")

这仍然有一些烦恼。我们重复了编写代码,它构建了其他字符串(在"\n" 上标记),并且每种情况都有一大堆类似的代码。事实上,重复可能导致了一个错误:out_file 中的单次出现没有换行符分隔符。让我们找出真正的不同之处:

for key, value in dict_in.iteritems():  # just items() in Python 3
    print("Processing key: " + key)
    #print(value)
    count, lines = value[0], value[1:]
    if count < 2:
        key_outf = out_file
    elif count == 2:
        key_outf = out_file2
    else:  #  elif count > 2:  # Test not needed
        key_outf = out_file3
    key_outf.writelines(line_in + "\n" for line_in in lines)

我已经离开了换行连接,因为将它们作为单独的调用混合起来更加复杂。该字符串是短暂的,它的目的是将换行符放在同一个位置:它使得在操作系统级别上一行被并发写入中断的可能性降低。

您会注意到这里有 Python 2 和 Python 3 的区别。如果首先在 Python 3 中运行,您的代码很可能并没有那么慢。存在一个名为six 的兼容性模块,用于编写更容易在其中任何一个中运行的代码;它可以让你使用例如six.viewkeys 和 six.iteritems 来避免这个问题。

【讨论】:

    【解决方案2】:

    您有多个点会减慢您的代码 - 无需将整个文件加载到内存中只是为了再次对其进行迭代,也无需在每次要进行查找时获取键列表( if key not in dict_in: ... 就足够了,而且速度非常快),您不需要保留行数,因为无论如何您都可以事后检查行长......仅举几例。

    我会将您的代码完全重组为:

    import collections
    
    dict_in = collections.defaultdict(list)  # save some time with a dictionary factory
    with open(file_in, "r") as f:  # open the file_in for reading
        for line in file_in:  # read the file line by line
            key = line.strip()[10:69]  # assuming this is how you get your key
            dict_in[key].append(line)  # add the line as an element of the found key
    # now that we have the lines in their own key brackets, lets write them based on frequency
    with open(file_ot, "w") as f1, open(file_ot2, "w") as f2, open(file_ot3, "w") as f3:
        selector = {1: f1, 2: f2}  # make our life easier with a quick length-based lookup
        for values in dict_in.values():  # use dict_in.itervalues() on Python 2.x
            selector.get(len(values), f3).writelines(values)  # write the collected lines
    

    而且你几乎不会比这更有效率,至少在 Python 中是这样。

    请记住,这不能保证 Python 3.7(或 CPython 3.6)之前的输出中的行顺序。但是,密钥本身的顺序将被保留。如果您需要在上述 Python 版本之前保留行顺序,则必须保留一个单独的键顺序列表并对其进行迭代以按顺序获取 dict_in 值。

    【讨论】:

    • 非常非常快,还有很多有用的cmets。谢谢!!
    • 非常整洁!记住我们存储所有项目,因此计数是长度,这使得代码更加整洁。
    【解决方案3】:

    您一次将一个非常大的文件加载到内存中。如果您实际上不需要这些行,而您只需要处理它,请使用 generator。它更节省内存。

    Counter 是一个集合,其中元素存储为字典键,其计数存储为字典值。您可以使用它来计算键的频率。然后只需遍历新的dict 并将密钥附加到相关文件:

    from collections import Counter
    
    keys = ['A1KEY1', 'A2KEY1', 'B1KEY2', 'C1KEY3', 'D1KEY3', 'E1KEY4']
    count = Counter(keys)
    
    
    with open('single.txt') as f1:
        with open('double.txt') as f2:
            with open('more_than_double.txt') as f3:
    
            for k, v in count.items():
                if v == 1:
                    f1.writelines(k)
                elif v == 2:
                    f2.writelines(k)
                else:
                    f3.writelines(k)
    

    【讨论】:

      猜你喜欢
      • 2015-09-21
      • 1970-01-01
      • 1970-01-01
      • 2021-07-01
      • 2019-04-21
      • 2016-06-04
      • 2015-01-04
      • 2013-12-03
      • 1970-01-01
      相关资源
      最近更新 更多