【问题标题】:The fastest way to loop through a list of 300k dictionaries?循环浏览 300k 字典列表的最快方法是什么?
【发布时间】:2020-11-22 23:42:44
【问题描述】:

我有一个这样的字典:

# I have a list of 300k dictionaries similar in format to the following
# I cannot assume the dictionaries are sorted by the "id" key
sentences = {"id": 1, "some_sentence" :"In general, the performance gains that indexes provide for read operations are worth the insertion penalty.",
             "another_sent": "She said, I had a dream you were playing for the Panthers.' I was like, that's weird because I'm in Indianapolis. But life has come full circle and her dream came true.”" }
# the double quotations are not a typo


# the re.findall is meant to split the the sentence into individual words excluding punctuations
temp = []
temp = re.findall(r"[\w']+|[.,!?;]", sentences.get('some_sentence'))
temp += re.findall(r"[\w']+|[.,!?;]", sentences.get('another_sent'))
# temp = ["She", "said", "had",...]

# to delete case-sensitive duplicates so so if "My" and "my" is used in the sentence only "my" is kept
words = list(set(t.lower() for t in temp))

# I need to remove words of length less than 3
for i in words:
    if len(i) < 3:
       words.remove(i)

# put the list of words back into the dict

sentences["Words"] = words. # O(1)

我有一个包含 300k 字典的列表,现在在我的 mac 上运行大约需要 53 秒 我真的不知道我还能做些什么来缩短时间

我尝试过的事情:

  • 我曾尝试使用 enumerate,但速度会慢一些
  • 我曾尝试使用 cython 库翻译成 C,但由于无法将“re.findall”翻译成 C,我没有得到足够的改进?

有什么想法吗?

【问题讨论】:

  • 我会说“添加更多 RAM”,但因为它是 mac .. 购买新的 mac? (说真的,使用您的本地系统状态监控应用程序来查看您是否正在最大限度地使用 cpu 或 ram)您正在做一些非常密集的事情,它可能只是很慢。
  • 你能解释一下你真正想要做什么的吗? “循环通过 300k dicts”不是它。你这样做是为了取得一些成就。说明您想要实现的目标。
  • 目前尚不清楚为什么您使用正则表达式 r"[\w']+|[.,!?;]" 进行标记只是为了丢弃所有标点符号。为什么不直接从r"[\w']+" 开始呢?
  • 除了上面的人正确指出的:数据如何以字典列表开头?您是否从文件、一些在线资源等开始?因为您可能在这里遇到 XY 问题,最好的解决方案不是您要询问的那个。
  • 你只显示一个字典,你有没有建立一个我们在这里看不到的 300k 句子字典的列表?这些都来自不同的文件吗?您可能会加快生成列表的速度。您可能会通过多处理池获得一些加速,但这里的细节是魔鬼。

标签: python performance dictionary cython performance-testing


【解决方案1】:

这可能会缩短执行时间,并使自己免受麻烦,因为修改您正在迭代的列表不是一个好主意 - 如果没有分析,很难说改进是微不足道的还是显着的:

替换:

words = list(set(t.lower() for t in temp))
for i in words:
    if len(i) < 3:
       words.remove(i)   # this is an expensive operation on longer lists

与:

words = [word for word in set(t.lower() for t in temp) if len(word) > 3]

【讨论】:

  • 似乎您可以将其添加到list(set(t.lower() for t in temp)) 以仅循环一次。
【解决方案2】:

您可以使用多个进程获得加速。由于您在 Mac 上,因此子进程可以查看父内存。通过将句子列表设置为全局变量并仅将其索引传递给子进程,您就有了一种合理精简的方式来获取子进程的数据。尽管如此,生成的单词列表需要传递回父级,这可能会否定池的优势。此方法不适用于在子进程空间中看不到全局变量的 Windows。

import multiprocessing as mp
import re

# I have a list of 300k dictionaries similar in format to the following
# I cannot assume the dictionaries are sorted by the "id" key
sentences = {"id": 1, "some_sentence" :"In general, the performance gains that indexes provide for read operations are worth the insertion penalty.",
                 "another_sent": "She said, I had a dream you were playing for the Panthers.' I was like, that's weird because I'm in Indianapolis. But life has come full circle and her dream came true.”" }
# the double quotations are not a typo

# imagine this is the 300k list of dicts
sentence_list = [sentences]

def worker(index):
    sentences = sentence_list[index]
    # the re.findall is meant to split the the sentence into individual words excluding punctuations
    temp = []
    temp = re.findall(r"[\w']+|[.,!?;]", sentences.get('some_sentence'))
    temp += re.findall(r"[\w']+|[.,!?;]", sentences.get('another_sent'))
    # temp = ["She", "said", "had",...]

    # to delete case-sensitive duplicates so so if "My" and "my" is used in the sentence only "my" is kept
    words = list(set(t.lower() for t in temp))

    # I need to remove words of length less than 3
    for i in words:
        if len(i) < 3:
            words.remove(i)

    return index, words

with mp.Pool() as pool:
    for index, words in pool.imap_unordered(worker, range(len(sentence_list))):
        sentence_list[index]["word"] = words

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-01-28
    • 2021-02-17
    • 1970-01-01
    • 2016-08-15
    • 2015-07-25
    相关资源
    最近更新 更多