【发布时间】:2020-11-22 23:42:44
【问题描述】:
我有一个这样的字典:
# I have a list of 300k dictionaries similar in format to the following
# I cannot assume the dictionaries are sorted by the "id" key
sentences = {"id": 1, "some_sentence" :"In general, the performance gains that indexes provide for read operations are worth the insertion penalty.",
"another_sent": "She said, I had a dream you were playing for the Panthers.' I was like, that's weird because I'm in Indianapolis. But life has come full circle and her dream came true.”" }
# the double quotations are not a typo
# the re.findall is meant to split the the sentence into individual words excluding punctuations
temp = []
temp = re.findall(r"[\w']+|[.,!?;]", sentences.get('some_sentence'))
temp += re.findall(r"[\w']+|[.,!?;]", sentences.get('another_sent'))
# temp = ["She", "said", "had",...]
# to delete case-sensitive duplicates so so if "My" and "my" is used in the sentence only "my" is kept
words = list(set(t.lower() for t in temp))
# I need to remove words of length less than 3
for i in words:
if len(i) < 3:
words.remove(i)
# put the list of words back into the dict
sentences["Words"] = words. # O(1)
我有一个包含 300k 字典的列表,现在在我的 mac 上运行大约需要 53 秒 我真的不知道我还能做些什么来缩短时间
我尝试过的事情:
- 我曾尝试使用 enumerate,但速度会慢一些
- 我曾尝试使用 cython 库翻译成 C,但由于无法将“re.findall”翻译成 C,我没有得到足够的改进?
有什么想法吗?
【问题讨论】:
-
我会说“添加更多 RAM”,但因为它是 mac .. 购买新的 mac? (说真的,使用您的本地系统状态监控应用程序来查看您是否正在最大限度地使用 cpu 或 ram)您正在做一些非常密集的事情,它可能只是很慢。
-
你能解释一下你真正想要做什么的吗? “循环通过 300k dicts”不是它。你这样做是为了取得一些成就。说明您想要实现的目标。
-
目前尚不清楚为什么您使用正则表达式
r"[\w']+|[.,!?;]"进行标记只是为了丢弃所有标点符号。为什么不直接从r"[\w']+"开始呢? -
除了上面的人正确指出的:数据如何以字典列表开头?您是否从文件、一些在线资源等开始?因为您可能在这里遇到 XY 问题,最好的解决方案不是您要询问的那个。
-
你只显示一个字典,你有没有建立一个我们在这里看不到的 300k 句子字典的列表?这些都来自不同的文件吗?您可能会加快生成列表的速度。您可能会通过多处理池获得一些加速,但这里的细节是魔鬼。
标签: python performance dictionary cython performance-testing