【发布时间】:2021-06-15 04:35:04
【问题描述】:
我是 python 新手,编写并编写了以下代码:
- 读取文本文件。
- 保存在列表中。
- 执行正则表达式。
- 将列表转换为字符串
- 删除不需要的特殊字符。
- 将内容放回列表中。
- 使用列表中的计数器,然后将它们打包到字典中。
- 最后使用 Pandas 绘制键和值。
如你所知,我的 Python 经验非常低。我的代码非常适合较小的文件,但是当我使用 700 MB 的文件时,它似乎一直在运行!
如何优化我的代码?
这是我的输入文件格式。
74M2S
73M
74M2S
*
73M
75M1S
这是我的代码:
import matplotlib.pyplot as plt
import re
import pandas as pd
from collections import Counter
f = open('/PathTpFile/MyFILE.txt','r+')
listToStr: str
str2: str
mylist1 = []
for line in f.readlines():
mylist1.append([re.findall(r'[\d]+M', line)])
mylist1.sort(reverse=True)
listToStr = ' '.join(map(str, mylist1))
specialChars = "M[]'"
for specialChar in specialChars:
listToStr = listToStr.replace(specialChar, '')
words: list = listToStr.split()
counts = Counter(words)
dict(counts)
print(counts)
f.close()
keys = counts.keys()
values = counts.values()
print(counts.keys())
print(counts.values())
plt.bar(keys, values)
plt.savefig("out.png")
【问题讨论】:
-
为什么要将
findall()结果包装在另一个列表中:[re.findall(r'[\d]+M', line)]? -
为什么使用
[\d]+M而不是\d+M? -
为什么每次循环都重新分配
words?最后做一次。 -
你每次都在通过
for line循环做很多事情,应该在最后完成。 -
如果您的目标只是计算字数,为什么还需要对任何内容进行排序?
标签: regex string list substring counter