【发布时间】:2017-04-21 10:52:58
【问题描述】:
我有一个包含文章列表的熊猫数据框;出口、发布日期、链接等。此数据框中的一列是关键字列表。例如,在关键字列中,每个单元格都包含一个类似 [drop, right, states, laws] 的列表。
我的最终目标是计算每天每个唯一单词的出现次数。我面临的挑战是将关键字从列表中分离出来,然后将它们与它们发生的日期相匹配。 ...假设这甚至是最合乎逻辑的第一步。
目前我在下面的代码中有一个解决方案,但是我是 python 新手,在思考这些事情时,我仍然以 Excel 思维方式思考。下面的代码有效,但速度很慢。有没有快速的方法来做到这一点?
# Create a list of the keywords for articles in the last 30 days to determine their quantity
keyword_list = stories_full_recent_df['Keywords'].tolist()
keyword_list = [item for sublist in keyword_list for item in sublist]
# Create a blank dataframe and new iterator to write the keyword appearances to
wordtrends_df = pd.DataFrame(columns=['Captured_Date', 'Brand' , 'Coverage' ,'Keyword'])
r = 0
print("Creating table on keywords: {:,}".format(len(keyword_list)))
print(time.strftime("%H:%M:%S"))
# Write the keywords out into their own rows with the dates and origins in which they occur
while r <= len(keyword_list):
for i in stories_full_recent_df.index:
words = stories_full_recent_df.loc[i]['Keywords']
for word in words:
wordtrends_df.loc[r] = [stories_full_recent_df.loc[i]['Captured_Date'], stories_full_recent_df.loc[i]['Brand'],
stories_full_recent_df.loc[i]['Coverage'], word]
r += 1
print(time.strftime("%H:%M:%S"))
print("Keyword compilation complete.")
一旦我在它自己的行上找到了每个单词,我就可以简单地使用 .groupby() 来计算每天出现的次数。
# Group and count the keywords and days to find the day with the least of each word
test_min = wordtrends_df.groupby(('Keyword', 'Captured_Date'), as_index=False).count().sort_values(by=['Keyword','Brand'], ascending=True)
keyword_min = test_min.groupby(['Keyword'], as_index=False).first()
目前这个列表中大约有 100,000 个单词,我需要一个小时来浏览该列表。我很想有一个更快的方法来做这件事。
【问题讨论】:
标签: python pandas count frequency