【发布时间】:2019-11-25 03:14:46
【问题描述】:
我有一个包含 +- 130k 条推文的数据框,旁边有一个标签(1 = 正面,0 = 负面)。从这个数据框中,我想提取与电影相关的推文。为此,我想出了一个与电影相关的单词列表:
movie_related_words = ["movie", "movies", "watch",
"watching", "film", "cinema",
"actor", "video", "thriller",
"horror", "dvd", "bluray", "soundtrack",
"director", "remake", "blockbuster"]
经过一些预处理后,数据框中的推文被标记化,因此我的数据框的文本列包含 推文列表,其中每个单词都是一个单独的列表元素。供您参考,请在下面找到我的数据框的三个随机元素:
[well, time, for, bed, 500, am, comes, early, nice, chatting, with, everyone, have, a, good, evening, and, rest, of, the, weekend, whats, left, of, it]
[tekkah, defyingsantafe, umm, dont, forget, that, youre, all, gay, socialist, atheists]
[s, mom, nearly, got, ran, over, by, a, truck, on, her, bike, and, dropped, her, work, bag, with, all, her, information, which, was, then, stolen, fb]
我想过滤推文,当给定推文的任何字词(因此是列表的元素)在movie_related_words列表,我想保留那个观察,如果没有,我想丢弃它。
我尝试过像这样应用 lambda 表达式:
def filter_movies(text):
movie_filtered = "".join([i for i in text if i in movie_related_words])
return movie_filtered
twitter_loaded_df['text'] = twitter_loaded_df['text'].apply(lambda x : filter_movies(x))
但这给了我一个奇怪的结果。任何有关如何实现这一目标的指导将不胜感激。一种pythonic /有效的方式将导致我对你的永恒爱。我希望为此目的存在某种熊猫功能,但我还没有找到它......
【问题讨论】:
标签: python pandas dataframe text