【问题标题】:Improving for loop - Trying to compare 2 lists of dicts改进 for 循环 - 尝试比较 2 个 dicts 列表
【发布时间】:2019-10-02 13:38:27
【问题描述】:

我会尽量让自己清楚:我有 50k 条推文我想对其进行文本挖掘,并且我想改进我的代码。数据如下所示 (sample_data)。

我有兴趣对已清理和标记化的单词进行词形还原(这是 twToken 键的值)

sample_data = [{'twAuthor': 'Jean Lassalle',
                'twMedium': 'iPhone',
                'nFav': None,
                'nRT': '33',
                'isRT': True,
                'twText': ' RT @ColPeguyVauvil : @jeanlassalle "allez aux bouts de vos rêves" ',
                'twParty': 'Résistons!',
                'cleanText': ' rt colpeguyvauvil jeanlassalle allez aux bouts de vos rêves ',
                'twToken': ['colpeguyvauvil', 'jeanlassalle', 'allez', 'bouts', 'rêves']},
               {'twAuthor': 'Jean-Luc Mélenchon',
                'twMedium': 'Twitter Web Client',
                'nFav': '806',
                'nRT': '375',
                'isRT': False,
                'twText': ' (2/2) Ils préfèrent créer une nouvelle majorité cohérente plutôt que les alliances à géométrie variable opportunistes de leur direction. ',
                'twParty': 'La France Insoumise',
                'cleanText': ' 2 2 ils préfèrent créer une nouvelle majorité cohérente plutôt que les alliances à géométrie variable opportunistes de leur direction ',
                'twToken': ['2', '2', 'préfèrent', 'créer', 'nouvelle', 'majorité', 'cohérente', 'plutôt', 'alliances', 'géométrie', 'variable', 'opportunistes', 'direction']},
               {'twAuthor': 'Nathalie Arthaud',
                'twMedium': 'Android',
                'nFav': '37',
                'nRT': '24',
                'isRT': False,
                'twText': ' #10mai Commemoration fin de l esclavage. Reste à supprimer l esclavage salarial defendu par #Macron et Hollande ',
                'twParty': 'Lutte Ouvrière',
                'cleanText': ' 10mai commemoration fin de l esclavage reste à supprimer l esclavage salarial defendu par macron et hollande ',
                'twToken': ['10mai', 'commemoration', 'fin', 'esclavage', 'reste', 'supprimer', 'esclavage', 'salarial', 'defendu', 'macron', 'hollande']
               }]

但是,Python 中没有可靠的法语词形还原器。所以我使用了一些资源来拥有我自己的法语单词词法词典。字典看起来像这样:

sample_lemmas = [{"ortho":"rêves","lemme":"rêve","cgram":"NOM"},
                 {"ortho":"opportunistes","lemme":"opportuniste","cgram":"ADJ"},
                 {"ortho":"préfèrent","lemme":"préférer","cgram":"VER"},
                 {"ortho":"nouvelle","lemme":"nouveau","cgram":"ADJ"},
                 {"ortho":"allez","lemme":"aller","cgram":"VER"},
                 {"ortho":"défendu","lemme":"défendre","cgram":"VER"}]

所以ortho是单词的书面形式(例如processed),lemme是单词的词形化形式(例如process) cgram 是单词的语法类别(例如 VER 表示动词)。

所以我想做的是为每条推文创建一个twLemmas 键,这是从twToken 列表派生的引理列表。所以我循环遍历sample_data 中的每条推文,然后循环遍历twToken 中的每个标记,查看标记是否存在于我的引理字典sample_lemmas 中,如果存在,我从sample_lemmas 字典中检索引理并将其添加到将在每个 twLemmas 键中提供的列表中。如果没有,我只需将这个词添加到列表中。

我的代码如下:

list_of_ortho = []                      #List of words used to compare if a token doesn't exist in my lemmas dictionary
for wordDict in sample_lemmas:          #This loop feeds this list with each word
    list_of_ortho.append(wordDict["ortho"])

for elemList in sample_data:            #Here I iterate over each tweet in my data
    list_of_lemmas = []                 #This is the temporary list which will be the value to each twLemmas key
    for token in elemList["twToken"]:   #Here, I iterate over each token/word of a tweet
        for wordDict in sample_lemmas:
            if token == wordDict["ortho"]:
                list_of_lemmas.append(wordDict["lemme"])
        if token not in list_of_ortho:  #And this is to add a word to my list if it doesn't exist in my lemmas dictionary
            list_of_lemmas.append(token)
    elemList["lemmas"] = list_of_lemmas

sample_data

循环运行良好,但大约需要 4 小时才能完成。现在我知道我既不是程序员也不是 Python 专家,而且我知道无论如何都需要时间来完成。然而,这就是为什么我想问你是否有人对如何改进我的代码有更好的想法?

如果有人能花时间理解我的代码并帮助我,谢谢。我希望我说得够清楚(抱歉,英语不是我的母语)。

【问题讨论】:

  • 另外,在 Python 中,我们使用 lower_case_with_underscores 而不是 camelCase 命名函数和变量
  • 您还需要一个查找字典而不是 list_of_ortho。它正在为应该是一次检查的内容做一个大循环。
  • 哇,感谢所有这些 cmets !不幸的是,我现在没有时间尝试,但我今晚会尝试... @KennyOstrom:谢谢,我知道部分原因是我愚蠢地失去了计算性能。 //@鲍里斯:也谢谢你!我忽略了列表和集合之间的区别……关于约定,我被告知实际上相反的 u.u”(有人告诉我下划线约定是针对 R,而不是 Python……)。

标签: python dictionary for-loop lemmatization


【解决方案1】:

使用将正交映射到词根的字典:

ortho_to_lemme = {word_dict["ortho"]: word_dict["lemme"] for word_dict in sample_lemmas}
for tweet in sample_data:
    tweet["twLemmas"] = [
        ortho_to_lemme.get(token, token) for token in tweet["twToken"]
    ]

【讨论】:

  • 这里的关键是在 for 循环之外一次性设置字典,当然还有它的 O(1) 查找。原始代码很慢,因为它循环遍历所有引理,每个标记多次。
  • 我尝试使用我的数据子集,您的回复给了我令人印象深刻的结果:-原始循环经过的时间:30.78663806700206-您的循环经过的时间:0.23488750499745947
  • 编辑:刚刚在整个数据集上运行它,花了我 1.9920109320009942 秒而不是 ~3.5 小时!非常感谢!
  • 感谢您如此巧妙地提出您的问题并提供带有示例数据的可运行代码。
猜你喜欢
  • 2013-07-18
  • 2020-08-17
  • 2013-08-15
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-08-18
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多