【问题标题】:How to use pyspellchecker to autocorrect spelling errors in a pandas column?如何使用 pyspellchecker 自动更正 pandas 列中的拼写错误?
【发布时间】:2023-02-11 05:15:03
【问题描述】:

我有以下数据框:

df = pd.DataFrame({'id':[1,2,3],'text':['a foox juumped ovr the gate','teh car wsa bllue','why so srious']})

我想使用 pyspellchecker 库生成一个包含固定拼写错误的新列。

我尝试了以下但没有更正任何拼写错误:

import pandas as pd
from spellchecker import SpellChecker

spell = SpellChecker()

def correct_spelling(word):
    corrected_word = spell.correction(word)
    if corrected_word is not None:
        return corrected_word
    else:
        return word

df['corrected_text'] = df['text'].apply(correct_spelling)

下面是预期输出的数据框

pd.DataFrame({'id':[1,2,3],'text':['a foox juumped ovr the gate','teh car wsa bllue','why so srious'],
              'corrected_text':['a fox jumped over the gate','the car was blue','why so serious']})

【问题讨论】:

  • 您将整个短语(多个单词)传递给 correction() 函数,而它接受一个单词。
  • 不要写“没有工作”的问题。相反,显示或描述您获得的结果。另外,尝试阅读How to debug small programs。

标签: python pandas spell-checking autocorrect pyspellchecker


【解决方案1】:

我对这个包一无所知(如何修复准确性),但您可以将每一行中的字符串拆分成一个列表,然后遍历列表列表。此示例使用列表理解:

df["text"] = [[spell.correction(word) for word in row] for row in df["text"].str.split(" ").to_list()]
df["text"] = df["text"].apply(lambda x: " ".join(x))

输出(如您所见,您需要提高准确性):

   id                       text
0   1  a food jumped or the gate
1   2           the car was blue
2   3             why so serious

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-04-29
    • 2011-11-10
    • 1970-01-01
    • 1970-01-01
    • 2012-09-09
    • 1970-01-01
    相关资源
    最近更新 更多