【问题标题】:Spell Correction with Python (pyspellchecker)使用 Python 进行拼写纠正 (pyspellchecker)
【发布时间】:2022-01-14 20:26:21
【问题描述】:

我想使用 python 构建拼写校正,我尝试使用 pyspellchecker,因为我必须构建自己的字典,而且我认为 pyspellchecker 很容易与我们自己的模型或字典一起使用。我的问题是,如何在 case_sensitive 为 On 的情况下加载和返回我的单词? 我试过这个:

spell = SpellChecker(language=None, case_sensitive=True)

但是当我加载我的文件时,我的文件中包含许多文本,例如 'Hello' 的代码:

spell.word_frequency.load_text_file('myfile.txt')

当我开始用spell.correction('Hello') 拼写时,它的返回'hello'(小写)。 你知道如何构建我们自己的模型或字典,而我们的字母不会减少或保持大写吗?

或者如果您有使用我们自己的模型进行拼写检查的建议,请告诉我,谢谢!

【问题讨论】:

    标签: python nlp spell-checking


    【解决方案1】:

    试试这个:

    from spellchecker import SpellChecker
    
    spell = SpellChecker(language=None, case_sensitive=True)
    a = spell.word_frequency.load_words(["Hello", "HELLO", "I", "AM", "Alok", "Mishra"])
    
    # find those words that may be misspelled
    misspelled = spell.unknown(["helo", "Alk", "Mishr"])
    
    for word in misspelled:
        # Get the one `most likely` answer
        print(spell.correction(word))
    
        # Get a list of `likely` options
        print(spell.candidates(word))
    

    输出:

    Alok
    {'Alok'}
    Hello
    {'Hello'}
    Mishra
    {'Mishra'}
    

    【讨论】:

    • 您好,感谢您的解决方案。有效!但是,如果我们要保存我们的模型,然后加载我们之前保存的模型,我的单词会被 pyspellchecker 自动变为小写。我使用spell.export('my_custom_dictionary.gz', gzipped=True) 导出我的模型并使用spell.word_frequency.load_text_file('./path-to-my-model.gz') 加载我的模型
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-02-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2012-05-30
    相关资源
    最近更新 更多