【发布时间】:2021-11-28 00:34:19
【问题描述】:
我正在按照这里的教程:https://www.analyticsvidhya.com/blog/2019/08/comprehensive-guide-language-model-nlp-python-code/#h2_5 创建语言模型。我正在关注 N-gram 语言模型。
这是完成的代码:
from nltk.corpus import reuters
from nltk import bigrams, trigrams
from collections import Counter, defaultdict
# Create a placeholder for model
model = defaultdict(lambda: defaultdict(lambda: 0))
# Count frequency of co-occurance
for sentence in reuters.sents():
for w1, w2, w3 in trigrams(sentence, pad_right=True, pad_left=True):
model[(w1, w2)][w3] += 1
# Let's transform the counts to probabilities
for w1_w2 in model:
total_count = float(sum(model[w1_w2].values()))
for w3 in model[w1_w2]:
model[w1_w2][w3] /= total_count
input = input("Hi there! Please enter an incomplete sentence and I can help you\
finish it!\n").lower().split()
print(model[tuple(input)])
为了从模型中获取输出,网站这样做:print(dict(model["the", "price"])) 但我想从用户输入的句子中生成输出。当我写print(model[tuple(input)]) 时,它给了我一个空的defaultdict。
忽略这个(留作历史):
如何给它我从输入创建的列表?
model是一个 字典,我读过使用列表作为键不是一个好主意 但这正是他们正在做的事情?我假设我的没有 工作,因为我列出一个列表?我是否必须遍历 单词得到结果?作为旁注,这个模型是否将句子作为一个整体来考虑 预测下一个词,还是只预测最后一个个词?
【问题讨论】:
-
1.字典使用的是元组,而不是列表(请参阅
model[(w1, w2)][w3]...。2. 从对trigrams的调用中,我只能得出结论,它使用三元组,即:计算一个单词出现前两个单词的概率。 -
@JuanR 我完全错过了它使用元组!是的,trigram 也会暗示这一点。感谢您指出所有这些!
标签: python dictionary prediction n-gram language-model