【问题标题】:Split a list of sentence to words and append them to a dict将句子列表拆分为单词并将它们附加到字典
【发布时间】:2021-03-29 14:48:31
【问题描述】:

我正在尝试将句子拆分为单词并将每个单词作为句子的值。所以:

{sent 1: ["hi","how"], sent 2:["hello","i'm"]}

这是我的代码:

def get_maps(corpus):
vocab= {}
ind = len(corpus)
for line in corpus.values:
    for line[0] in line:
        print(line[0])
        for i in range(ind):
            sents = []
       
            
            vocab[i].append(line[0].lower().split())
        

return vocab

但此代码返回一个关键错误。当我只是等同于

vocab[i] = (line[0].lower().split())

我为每个键得到相同的值。

我的语料库是句子的数据库。

【问题讨论】:

  • 如果您可以在问题中添加几行dataframe 会更好..

标签: python dictionary


【解决方案1】:

在您的解决方案中,您正在用多个 for 循环覆盖前面的句子。这是一个只有一个循环的工作示例:

import pandas as pd
sents = ["hi my name is retroflake", "this is a sample test case"]
df=pd.DataFrame(sents)
d={}
counter = 0
for sent in df.values:
    counter +=1
    d['sent {}'.format(counter)]=list(sent[0].split())
print(d)

这是上面 sn-p 的输出:

{'sent 1': ['hi', 'my', 'name', 'is', 'retroflake'], 'sent 2': ['this', 'is', 'a', '样本','测试','案例']}

【讨论】:

    【解决方案2】:

    您不能追加到您的字典中不存在的列表。 您必须使用 defaultdict,将 list 关键字作为 default_dict 参数传入。 或者你必须在 try-except 语句块中使用普通的字典

    import collections
    def get_maps1(corpus):
        vocab=collections.defaultdict(list)
    
        ind = len(corpus)
        for line in list(corpus.values()):
            for line[0] in line:
    
                for i in range(ind):
                    sents = []
    
    
                    vocab[i].append(line[0].lower().split())
    

    检查第 4 条语句。您必须使 dict value 可迭​​代 inroder 循环。只需调用 values 将仅返回内置方法 'values'

    如果你想使用普通的字典,你可以使用 try-except 块

    d1={1: ["hi","how"], 2:["hello","i'm"]}
    vocab={}
    for line in list(d1.values()):
            ind=len(d1)
            for line[0] in line:
                print(line[0])
                for i in range(ind):
                    sents = []
    
                    try:
                        vocab[i].append(line[0].lower().split())
    
                    except:
                        #since key is absent a new key is created for first found list item
                        vocab[i]=[line[0].lower().split()]
    
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-11-16
      • 1970-01-01
      • 2023-03-24
      • 2021-09-09
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多