【问题标题】:How can I turn pandas dataframe into an ordered list with many to one relationship?如何将 pandas 数据框转换为具有多对一关系的有序列表?
【发布时间】:2017-03-08 05:27:55
【问题描述】:

我目前有一个 pandas 数据框,其中有许多答案加入到一个问题上,所以我试图将它变成一个列表,以便我可以做余弦相似度。

目前我有dataframe,其中问题通过parent_id = q_id与答案相连,如图:

many answers to one question dataframe

print (df)
   q_id      q_body  parent_id    a_body
0     1  question 1          1  answer 1
1     1  question 1          1  answer 2
2     1  question 1          1  answer 3
3     2  question 2          2  answer 1
4     2  question 2          2  answer 2

我正在寻找的产品是:

(“问题 1”、“答案 1”、“答案 2”、“答案 3”)

(“问题 2”、“答案 1”、“答案 2”)

任何帮助将不胜感激!非常感谢你。

【问题讨论】:

    标签: python list pandas many-to-one


    【解决方案1】:

    我认为你需要groupby 和apply:

    #output is tuple with question value
    df = df.groupby('q_body')['a_body'].apply(lambda x: tuple([x.name] + list(x)))
    print (df)
    q_body
    question 1    (question 1, answer 1, answer 2, answer 3)
    question 2              (question 2, answer 1, answer 2)
    Name: a_body, dtype: object
    
    #output is list with question value
    df = df.groupby('q_body')['a_body'].apply(lambda x: [x.name] + list(x))
    print (df)
    q_body
    question 1    [question 1, answer 1, answer 2, answer 3]
    question 2              [question 2, answer 1, answer 2]
    Name: a_body, dtype: object
    
    #output is list without question value
    df = df.groupby('q_body')['a_body'].apply(list)
    print (df)
    q_body
    question 1    [answer 1, answer 2, answer 3]
    question 2              [answer 1, answer 2]
    Name: a_body, dtype: object
    
    #grouping by parent_id without question value
    df = df.groupby('parent_id')['a_body'].apply(list)
    print (df)
    parent_id
    1    [answer 1, answer 2, answer 3]
    2              [answer 1, answer 2]
    Name: a_body, dtype: object
    
    #output is string, values are concanecated by ,
    df = df.groupby('parent_id')['a_body'].apply(', '.join)
    print (df)
    parent_id
    1    answer 1, answer 2, answer 3
    2              answer 1, answer 2
    Name: a_body, dtype: object
    

    但如果需要输出为列表添加tolist:

    L = df.groupby('q_body')['a_body'].apply(lambda x: tuple([x.name] + list(x))).tolist()
    print (L)
    [('question 1', 'answer 1', 'answer 2', 'answer 3'), ('question 2', 'answer 1', 'answer 2')]
    

    【讨论】:

    • 谢谢jezrael,现在会更多地使用lambda。
    • 很高兴能为您提供帮助。美好的一天。
    【解决方案2】:
    df = pd.DataFrame([
            ['question 1', 'answer 1'],
            ['question 1', 'answer 2'],
            ['question 1', 'answer 3'],
            ['question 2', 'answer 1'],
            ['question 2', 'answer 2'],
        ], columns=['q_body', 'a_body'])
    
    print(df)
    
           q_body    a_body
    0  question 1  answer 1
    1  question 1  answer 2
    2  question 1  answer 3
    3  question 2  answer 1
    4  question 2  answer 2
    

    apply(list)

    df.groupby('q_body').a_body.apply(list)
    
    q_body
    question 1    [answer 1, answer 2, answer 3]
    question 2              [answer 1, answer 2]
    

    【讨论】:

      【解决方案3】:

      看看对你有没有帮助

      result = df.groupby('q_id').agg({'q_body': lambda x: x.iloc[0], 'a_body': lambda x: ', '.join(x)})
      result['output'] = result.q_body + ', ' + result.a_body                                                                                
      

      这将创建一个具有所需结果的新列输出。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2022-01-07
        • 2021-11-17
        • 2021-01-14
        • 1970-01-01
        • 2014-04-04
        相关资源
        最近更新 更多