【问题标题】:python - How to concatenate consecutive rows without losing the order of the records in pandas?python - 如何在不丢失熊猫记录顺序的情况下连接连续行?
【发布时间】:2020-04-11 02:15:08
【问题描述】:

当连续行中变量的类别相同时,我试图根据分类变量连接两行。下面是我的数据,例如:

   SNo  user    Text
0   1   Sam     Hello
1   1   John    Hi
2   1   Sam     How are you?
3   1   John    I am good
4   1   John    How about you?
5   1   John    How is it going?
6   1   Sam     Its going good
7   1   Sam     Thanks
8   2   Mary    Howdy?
9   2   Jake    Hey!!
10  2   Jake    What a surprise
11  2   Mary    Good to see you here :)
12  2   Jake    Ha ha. Hectic life
13  2   Mary    I know right..
14  2   Mary    How's Amy doing?
15  2   Mary    How are the kids?
16  2   Jake    All is good! :)

在这里,如果我之前的 user 列值与我当前的 user 列值相同,但与该列中的下一个值不同,那么,我将连接列 Text 中的值那个用户。我需要这样做,直到该特定用户不再多次出现。下面给出了一个示例输出:

SNo user    Text
1   Sam     Hello
1   John    Hi
1   Sam     How are you?
1   John    I am good-How about you?-How is it going?
1   Sam     Its going good-Thanks
2   Mary    Howdy?
2   Jake    Hey!!-What a surprise
2   Mary    Good to see you here :)
2   Jake    Ha ha. Hectic life
2   Mary    I know right..-How's Amy doing?-How are the kids?
2   Jake    All is good! :)

我尝试使用df.groupby() 然后.agg() 完成连接,但无法对其应用上述条件。因此,输出将组合所有出现的用户进行聊天。

df = sample_data.groupby(["SNo","user"]).agg({'Text': '-'.join}).reset_index() # incorrect though
df

此外,我试图避免像瘟疫一样的for 循环并尝试矢量化解决方案。


样本数据:

data_dict = {'S. No.': [1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2], 'user': ['Sam', 'John', 'Sam', 'John', 'John', 'John', 'Sam', 'Sam', 'Mary', 'Jake', 'Jake', 'Mary', 'Jake ', 'Mary', 'Mary', 'Mary', 'Jake'], 'Text': ['Hello', 'Hi', 'How are you?', 'I am good', 'How about you?', 'How is it going?', 'Its going good', 'Thanks', 'Howdy?', 'Hey!!', 'What a surprise', 'Good to see you here :)', 'Ha ha. Hectic life', 'I know right..', "How's Amy doing?", 'How are the kids?', 'All is good! :)']}

sample_data = pd.DataFrame(data_dict)

【问题讨论】:

    标签: python pandas dataframe concatenation


    【解决方案1】:

    您希望将user 与其shift 和cumsum 进行比较以进行更改。然后你可以分组:

    blocks = df['user'].ne(df['user'].shift()).cumsum()
    (df.groupby(['SNo', blocks])
      .agg({'user':'first','Text': '-'.join})
      .reset_index('user', drop=True)
    )
    

    输出:

         user                                               Text
    SNo                                                         
    1     Sam                                              Hello
    1    John                                                 Hi
    1     Sam                                       How are you?
    1    John          I am good-How about you?-How is it going?
    1     Sam                              Its going good-Thanks
    2    Mary                                             Howdy?
    2    Jake                              Hey!!-What a surprise
    2    Mary                            Good to see you here :)
    2    Jake                                 Ha ha. Hectic life
    2    Mary  I know right..-How's Amy doing?-How are the kids?
    2    Jake                                    All is good! :)
    

    【讨论】:

    • 第一次看到.shift()属性!
    • 我有一个类似的问题,但我必须合并至少 40 列的文本,就像您对一个名为“文本”的列所做的那样。我怎样才能做到这一点?我已经尝试过上述解决方案,但它可以用于单个列。
    猜你喜欢
    • 2020-12-27
    • 1970-01-01
    • 2012-05-30
    • 1970-01-01
    • 2021-09-12
    • 1970-01-01
    • 2015-01-06
    • 1970-01-01
    • 2013-12-02
    相关资源
    最近更新 更多