【问题标题】:"Tricky" grouping in DataFrameDataFrame 中的“棘手”分组
【发布时间】:2021-09-03 23:46:17
【问题描述】:

有熊猫数据框为:

number      Num_1   Num_2   Num_3   col_1   col_2   col_3   col_4
915         14      5       8       1       0       1       15:46:14.826872
915         14      5       8       1       1       1                       
915         14      5       8       1       1       0                           
787         12      6       8       1       0       1       15:46:36.168584 
787         12      6       8       1       0       0

Num_1,Num_2,Num_3 - 是字符串值,对于具有相同“数字”的行是相同的 col_1,col_2,col_3 - 是布尔值。 col_4 - 字符串值

如何对该结果进行一些“分组”:

number      Num_1   Num_2   Num_3   col_1   col_2   col_3   col_4
915         14      5       8       1       1       1       15:46:14.826872
787         12      6       8       1               1       15:46:36.168584 

对于 col_1,col_2,col_3 我需要布尔“或”

【问题讨论】:

  • 您好,请让问题更清楚,并生成您使用的示例代码。另外,请查看fillna,因为对空字段进行分组可能会失败。

标签: pandas dataframe


【解决方案1】:

由于col_1col_2col_3只是0和1,所以可以使用maxagg函数实现OR条件:

df.groupby(['number', 'Num_1', 'Num_2', 'Num_3'], as_index=False).max()

   number  Num_1  Num_2  Num_3  col_1  col_2  col_3            col_4
0     787     12      6      8      1      0      1  15:46:36.168584
1     915     14      5      8      1      1      1  15:46:14.826872

或者:

df.groupby(['number', 'Num_1', 'Num_2', 'Num_3'], as_index=False).agg('max')

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-03-04
    • 2010-11-18
    • 1970-01-01
    相关资源
    最近更新 更多