【问题标题】:Group by data to dataframe columns with pandas使用熊猫按数据分组到数据框列
【发布时间】:2017-02-15 12:11:35
【问题描述】:

我有以下数据框。描述每个用户居住的城市

       City     Name    Date
0   Seattle    Alice    2017
1   Seattle      Bob    2011
2  Portland  Mallory    2010
3   Seattle  Mallory    2016
4   Memphis      Bob    2012
5  Portland  Mallory    2013

你可以通过 pandas 实现以下目标吗?

     Name     City1    Date1   City2   Date2   City3    Date3
0   Alice     Seattle  2017    NaN     NaN     NaN      NaN
1   Bob       Seattle  2011    Memphis 2012    NaN      NaN
2   Mallory   Portland 2010    Seattle 2016    Portland 2013

非常感谢!

【问题讨论】:

    标签: python dataframe bigdata


    【解决方案1】:

    您可以将groupby 与自定义函数一起使用,其中创建新的DataFrame,然后unstack,将MultiIndex 的第二级按sort_index 排序,最后使用join 将其删除:

    df1 = df.groupby('Name')['City','Date']
            .apply(lambda x: pd.DataFrame(x.values, 
                                          columns = ['City','Date'], 
                                          index = np.arange(1, len(x) + 1).astype(str)))
            .unstack()
    df1 = df1.sort_index(axis=1, level=1).replace({None:np.nan})
    df1.columns = df1.columns.map(''.join)
    print (df1)
                City1  Date1    City2   Date2     City3   Date3
    Name                                                       
    Alice     Seattle   2017      NaN     NaN       NaN     NaN
    Bob       Seattle   2011  Memphis  2012.0       NaN     NaN
    Mallory  Portland   2010  Seattle  2016.0  Portland  2013.0
    

    【讨论】:

      猜你喜欢
      • 2017-07-08
      • 1970-01-01
      • 2018-07-19
      • 2022-01-12
      • 2013-02-28
      • 2019-10-14
      • 1970-01-01
      • 1970-01-01
      • 2016-06-04
      相关资源
      最近更新 更多