【问题标题】:How to groupby().transform() to value_counts() in pandas?如何在熊猫中将 groupby().transform() 转换为 value_counts()?
【发布时间】:2018-06-02 14:02:05
【问题描述】:

我正在处理带有商品价格的 pandas 数据框 df1。

  Item    Price  Minimum Most_Common_Price
0 Coffee  1      1       2
1 Coffee  2      1       2
2 Coffee  2      1       2
3 Tea     3      3       4
4 Tea     4      3       4
5 Tea     4      3       4

我创建Minimum 使用:

df1["Minimum"] = df1.groupby(["Item"])['Price'].transform(min)

如何创建Most_Common_Price?

df1["Minimum"] = df1.groupby(["Item"])['Price'].transform(value_counts()) # Doesn't work

目前,我使用多步骤方法:

for item in df1.Item.unique().tolist(): # Pseudocode
 df1 = df1[df1.Price == Item]           # Pseudocode
 df1.Price.value_counts().max()         # Pseudocode

这是矫枉过正。一定有更简单的方法,最好是一行

如何在 pandas 中 groupby().transform() 到 value_counts()?

【问题讨论】:

    标签: python pandas dataframe group-by pandas-groupby


    【解决方案1】:

    您可以将groupby + transform 与value_counts 和idxmax 一起使用。

    df['Most_Common_Price'] = (
        df.groupby('Item')['Price'].transform(lambda x: x.value_counts().idxmax()))
    
    df
    
         Item  Price  Minimum  Most_Common_Price
    0  Coffee      1        1                  2
    1  Coffee      2        1                  2
    2  Coffee      2        1                  2
    3     Tea      3        3                  4
    4     Tea      4        3                  4
    5     Tea      4        3                  4
    

    一项改进涉及使用pd.Series.map,

    # Thanks, Vaishali!
    df['Item'] = (df['Item'].map(df.groupby('Item')['Price']
                            .agg(lambda x: x.value_counts().idxmax()))
    df
    
         Item  Price  Minimum  Most_Common_Price
    0  Coffee      1        1                  2
    1  Coffee      2        1                  2
    2  Coffee      2        1                  2
    3     Tea      3        3                  4
    4     Tea      4        3                  4
    5     Tea      4        3                  4
    

    【讨论】:

    • @sudonym 注意,这个方法也适用于对象:-)
    • @Wen 谢谢,这是我没有想到的重要考虑因素!
    • @sudonym 例如,如果Price 是一列字符串,并且您想找到每个组中计数最高的字符串,这仍然可以工作。而mode 仅适用于数字。
    • @sudonym pandas 对象的类型类似于numpy,它们可以包含原始类型的数组,例如numpy.int64 或 dtype=object,在这种情况下,它们可以包含 any python 对象。请注意,这通常是以牺牲效率为代价的。
    • 使用地图代替变换,性能进一步提高。 df['Item'].map(df.groupby('Item').Price.agg(lambda x: x.value_counts().idxmax()))
    【解决方案2】:

    如果您想要最常见的元素(即模式),一个不错的方法是使用pd.Series.mode。

    In [32]: df
    Out[32]:
         Item  Price  Minimum
    0  Coffee      1        1
    1  Coffee      2        1
    2  Coffee      2        1
    3     Tea      3        3
    4     Tea      4        3
    5     Tea      4        3
    
    In [33]: df['Most_Common_Price'] = df.groupby(["Item"])['Price'].transform(pd.Series.mode)
    
    In [34]: df
    Out[34]:
         Item  Price  Minimum  Most_Common_Price
    0  Coffee      1        1                  2
    1  Coffee      2        1                  2
    2  Coffee      2        1                  2
    3     Tea      3        3                  4
    4     Tea      4        3                  4
    5     Tea      4        3                  4
    

    正如@Wen 所说,pd.Series.mode 可以返回 pd.Series 的值,所以只需抓住第一个:

    Out[67]:
         Item  Price  Minimum
    0  Coffee      1        1
    1  Coffee      2        1
    2  Coffee      2        1
    3     Tea      3        3
    4     Tea      4        3
    5     Tea      4        3
    6     Tea      3        3
    
    In [68]: df[df.Item =='Tea'].Price.mode()
    Out[68]:
    0    3
    1    4
    dtype: int64
    
    In [69]: df['Most_Common_Price'] = df.groupby(["Item"])['Price'].transform(lambda S: S.mode()[0])
    
    In [70]: df
    Out[70]:
         Item  Price  Minimum  Most_Common_Price
    0  Coffee      1        1                  2
    1  Coffee      2        1                  2
    2  Coffee      2        1                  2
    3     Tea      3        3                  3
    4     Tea      4        3                  3
    5     Tea      4        3                  3
    6     Tea      3        3                  3
    

    【讨论】:

    • 一个小改动df.groupby(["Item"])['Price'].transform(lambda x : x.mode()[0]),以防有两个相同的:-)
    • 是否有可能 Pandas 改变了您对第一个解决方案的评估? df.groupby(["Item"])['Price'].transform(pd.Series.mode) 在我的机器上返回 ValueError: Length of passed values is 1, index implies 3。
    • @00schneider:那是因为您的情况下的模式应用于两个或多个值;使用 BENY 的建议
    猜你喜欢
    • 2019-01-18
    • 1970-01-01
    • 2019-02-25
    • 2020-12-06
    • 2021-12-01
    • 2021-12-02
    • 1970-01-01
    • 2015-11-10
    相关资源
    最近更新 更多