【问题标题】:Calculate average price and number of unique buyers for each seller and visualize计算每个卖家的平均价格和唯一买家数量并可视化
【发布时间】:2019-12-11 06:08:30
【问题描述】:

我有以下数据框:

      seller_id| buyer_id| quantity| price | transaction_date

          1432 | 344     |   40    | 3420  | 2015-02-01
          1432 | 356     |   41    | 3420  | 2015-02-01
          1432 | 354     |   41    | 3420  | 2015-02-03
          1456 | 354     |   41    | 3420  | 2015-02-04  
          1498 | 354     |   41    | 3420  | 2015-02-04  

对于每个卖家,我想找到唯一买家的数量以及所有唯一卖家销售的平均购买价格。然后我想可视化这些数据,因为我有 200 个独特的卖家,以查看系统中所有卖家的所有销售额的分布。

因此,对于 Seller_id 1432,平均销售额为 3420 美元

我试过df.groupby(['seller_id', 'buyer_id'])['amount'].transform(avg)

它没有返回我想要的结果。

【问题讨论】:

  • df.groupby(['seller_id'])['amount'].mean()?
  • 您还尝试过什么?你读过 Pandas 文档吗?

标签: python pandas pandas-groupby


【解决方案1】:

创建你的熊猫DataFrame

import pandas as pd
from io import StringIO

data = \
"""seller_id|buyer_id|quantity|price|transaction_date
1432 | 344     |   40    | 3420  | 2015-02-01
1432 | 356     |   41    | 3420  | 2015-02-01
1432 | 354     |   41    | 3420  | 2015-02-03
1456 | 354     |   41    | 3420  | 2015-02-04  
1498 | 354     |   41    | 3420  | 2015-02-04  
"""
df = pd.read_csv(StringIO(data), sep='|')

然后根据'seller_id' 创建您的统计信息

sellers_stats = df.groupby(['seller_id']).agg({'price': 'mean', 'buyer_id': pd.Series.nunique}).reset_index()
sellers_stats.columns = ['seller_id', 'avg_price', 'unique_buyers']
sellers_stats

然后你就可以绘图了

import matplotlib.pyplot as plt

fig, ax = plt.subplots()

x = [str(x) for x in sellers_stats['seller_id']]
y = [x for x in sellers_stats['unique_buyers']]
ax.plot(x, y)
ax.set(title='Unique buyers per seller')

plt.show()

和

fig, ax = plt.subplots()

x = [str(x) for x in sellers_stats['seller_id']]
y = [x for x in sellers_stats['avg_price']]
ax.plot(x, y)
ax.set(title='Average amount per seller')

plt.show()

或者你可以做一个单独的情节

import matplotlib.pyplot as plt

width = .35 # width of a bar
sellers_stats['avg_price'].plot(color='blue', kind='bar', width=width)
sellers_stats['unique_buyers'].plot(color='red', secondary_y=True)

ax = plt.gca()
plt.xlim([-width, len(sellers_stats['seller_id'])-width])
ax.set_xticklabels([str(x) for x in sellers_stats['seller_id']])

plt.show()

【讨论】:

    【解决方案2】:
    1. df.groupby(['seller_id','buyer_id']).count()

    2. df.groupby(['seller_id','buyer_id']).mean()['amount']

    哦,我以为这就是你想要的,这应该为你工作 data.groupby(['seller_id']).buyer_id.nunique()

    【讨论】:

    • 对于 1. 这以嵌套的方式为我提供了给定卖家的所有买家 ID,并且没有给我实际的计数。因此,对于每个卖家 ID,我列出了所有唯一买家,但没有计数。
    【解决方案3】:

    这是返回两个请求的数据帧的最小完整示例:

    import pandas as pd
    from io import StringIO
    
    data = \
    """seller_id|buyer_id|quantity|amount|transaction_date
    1432 | 344     |   40    | 3420  | 2015-02-01
    1432 | 356     |   41    | 3420  | 2015-02-01
    1432 | 354     |   41    | 3420  | 2015-02-03
    1456 | 354     |   41    | 3420  | 2015-02-04  
    1498 | 354     |   41    | 3420  | 2015-02-04  
    """
    df = pd.read_csv(StringIO(data), sep='|')
    
    nbs = df.groupby(['seller_id']).buyer_id.nunique()
    print("Number of buyer per seller")
    print(nbs)
    
    avgpp = df.groupby(['seller_id']).mean()['amount']
    print("Average purchase price across all the unique seller's sales")
    print(avgpp)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2019-09-06
      • 1970-01-01
      • 2023-03-21
      • 2020-06-15
      • 1970-01-01
      • 2016-03-02
      • 1970-01-01
      • 2013-09-13
      相关资源
      最近更新 更多