【问题标题】:replace pandas column values with counts用计数替换熊猫列值
【发布时间】:2020-06-23 02:06:44
【问题描述】:

Pandas GroupBy 并用标准化计数替换值

样本 DF:

df = pd.DataFrame(np.random.randint(0,20,size=(10,3)),columns=["c1","c2","c3"])
df["r1"]=["Apple","Mango","Apple","Mango","Mango","Mango","Apple","Mango","Apple","Apple"]
df["r2"]=["Orange","lemon","lemon","Orange","lemon","Orange","lemon","lemon","Orange","lemon"]
df["date"] = ["2002-01-01","2002-01-01","2002-01-01","2002-01-01","2002-01-01",
              "2002-01-01","2002-02-01","2002-02-01","2002-02-01","2002-02-01"]
df["date"] = pd.to_datetime(df["date"])
df

DF:

    c1      c2      c3     r1        r2       date
0   10       2       0     Apple    Orange  2002-01-01
1   10      10      13     Mango    lemon   2002-01-01
2   0       12       0     Apple    lemon   2002-01-01
3   1       13       8     Mango    Orange  2002-01-01
4   6        5       9     Mango    lemon   2002-01-01
5   3       18      13     Mango    Orange  2002-01-01
6   2        6       7     Apple    lemon   2002-02-01
7   0        4       7     Mango    lemon   2002-02-01
8   1       10      19     Apple    Orange  2002-02-01
9   11      18       2     Apple    lemon   2002-02-01

我正在尝试按date 列分组,并用标准化计数替换选定的列。

例如:

在2002-01-01 组中,r1 列中的值Apple 将替换为0.3,因为在该组中有6 记录和2records 有Apple,所以2/6 和@ 987654332@ 将替换为 4/6 即 0.6

熊猫解决方案:

df.groupby("date")[["r1","r2"]].apply(lambda x: x.map(x.value_counts()))

错误:

AttributeError: 'DataFrame' object has no attribute 'map'

有没有一种 pandas 方法可以代替迭代的 iterrows 解决方案。

【问题讨论】:

  • 你能提供你想要的输出吗?

标签: python pandas pandas-groupby


【解决方案1】:

我们可以value_counts + normalize

df['New']=df.groupby(['date']).r1.value_counts(normalize=True).reindex(pd.MultiIndex.from_frame(df[['date','r1']])).values
df
   c1  c2  c3     r1      r2       date       New
0   1   8   2  Apple  Orange 2002-01-01  0.333333
1   8   1   7  Mango   lemon 2002-01-01  0.666667
2   0  14   8  Apple   lemon 2002-01-01  0.333333
3  11  13  10  Mango  Orange 2002-01-01  0.666667
4  15   4  15  Mango   lemon 2002-01-01  0.666667
5  13   7   7  Mango  Orange 2002-01-01  0.666667
6   7   0  14  Apple   lemon 2002-02-01  0.750000
7  13   5  11  Mango   lemon 2002-02-01  0.250000
8  19  17  11  Apple  Orange 2002-02-01  0.750000
9   8   1   9  Apple   lemon 2002-02-01  0.750000

【讨论】:

  • 谢谢这很好用。所以我不能在 1 go 中为 r1 and r2 两列都做是吗?
  • @vikky 对于两列您需要找到组合百分比?
  • 是的,精确解同时应用于两列。我可以这样做吗?
  • @vikky 你可以 df['New']=df[['r1','r2']].apply(tuple ,1);df.groupby(['date'])。 New.value_counts(normalize=True).reindex(pd.MultiIndex.from_frame(df[['date','New']])).values
【解决方案2】:

您可以使用transform 方法获取每个组的大小,并将此值分配给原始数据帧的每一行。

In [11]: df.groupby(['date', 'r1'])['c1'].transform(len)/df.groupby(['date'])['c1'].transform(len)                                                    
Out[11]: 
0    0.333333
1    0.666667
2    0.333333
3    0.666667
4    0.666667
5    0.666667
6    0.750000
7    0.250000
8    0.750000
9    0.750000
Name: c1, dtype: float64

如果你需要得到四舍五入的值,只需使用round 方法。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2022-06-24
    • 2022-12-23
    • 2023-03-06
    • 2022-01-11
    • 2018-11-04
    • 2023-02-25
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多