【问题标题】:Replace the unique values in a DataFrame column with their count [duplicate]用它们的计数替换 DataFrame 列中的唯一值[重复]
【发布时间】:2018-03-05 03:53:17
【问题描述】:

我有一个这样的DataFrame:

Index Label
0     ABCD
1     EFGH
2     ABCD
3     ABCD
4     EFGH
5     ABCD
6     IJKL
7     IJKL
8     ABCD
9     EFGH

因此,“ABCD”出现 5 次,“EFGH”出现 3 次,“IJKL”出现两次。我想计算每个标签的出现次数并用它们的计数替换单个标签,以获得以下信息:

Index Label
0     5
1     3
2     5
3     5
4     3
5     5
6     2
7     2
8     5
9     3

最好的方法是什么? 谢谢!

【问题讨论】:

    标签: python pandas dataframe unique


    【解决方案1】:

    使用由Series创建的mapvalue_counts:

    df['Label'] = df['Label'].map(df['Label'].value_counts())
    print (df)
       Label
    0      5
    1      3
    2      5
    3      5
    4      3
    5      5
    6      2
    7      2
    8      5
    9      3
    

    transform + size 的另一种解决方案:

    df['Label'] = df.groupby('Label')['Label'].transform('size')
    print (df)
    
       Label
    0      5
    1      3
    2      5
    3      5
    4      3
    5      5
    6      2
    7      2
    8      5
    9      3
    

    【讨论】:

    • 你确定吗?我认为总是需要size,如果需要排除NaNs 需要count(很少使用)
    【解决方案2】:

    使用groupby 和transform:

    print(df)
          Label
    Index      
    0      ABCD
    1      EFGH
    2      ABCD
    3      ABCD
    4      EFGH
    5      ABCD
    6      IJKL
    7      IJKL
    8      ABCD
    9      EFGH
    
    df['Label'] = df.groupby('Label').Label.transform('count')
    print(df)
           Label
    Index       
    0          5
    1          3
    2          5
    3          5
    4          3
    5          5
    6          2
    7          2
    8          5
    9          3
    

    如果您的列没有 NaNs,size 和 count 返回相同的值。否则,size 包含NaNs,所以避免使用它。


    使用Counter的另一种方式:

    from collections import Counter
    
    df['Label'] = df.Label.map(Counter(df.Label))
    print(df)
           Label
    Index       
    0          5
    1          3
    2          5
    3          5
    4          3
    5          5
    6          2
    7          2
    8          5
    9          3
    

    【讨论】:

    • @P.Prunesquallor 感谢您的支持。
    • @P.Prunesquallor 另外,如果您使用的是 groupby 解决方案,请不要像 jezrael 的解决方案那样使用 size。
    • 我不明白Otherwise, size includes NaNs, so avoid using it.为什么要避免?我认为这两个函数都很好 - 我认为函数 count 是最好的不使用,只有在需要明确排除 NaN 时才使用。我认为没有理由避免使用size,因为如果我知道我有一些 NaN(而且我认为数据中没有 NaN,尤其是浮点数据),那就太好了。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2020-12-17
    • 2017-05-30
    • 2017-11-08
    • 2019-07-16
    • 1970-01-01
    • 2020-02-06
    • 2021-12-02
    相关资源
    最近更新 更多