【问题标题】:Grouping by a set in pandas按熊猫中的一组分组
【发布时间】:2021-12-23 06:18:12
【问题描述】:

我有一个例子df:

import pandas as pd 
import numpy as np


df = pd.DataFrame({'name':['Josh', 'Paul','Ivy','Mark'],
                   'orderId':[1,2,3,4],
                   'purchases':[['sofa','sofa','chair'],
                                ['chair','sofa'],
                                ['sofa','chair'],
                                ['sofa','chair','chair']]})

4 人购买了相同的商品 - sofa & chair,但数量不同,但总的来说,他们都购买了 sofa 和 chair - 只有 1 个不同产品的组合,将其视为 set(purchases) .

我想回答每种购买组合被购买了多少次 - 我们知道这是4,因为4 人们购买了同一组商品。

所以我认为这是一个伪代码,我按每个 purchase 值的集合进行分组:

df = df.groupby(set('purchases')).agg({'orderId':pd.Series.nunique})

但我得到一个预期的错误:

TypeError: 'set' 对象不可调用

我想知道通过值的set 而非实际值(在本例中为列表)实现分组的最佳方法是什么。

当我尝试简单地按purchases分组时

df = df.groupby('purchases').agg({'orderId':pd.Series.nunique})

我明白了:

TypeError: unhashable type: 'list'\

我尝试将列表更改为元组:

df = pd.DataFrame({'name':['Josh', 'Paul','Ivy','Mark'],
                   'orderId':[1,2,3,4],
                   'purchases':[('sofa','sofa','chair'),
                                ('chair','sofa'),
                                ('sofa','chair'),
                                ('sofa','chair','chair')]})

然后

df = df.groupby('purchases').agg({'purchases':lambda x:{y for y in x}}) # or set(x)

但这给了

                        purchases
purchases   
(chair, sofa)           {(chair, sofa)}
(sofa, chair)           {(sofa, chair)}
(sofa, chair, chair)    {(sofa, chair, chair)}
(sofa, sofa, chair)     {(sofa, sofa, chair)}

集合内还有一个元组,因为它在寻找相同的元组而不是在元组内部?

我试过了:

df['purchases_unique'] = df['purchases'].apply(lambda x: set(x))
df['# of times bought'] = df.apply(lambda x: x.value_counts())

但我明白了:

TypeError: unhashable type: 'set' 虽然 Jupyter 笔记本仍然通过日志消息提供答案:

Exception ignored in: 'pandas._libs.index.IndexEngine._call_map_locations'
Traceback (most recent call last):
  File "pandas\_libs\hashtable_class_helper.pxi", line 1709, in pandas._libs.hashtable.PyObjectHashTable.map_locations
TypeError: unhashable type: 'set'
{sofa, chair}    4

所以总而言之,我都在寻找答案,如何将值 4 分配给每一行,以便结果如下所示:

name        orderId         purchases                   # of times bought
Josh        1               (sofa, sofa, chair)         4
Paul        2               (chair, sofa)               4
Ivy         3               (sofa, chair)               4
Mark        4               (sofa, chair, chair)        4

如果 python 能够评估 {'chair', 'sofa'} == {'sofa', 'chair'},为什么 pandas 不允许按 set 分组?

【问题讨论】:

  • 注意。你知道set('purchases') 给{'a', 'c', 'e', 'h', 'p', 'r', 's', 'u'} 吗? ;)
  • 是的,我提到这是伪的。

标签: python pandas group-by set


【解决方案1】:

用途:

df["times_bought"] = df.groupby(df["purchases"].apply(frozenset))["purchases"].transform("count")
print(df)

输出

   name  orderId             purchases  times_bought
0  Josh        1   [sofa, sofa, chair]             4
1  Paul        2         [chair, sofa]             4
2   Ivy        3         [sofa, chair]             4
3  Mark        4  [sofa, chair, chair]             4

表达式:

df["purchases"].apply(frozenset)

将购买中的每个列表转换为frozenset:

0    (chair, sofa)
1    (chair, sofa)
2    (chair, sofa)
3    (chair, sofa)

来自文档(重点是我的):

frozenset 类型是不可变和可散列的——它的内容不能 创建后更改;因此它可以用作字典 键或作为另一个集合的元素。

鉴于.apply 之后的元素是不可变和可散列的,它们可以在DataFrame.groupby 中使用。

替代品

最终,您需要将购买中的每个元素映射到相同的标识符,考虑到您的问题的限制。所以你可以直接使用frozenset的哈希函数如下:

def _hash(lst):
    import sys
    uniques = set(lst)
    # https://stackoverflow.com/questions/20832279/python-frozenset-hashing-algorithm-implementation
    MAX = sys.maxsize
    MASK = 2 * MAX + 1
    n = len(uniques)
    h = 1927868237 * (n + 1)
    h &= MASK
    for x in uniques:
        hx = hash(x)
        h ^= (hx ^ (hx << 16) ^ 89869747)  * 3644798167
        h &= MASK
    h = h * 69069 + 907133923
    h &= MASK
    if h > MAX:
        h -= MASK + 1
    if h == -1:
        h = 590923713
    return h


df["times_bought"] = df.groupby(df["purchases"].apply(_hash))["purchases"].transform("count")
print(df)

输出

   name  orderId             purchases  times_bought
0  Josh        1   [sofa, sofa, chair]             4
1  Paul        2         [chair, sofa]             4
2   Ivy        3         [sofa, chair]             4
3  Mark        4  [sofa, chair, chair]             4

第二种选择是使用(作为 groupby 的参数):

df["purchases"].apply(lambda x: tuple(sorted(set(x))))

这将找到唯一的元素,对它们进行排序,最后将它们转换为可散列表示(元组)。

【讨论】:

  • 您能解释一下为什么 frozonset 有效,而 set 无效?
  • @JonasPalačionis frozensets 与 sets 一样,但不可变且可散列,这是 DataFrames/Series 索引的要求(在这种情况下,用作组键)。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2017-12-30
  • 1970-01-01
  • 2021-11-11
  • 2020-03-24
  • 2016-07-07
  • 2020-10-27
  • 2019-12-12
相关资源
最近更新 更多