【问题标题】:Pandas: The apply function I am using is giving me wrong results熊猫:我使用的应用功能给了我错误的结果
【发布时间】:2016-05-31 14:43:24
【问题描述】:

我有一个看起来像这样的数据集

     a_id b_received brand_id c_consumed type_received       date  output  \
0    sam       soap     bill        oil       edibles 2011-01-01       1   
1    sam        oil    chris        NaN       utility 2011-01-02       1   
2    sam      brush      dan       soap       grocery 2011-01-03       0   
3  harry        oil      sam      shoes      clothing 2011-01-04       1   
4  harry      shoes     bill        oil       edibles 2011-01-05       1   
5  alice       beer      sam       eggs     breakfast 2011-01-06       0   
6  alice      brush    chris      brush      cleaning 2011-01-07       1   
7  alice       eggs      NaN        NaN       edibles 2011-01-08       1   

我正在使用以下代码

 def probability(x):
    y=[]
    for i in range(len(x)):
        y.append(float(x[i])/float(len(x)))
    return y

 df2['prob']= (df2.groupby('a_id')
           .apply(probability(['output']))
           .reset_index(level='a_id', drop=True))

理想的结果应该是具有以下值的新列

    prob  
 0  0.333334  
 1  0.333334  
 2  0.0  
 3  0.5  
 4  0.5  
 5  0     
 6  0.333334     
 7  0.333334     

但我遇到了一个错误

y.append(float(x[i])/float(len(x)))
ValueError: could not convert string to float: output

列输出为 int 格式。我不明白为什么会出现此错误。

我正在尝试计算每个消费产品的人的输出概率,该产品由列输出给出。例如,如果 sam 收到了肥皂,并且肥皂也出现在“c_consumed”列中,则结果为 1,否则结果为 0。

现在,由于 sam 收到了 3 种产品,其中他消费了 2 种,因此消费每种产品的概率是 1/3。所以输出为 1 的概率应该是 0.333334,输出为 0 的概率应该是 0。

如何达到预期的效果?

【问题讨论】:

  • 我想你想要df2.groupby('a_id').output.apply(probability)。
  • "列输出是 int 格式。我不明白为什么会出现这个错误。"您没有传递“列输出”,而是传递列表['output']。

标签: python pandas group-by apply


【解决方案1】:

我认为您可以简单地将output 列与已经计算的分组.groupby('a_id')['output'] 一起传递给GroupBy 对象,然后使用函数probability,它只返回除列output 及其@987654328 @:

def probability(x):
    #print x
    return x / len(x)

df2['prob']= (df2.groupby('a_id')['output']
           .apply(probability)
           .reset_index(level='a_id', drop=True))

或者lambda:

df2['prob']= (df2.groupby('a_id')['output']
           .apply(lambda x: x / len(x) )
           .reset_index(level='a_id', drop=True))

transform 是更简单、更快速的解决方案:

df2['prob']= df2['output'] / df2.groupby('a_id')['output'].transform('count')
print df2
    a_id b_received brand_id c_consumed type_received        date  output  \
0    sam       soap     bill        oil       edibles  2011-01-01       1   
1    sam        oil    chris        NaN       utility  2011-01-02       1   
2    sam      brush      dan       soap       grocery  2011-01-03       0   
3  harry        oil      sam      shoes      clothing  2011-01-04       1   
4  harry      shoes     bill        oil       edibles  2011-01-05       1   
5  alice       beer      sam       eggs     breakfast  2011-01-06       0   
6  alice      brush    chris      brush      cleaning  2011-01-07       1   
7  alice       eggs      NaN        NaN       edibles  2011-01-08       1   

       prob  
0  0.333333  
1  0.333333  
2  0.000000  
3  0.500000  
4  0.500000  
5  0.000000  
6  0.333333  
7  0.333333  

时间安排:

In [505]: %timeit (df2.groupby('a_id')['output'].apply(lambda x: x / len(x) ).reset_index(level='a_id', drop=True))
The slowest run took 10.99 times longer than the fastest. This could mean that an intermediate result is being cached 
100 loops, best of 3: 1.73 ms per loop

In [506]: %timeit df2['output'] / df2.groupby('a_id')['output'].transform('count')
The slowest run took 5.03 times longer than the fastest. This could mean that an intermediate result is being cached 
1000 loops, best of 3: 449 µs per loop

【讨论】:

  • 感谢您的解决方案!
  • 请不要忘记accept 并支持解决方案。谢谢。
猜你喜欢
  • 1970-01-01
  • 2011-01-07
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2015-09-30
相关资源
最近更新 更多