【发布时间】:2015-09-29 15:29:19
【问题描述】:
我是 Pandas 的新手,我想知道在以下示例中我做错了什么。
我找到了一个示例 here,它解释了如何在应用 group by 而不是 series 后获取数据框。
df1 = pd.DataFrame( {
"Name" : ["Alice", "Bob", "Mallory", "Mallory", "Bob" , "Mallory"] ,
"City" : ["Seattle", "Seattle", "Baires", "Caracas", "Baires", "Caracas"] })
df1['size'] = df1.groupby(['City']).transform(np.size)
df1.dtypes #Why is size an object? shouldn't it be an integer?
df1[['size']] = df1[['size']].astype(int) #convert to integer
df1['avera'] = df1.groupby(['City'])['size'].transform(np.mean) #group by again
基本上,我想将相同的转换应用于我现在正在处理的庞大数据集,但我收到一条错误消息:
budgetbid['meanpb']=budgetbid.groupby(['jobid'])['probudget'].transform(np.mean) #can't upload this data for the sake of explanation
ValueError: Length mismatch: Expected axis has 5564 elements, new values have 78421 elements
因此,我的问题是:
- 如何克服这个错误?
- 为什么在应用 group by with size 而不是整数类型时得到对象类型?
-
假设我想从
df1获取一个数据框,其中包含独特的城市及其各自的count(*)。我知道我可以做类似的事情newdf=df1.groupby(['City']).size()
不幸的是,这是一个系列,但我想要一个包含两列的数据框,City 和全新的变量,比如说countcity。如何从本示例中的分组操作中获取数据框?
- 你能给我一个在熊猫中
select distinct等价的例子吗?
【问题讨论】: