【发布时间】:2019-03-15 19:25:48
【问题描述】:
我有以下数据框
import pandas as pd
newd = {'year': [2001, 2002, 2005, 2002, 2004, 2001, 2001, 2002, 2003, 2003, 2002, 2002, 2003, 2004, 2005, 2003, 2004, 2005, 2004, 2004 ],
'indviduals': [12, 23, 24, 28,30, 15, 17, 18, 18, 19, 12, 15, 12, 12, 12, 15, 15, 15, 12, 12],
'employers': ['a', 'b', 'c', 'd', 'e', 'a', 'a', 'b', 'b', 'c', 'b', 'a', 'c', 'd', 'e', 'a', 'a', 'a', 'a', 'b'] }
newdf=newdf=pd.DataFrame(newd)
我的预期结果(只是一个例子):
2001, a: [12, 15, 17] count:3 employerchanged: []
2002, b: [12, 23, 28] count:3 employerchanged: [12]
2002, a: [15] count:1
这在 SQL 中完成时很容易。但是如果个人“12”在 2001 年到 2002 年之间更换雇主,SQL 不会告诉我方法。
这是我迄今为止在 python 中尝试过的:
dic={}
listofUniqueYears= [i for i in newdf.year.unique()]
# 给了我独特年份的列表
dic={}
for i in listofUniqueYears:
dic[i]=defaultdict(dict)
print(dic)
我的问题是如何根据我提供的条件过滤行值,在这种情况下,我希望每个雇主每年都有员工数量、计数和更改的员工。
【问题讨论】:
-
每个人每年只有一个条目(行)吗?
-
是的,但那一年他们可能有两个雇主
-
示例打印(newdf[newdf['year']==2004])。个人 12 在 2004 年有两个雇主
标签: python-3.x pandas list dictionary