【发布时间】:2021-04-04 08:27:41
【问题描述】:
假设我有一个数据框,其中包含代表学生专业和学生 GPA 的多索引列。我想找到每一行的每个专业的班级排名。如果学生每个人只有一个专业,那么我可以将专业堆叠为 multiindex 和 groupby(level=1).rank() 的另一个级别。但是,因为学生可以改变他们的专业,所以我必须将这些作为每年的单独功能。什么是获得正确输出的有效方法?我知道我可以通过遍历行来强制它,但我希望避免这种情况,因为我可能正在处理大量数据。
样本df:
categories = ['major', 'gpa']
students = ['Joe', 'Bob', 'Sara']
columns = pd.MultiIndex.from_product([categories, students])
index = range(4)
students_df = pd.DataFrame([['english', 'math', 'math', 3.8, 2.2, 3.7],
['english', 'math', 'math', 3.5, 2.4, 3.9],
['english', 'english', 'math', 3.5, 3.6, 3.9],
['english', 'english', 'math', 3.7, 3.5, 3.8],
], index=range(4), columns=columns)
print(students_df)
major gpa
Joe Bob Sara Joe Bob Sara
0 english math math 3.8 2.2 3.7
1 english math math 3.5 2.4 3.9
2 english english math 3.5 3.6 3.9
3 english english math 3.7 3.5 3.8
预期输出
rank
Joe Bob Sara
0 1 2 1
1 1 2 1
2 2 1 1
3 1 2 1
【问题讨论】:
标签: python pandas pandas-groupby multi-index