可以通过循环遍历.groupby 对Series 或DataFrame 的操作产生的GroupBy 对象来解决此一般任务类别。
在这种特殊情况下,您还可以使用GroupBy.apply method,它对每个组执行计算并将结果连接在一起。
GroupBy 类的文档是 here。
我将首先介绍循环版本,因为对于尚未熟悉计算的“DataFrame 样式”的程序员来说,它可能更易于使用。但是,我建议尽可能使用.apply 版本。处理大型数据集时速度会更快,并且可能会消耗更少的内存。它也被认为是更“惯用”的风格,它将迫使您学习如何将代码分解为单独的函数。
使用循环
很多人没有意识到DataFrame.groupby(GroupBy 对象)的结果可以被迭代。此特定功能已记录在 here。
除此之外,逻辑还包括一个简单的 if 语句、一些 Pandas 子集和 concat function。
完整示例:
import io
import pandas as pd
data = pd.read_csv(io.StringIO('''
Part,Project,Quote,Price,isSelected
1,A,1,5.0,No
1,A,1,2.2,Yes
5,C,2,6.6,No
5,C,2,1.2,Yes
3,B,3,5.5,No
3,B,3,4.6,No
'''))
group_results = []
for _, group in data.groupby(['Part', 'Project', 'Quote']):
is_selected = group['isSelected'] == 'Yes'
if is_selected.any():
# Select the rows where 'isSelected' is True, and
# then select the first row from that output.
# Using [0] instead of 0 ensures that the result
# is still a DataFrame, and that it does not get
# "squeezed" down to a Series.
group_result = group.loc[is_selected].iloc[[0]]
else:
group_result = group
group_results.append(group_result)
results = pd.concat(group_results)
print(results)
输出:
Part Project Quote Price isSelected
1 1 A 1 2.2 Yes
4 3 B 3 5.5 No
5 3 B 3 4.6 No
3 5 C 2 1.2 Yes
使用.apply
GroupBy.apply 方法本质上为您完成了pd.concat 和列表附加部分。我们没有编写循环,而是编写了一个函数,我们将其传递给.apply:
import io
import pandas as pd
data = pd.read_csv(io.StringIO('''
Part,Project,Quote,Price,isSelected
1,A,1,5.0,No
1,A,1,2.2,Yes
5,C,2,6.6,No
5,C,2,1.2,Yes
3,B,3,5.5,No
3,B,3,4.6,No
'''))
groups = data.groupby(['Part', 'Project', 'Quote'], as_index=False)
def process_group(group):
is_selected = group['isSelected'] == 'Yes'
if is_selected.any():
# Select the rows where 'isSelected' is True, and
# then select the first row from that output.
# Using [0] instead of 0 ensures that the result
# is still a DataFrame, and that it does not get
# "squeezed" down to a Series.
group_result = group.loc[is_selected].iloc[[0]]
else:
group_result = group
return group_result
# Use .reset_index to remove the extra index layer created by Pandas,
# which is not necessary in this situation.
results = groups.apply(process_group).reset_index(level=0, drop=True)
print(results)
输出:
Part Project Quote Price isSelected
1 1 A 1 2.2 Yes
4 3 B 3 5.5 No
5 3 B 3 4.6 No
3 5 C 2 1.2 Yes