【发布时间】:2021-05-16 22:32:18
【问题描述】:
我有一个包含 customer_id、年份、订单和其他一些但不重要的列的 df。每次收到新订单时,我的代码都会创建一个新行,因此每个 customer_id 可以有不止一行。我想创建一个“实际”的新列,如果 customer_id 在 2020 年或 2021 年购买,则其中包括“真”。我的代码是:
#Run through all customers and check if they bought in 2020 or 2021
investors = df["customer_id"].unique()
df["actually"] = np.nan
for i in investors:
selected_df = df.loc[df["customer_id"] == i]
for year in selected_df['year'].unique():
if "2021" in str(year) or "2020" in str(year):
df.loc[df["customer_id"] == i, "actually"] = "True"
break
#Want just latest orders / customers
df = df.loc[df["actually"] == "True"]
这很好用,但速度很慢。我想使用 Pandas groupby 功能,但到目前为止还没有找到工作方法。我也避免循环。有人有想法吗?
【问题讨论】:
-
请与预期输出共享示例数据框
-
分享你的输入输出数据帧请阅读minimal reproducible example
标签: python pandas dataframe group-by