【问题标题】:Pandas create new column with groupby and avoid loopsPandas 使用 groupby 创建新列并避免循环
【发布时间】:2021-05-16 22:32:18
【问题描述】:

我有一个包含 customer_id、年份、订单和其他一些但不重要的列的 df。每次收到新订单时,我的代码都会创建一个新行,因此每个 customer_id 可以有不止一行。我想创建一个“实际”的新列,如果 customer_id 在 2020 年或 2021 年购买,则其中包括“真”。我的代码是:

#Run through all customers and check if they bought in 2020 or 2021
investors = df["customer_id"].unique()
df["actually"] = np.nan
for i in investors:
    selected_df = df.loc[df["customer_id"] == i]
    for year in selected_df['year'].unique():
        if "2021" in str(year) or "2020" in str(year):
            df.loc[df["customer_id"] == i, "actually"] = "True"
            break
#Want just latest orders / customers
df = df.loc[df["actually"] == "True"]

这很好用,但速度很慢。我想使用 Pandas groupby 功能,但到目前为止还没有找到工作方法。我也避免循环。有人有想法吗?

【问题讨论】:

标签: python pandas dataframe group-by


【解决方案1】:

您可以像这样创建列名“Actually”。

list1=df['Customer_id'][df.year==2020].unique()
list2=df['Customer_id'][df.year==2021].unique()
df['Actually']=df['Customer_id'].apply( lambda x : x in list1 or x in list2)

【讨论】:

    【解决方案2】:

    根据我对您的场景的理解,这是一个简单的代码:

    import pandas as pd
    
    # Sample data to recreate the scenarion
    data = {'customer_id': ['c1','c2','c1','c4','c3','c3'], 'year': [2019, 2018,2021,2012,2020,2021], 'order': ['A1','A2','A3','A4','A5','A6']}
    df = pd.DataFrame.from_dict(data)
    
    # Creating the new column (initially all false)
    df['actually'] = False
    
    # Filling only required rows with True
    df.loc[(df['year']==2020) | (df['year']==2021), 'actually'] = True
    
    print(df)
    

    这将产生:

      customer_id  year order  actually
    0          c1  2019    A1     False
    1          c2  2018    A2     False
    2          c1  2021    A3      True
    3          c4  2012    A4     False
    4          c3  2020    A5      True
    5          c3  2021    A6      True
    

    【讨论】:

      【解决方案3】:

      您可以使用 apply 方法,以避免循环:

      df['actually']=df['customer_id'].apply(lambda x: df[df.customer_id==x]['year'].str.contains('2020').any() or df[df.customer_id==x]['year'].str.contains('2021').any())
      

      【讨论】:

      • apply 不只是变相的循环吗?
      • 有点,但我认为比循环更好。我在这里看不到任何更快的方法,例如 np.where 因为条件不是逐行的
      猜你喜欢
      • 2021-10-09
      • 2020-06-10
      • 1970-01-01
      • 2013-12-05
      • 1970-01-01
      • 1970-01-01
      • 2021-11-25
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多