【问题标题】:Create new column by joining itself-table multiple times通过多次加入自身表来创建新列
【发布时间】:2020-02-12 05:14:01
【问题描述】:

我有一个包含大家庭成员列表的 pandas 数据框。

import pandas as pd

data = {'child':['Joe','Anna','Anna','Steffani','Bob','Rea','Dani','Dani','Selma','John','Kevin'],
             'parents':['Steffani','Bob','Steffani','Dani','Selma','Anna','Selma','John','Kevin','-','Robert'],
            }
df = pd.DataFrame(data)

从这个数据框中,我需要通过在右侧添加多个列来显示数据之间的关系来构建一个新表。右栏中的值显示了长辈关系。每列代表关系。如果我能画出图表,它可能是这样的:

child --> parents --> grandparents --> parents of grandparents --> grandparents of grandparents --> etc.

因此,数据帧的预期输出将是这样的:

    child       parents     A           B           C           D (etc)
---------------------------------------------------------------------------------
0   Joe         Steffani    Dani        Selma       Kevin       <If still possible>
1   Joe         Steffani    Dani        John        -
2   Anna        Bob         Selma       Kevin       Robert
3   Anna        Steffani    Dani        Selma       Kevin
4   Anna        Steffani    Dani        John        -
5   Steffani    Dani        Selma       Kevin       Robert
6   Steffani    Dani        John        -           -
7   Bob         Selma       Kevin       Robert      -
8   Rea         Anna        Bob         Selma       Kevin
9   Rea         Anna        Steffani    Dani        Selma
10  Rea         Anna        Steffani    Dani        John
11  Dani        Selma       Kevin       Robert      -
12  Dani        John        -           -           -
13  Selma       Kevin       Robert      -           -
14  John        -           -           -           -
15  Kevin       Robert      -           -           -

目前,我使用pandas.merge 手动构建新表。但是我需要做很多次,直到最后一列与左列没有长辈关系。 例如:

第 1 步

df2 = pd.merge(df, df, left_on='parents', right_on='child', how='left').fillna('-')
df2 = df2[['child_x','parents_x','parents_y']]
df2.columns = ['child','parents','A']

第 2 步

df3 = pd.merge(df2, df, left_on='A', right_on='child', how='left').fillna('-')
df3 = df3[['child_x','parents_x','A','parents_y']]
df3.columns = ['child','parents','A','B']

第 3 步

df4 = pd.merge(df3, df, left_on='B', right_on='child', how='left').fillna('-')
df4 = df4[['child_x','parents_x','A','B','parents_y']]
df4.columns = [['child','parents','A','B','C']]

第 4 步

如果C列中的值仍然具有长辈关系,则编写类似的代码为D列添加第6列。

问题:

由于我的dataframe中有大数据(超过10K的数据点),不一步步写代码如何解决?我不知道构建决赛桌需要多少步骤。

提前感谢您的帮助。

【问题讨论】:

    标签: python dataframe join merge


    【解决方案1】:

    考虑使用merge 的suffixes 参数与reduce 进行链合并,并对重复列名进行一些处理并删除中间子 列:

    def proc_build(x,y):
        temp = (pd.merge(x, y, left_on='parents', right_on='child', 
                         how='left', suffixes=['_',''])                     
                  .fillna('-'))
    
        return temp       
    
    final_df = (reduce(proc_build, [df, df, df, df])
                   .set_axis(['child', 'parents',
                              'child1', 'A', 
                              'child2', 'B',
                              'child3', 'C'], axis='columns', inplace=False)
                   .reindex(['child', 'parents'] + list('ABC'), axis='columns')
               )
    
    print(final_df)
    
    #        child   parents         A       B       C
    # 0        Joe  Steffani      Dani   Selma   Kevin
    # 1        Joe  Steffani      Dani    John       -
    # 2       Anna       Bob     Selma   Kevin  Robert
    # 3       Anna  Steffani      Dani   Selma   Kevin
    # 4       Anna  Steffani      Dani    John       -
    # 5   Steffani      Dani     Selma   Kevin  Robert
    # 6   Steffani      Dani      John       -       -
    # 7        Bob     Selma     Kevin  Robert       -
    # 8        Rea      Anna       Bob   Selma   Kevin
    # 9        Rea      Anna  Steffani    Dani   Selma
    # 10       Rea      Anna  Steffani    Dani    John
    # 11      Dani     Selma     Kevin  Robert       -
    # 12      Dani      John         -       -       -
    # 13     Selma     Kevin    Robert       -       -
    # 14      John         -         -       -       -
    # 15     Kevin    Robert         -       -       -
    

    要扩展另一列,例如 D,请将另一个 df 添加到 reduce 的 iterable 参数以及 set_axis 和 reindex 中的其他列表项,特别是['child4', 'D'] 和list('ABCD')。虽然有一些方法可以使这些项目动态化,但reduce 可能会变得很昂贵,因此应该在处理时强调一些声明性。

    final_df = (reduce(proc_build, [df] * 5)
                   .set_axis(['child', 'parents',
                              'child1', 'A', 
                              'child2', 'B',
                              'child3', 'C',
                              'child4', 'D'], axis='columns', inplace=False)
                   .reindex(['child', 'parents'] + list('ABCD'), axis='columns')
               )
    
    print(final_df)
    
    #        child   parents         A       B       C       D
    # 0        Joe  Steffani      Dani   Selma   Kevin  Robert
    # 1        Joe  Steffani      Dani    John       -       -
    # 2       Anna       Bob     Selma   Kevin  Robert       -
    # 3       Anna  Steffani      Dani   Selma   Kevin  Robert
    # 4       Anna  Steffani      Dani    John       -       -
    # 5   Steffani      Dani     Selma   Kevin  Robert       -
    # 6   Steffani      Dani      John       -       -       -
    # 7        Bob     Selma     Kevin  Robert       -       -
    # 8        Rea      Anna       Bob   Selma   Kevin  Robert
    # 9        Rea      Anna  Steffani    Dani   Selma   Kevin
    # 10       Rea      Anna  Steffani    Dani    John       -
    # 11      Dani     Selma     Kevin  Robert       -       -
    # 12      Dani      John         -       -       -       -
    # 13     Selma     Kevin    Robert       -       -       -
    # 14      John         -         -       -       -       -
    # 15     Kevin    Robert         -       -       -       -
    

    【讨论】:

    • 如果我不知道在这个过程中是否需要列 E、F 等怎么办?但至少你的代码是有帮助的。谢谢
    【解决方案2】:

    这是我的一个粗略解决方案。你应该优化它。

    • 正在加载所有数据帧
    • 将所有数据框的名称保存在列表中
    list_data = [data1,data2]
    list_df = []
    i = 0
    for data in list_data:
        vars()[f'df{i}'] = pd.DataFrame(data)
        list_df.append(f'df{i}')
        i += 1
    
    • 然后创建2个代理变量;
      • df_family : 这将是一个输出
      • last_df :为了打破循环,如果父列中的每一行都是'-',但列表中还剩下数据框。
    last_df = False
    df_family = pd.DataFrame()
    
    
    • 此部分将根据需要将数据框合并在一起。我还将名称更改为 1,2,...,n,以便您轻松重命名。
    for df in list_df:
        if last_df:
            break
    
        if (eval(df)['parents'] == '-').all():
            last_df = True
    
        if df_family.empty:
            df_family = eval(df)
        else:
            df_family = pd.merge(df_family,eval(df), how = 'left', left_on = df_family.columns[-1], right_on = eval(df).columns[0])
            df_family.drop(columns = [eval(df).columns[0]], axis = 1, inplace = True)
    
        list_cols = [i for i in range(df_family.shape[1])]
        df_family.columns = list_cols
    

    【讨论】:

    • 在 pandas 中,如果你不得不调用 pd.DataFrame(),你就不是在优化处理对象。在循环内使用append 或merge 增长系列/数据框是very expensive operation。而是构建一个熊猫对象的列表/字典,以便在循环外连接一次。
    • @Parfait 谢谢你的建议。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2017-06-23
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多