【问题标题】:Based on Partial string Match fill one data frame column from another dataframe基于部分字符串匹配从另一个数据帧填充一个数据帧列
【发布时间】:2020-08-31 19:13:27
【问题描述】:

我是 python 编程的新手。我有两个数据框 df1 包含标签(180k 行)和 df2 包含设备名称(1600 行)

df1:

          Line                TagName                CLASS 
187877    PT_WOA  .ZS01_LA120_T05.SB.S2384_LesSwL     10
187878    PT_WOA  .ZS01_RB2202_T05.SB.S2385_FLOK      10
187879    PT_WOA  .ZS01_LA120_T05.SB._CBAbsHy         10
187880    PT_WOA  .ZS01_LA120_T05.SB.S3110_CBAPV      10
187881    PT_WOA  .ZS01_LARB2204.SB.S3111_CBRelHy     10

df2:

EquipmentNo EquipmentDescription    Equipment
1311256        Lifting table         LA120
1311257        Roller bed            RB2200
1311258        Lifting table         LT2202
1311259        Roller bed            RB2202
1311260        Roller bed            RB2204

df2.Equipment 位于 df1.TagName 中的字符串之间。我需要根据 df2 设备是否在 df1 标记名中进行匹配,然后 df2(设备描述和设备编号)必须与 df1 匹配。

最终输出应该是

        Line                TagName                quipmentdescription   EquipmentNo 
187877  PT_WOA  .ZS01_LA120_T05.SB.S2384_LesSwL     Lifting table        1311256
187878  PT_WOA  .ZS01_RB2202_T05.SB.S2385_FLOK      Roller bed           1311259  
187879  PT_WOA  .ZS01_LA120_T05.SB._CBAbsHy         Lifting table        1311256 
187880  PT_WOA  .ZS01_LA120_T05.SB.S3110_CBAPV      Lifting table        1311256
 187881 PT_WOA  .ZS01_LARB2204.SB.S3111_CBRelHy     Roller bed           1311260

我已经试过了

cols= df2['Equipment'].tolist()
Xs=[]
for i in cols:
    Test = df1.loc[df1.TagName.str.contains(i)] 
    Test['Equip']=i
    Xs.append(Test)

然后根据“设备”合并xs和df2

但是我收到了这个错误

第一个参数必须是字符串或编译模式

【问题讨论】:

    标签: python string pandas dataframe


    【解决方案1】:
    • 给定df1和df2如下:

    df1

    |    | Line   | TagName                         |   CLASS |
    |---:|:-------|:--------------------------------|--------:|
    |  0 | PT_WOA | .ZS01_LA120_T05.SB.S2384_LesSwL |      10 |
    |  1 | PT_WOA | .ZS01_RB2202_T05.SB.S2385_FLOK  |      10 |
    |  2 | PT_WOA | .ZS01_LA120_T05.SB._CBAbsHy     |      10 |
    |  3 | PT_WOA | .ZS01_LA120_T05.SB.S3110_CBAPV  |      10 |
    |  4 | PT_WOA | .ZS01_LARB2204.SB.S3111_CBRelHy |      10 |
    

    df2

    |    |   EquipmentNo | EquipmentDescription   | Equipment   |
    |---:|--------------:|:-----------------------|:------------|
    |  0 |       1311256 | Lifting table          | LA120       |
    |  1 |       1311257 | Roller bed             | RB2200      |
    |  2 |       1311258 | Lifting table          | LT2202      |
    |  3 |       1311259 | Roller bed             | RB2202      |
    |  4 |       1311260 | Roller bed             | RB2204      |
    
    1. 在df2 中从Equipment 中找到唯一的equipment
    equipment = df2.Equipment.unique().tolist()
    
    1. 通过在equipment 中查找匹配项,在df1 中创建Equipment 列
    df1['Equipment'] = df1['TagName'].apply(lambda x: ''.join([part for part in equipment if part in x]))
    
    1. 合并Equipment 为最终形式
      • 如果您不想在df_final 中添加Equipment 列,请将.drop(columns=['Equipment']) 添加到下一行代码的末尾。
    df_final = df1[['Line', 'TagName', 'Equipment']].merge(df2, on='Equipment')
    

    df_final

    |    | Line   | TagName                         | Equipment   |   EquipmentNo | EquipmentDescription   |
    |---:|:-------|:--------------------------------|:------------|--------------:|:-----------------------|
    |  0 | PT_WOA | .ZS01_LA120_T05.SB.S2384_LesSwL | LA120       |       1311256 | Lifting table          |
    |  1 | PT_WOA | .ZS01_LA120_T05.SB._CBAbsHy     | LA120       |       1311256 | Lifting table          |
    |  2 | PT_WOA | .ZS01_LA120_T05.SB.S3110_CBAPV  | LA120       |       1311256 | Lifting table          |
    |  3 | PT_WOA | .ZS01_RB2202_T05.SB.S2385_FLOK  | RB2202      |       1311259 | Roller bed             |
    |  4 | PT_WOA | .ZS01_LARB2204.SB.S3111_CBRelHy | RB2204      |       1311260 | Roller bed             |
    

    【讨论】:

      【解决方案2】:

      我会这样做:

      1. 创建一个新列 indexes,其中对于 df2 中的每个 Equipment 查找 df1 中的索引列表,其中 df1.TagName 包含 Equipment。

      2. 通过使用stack() 和reset_index() 为每个项目创建一行来展平indexes

      3. 将扁平化 df2 与 df1 结合起来,获取所需的所有信息
      from io import StringIO
      import numpy as np
      import pandas as pd
      df1=StringIO("""Line;TagName;CLASS
      187877;PT_WOA;.ZS01_LA120_T05.SB.S2384_LesSwL;10
      187878;PT_WOA;.ZS01_RB2202_T05.SB.S2385_FLOK;10
      187879;PT_WOA;.ZS01_LA120_T05.SB._CBAbsHy;10
      187880;PT_WOA;.ZS01_LA120_T05.SB.S3110_CBAPV;10
      187881;PT_WOA;.ZS01_LARB2204.SB.S3111_CBRelHy;10""")
      df2=StringIO("""EquipmentNo;EquipmentDescription;Equipment
      1311256;Lifting table;LA120
      1311257;Roller bed;RB2200
      1311258;Lifting table;LT2202
      1311259;Roller bed;RB2202
      1311260;Roller bed;RB2204""")
      df1=pd.read_csv(df1,sep=";")
      df2=pd.read_csv(df2,sep=";")
      
      df2['indexes'] = df2['Equipment'].apply(lambda x: df1.index[df1.TagName.str.contains(str(x)).tolist()].tolist())
      indexes = df2.apply(lambda x: pd.Series(x['indexes']),axis=1).stack().reset_index(level=1, drop=True)
      indexes.name = 'indexes'
      df2 = df2.drop('indexes', axis=1).join(indexes).dropna()
      df2.index = df2['indexes']
      matches = df2.join(df1, how='inner')
      print(matches[['Line','TagName','EquipmentDescription','EquipmentNo']])
      

      输出:

                Line                          TagName EquipmentDescription  EquipmentNo
      187877  PT_WOA  .ZS01_LA120_T05.SB.S2384_LesSwL        Lifting table      1311256
      187879  PT_WOA      .ZS01_LA120_T05.SB._CBAbsHy        Lifting table      1311256
      187880  PT_WOA   .ZS01_LA120_T05.SB.S3110_CBAPV        Lifting table      1311256
      187878  PT_WOA   .ZS01_RB2202_T05.SB.S2385_FLOK           Roller bed      1311259
      187881  PT_WOA  .ZS01_LARB2204.SB.S3111_CBRelHy           Roller bed      1311260
      

      【讨论】:

      • 得到这个错误 "'float' 对象没有属性 'replace'" 1 Test['indexes'] = Test['Equip'].apply(lambda x: Data.index[Data.TagName .str.contains(x.replace(" ","")).tolist()].tolist())
      • @Nandan 看来您的“装备”列是浮点类型,应该是字符串类型
      • 尝试复制和粘贴我的解决方案,只是为了检查它是否适合您。如果是这样,可能您必须仔细检查原始数据集的类型
      • 这很奇怪,我编辑了我的答案...尝试调用 str(x) 而不是 x.replace(" ","")... 但 pandas 正在将此列转换为 float跨度>
      • indexes = df2.apply(lambda x: pd.Series(x['indexes']),axis=1).stack().reset_index(level=1, drop=True) 导致DeprecationWarning:
      【解决方案3】:

      初始化提供的数据帧:

      import numpy as np
      import pandas as pd
      
      df1 = pd.DataFrame([['PT_WOA', '.ZS01_LA120_T05.SB.S2384_LesSwL', 10],
                          ['PT_WOA', '.ZS01_RB2202_T05.SB.S2385_FLOK', 10],
                          ['PT_WOA', '.ZS01_LA120_T05.SB._CBAbsHy', 10],
                          ['PT_WOA', '.ZS01_LA120_T05.SB.S3110_CBAPV', 10],
                          ['PT_WOA', '.ZS01_LARB2204.SB.S3111_CBRelHy', 10]],
                         columns = ['Line', 'TagName', 'CLASS'],
                         index = [187877, 187878, 187879, 187880, 187881])
      
      df2 = pd.DataFrame([[1311256, 'Lifting table', 'LA120'],
                          [1311257, 'Roller bed', 'RB2200'],
                          [1311258, 'Lifting table', 'LT2202'],
                          [1311259, 'Roller bed', 'RB2202'],
                          [1311260, 'Roller bed', 'RB2204']],
                        columns = ['EquipmentNo', 'EquipmentDescription', 'Equipment'])
      

      我建议如下:

      # create a copy of df1, dropping the 'CLASS' column
      df3 = df1.drop(columns=['CLASS'])
      
      # add the columns 'EquipmentDescription' and 'Equipment' filled with numpy NaN's
      df3['EquipmentDescription'] = np.nan
      df3['EquipmentNo'] = np.nan
      
      # for each row in df3, iterate over each row in df2
      for index_df3, row_df3 in df3.iterrows():
          for index_df2, row_df2 in df2.iterrows():
      
              # check if 'Equipment' is in 'TagName'
              if df2.loc[index_df2, 'Equipment'] in df3.loc[index_df3, 'TagName']:
      
                  # set 'EquipmentDescription' and 'EquipmentNo'
                  df3.loc[index_df3, 'EquipmentDescription'] = df2.loc[index_df2, 'EquipmentDescription']
                  df3.loc[index_df3, 'EquipmentNo'] = df2.loc[index_df2, 'EquipmentNo']
      
      
      # conver the 'EquipmentNo' to type int
      df3['EquipmentNo'] = df3['EquipmentNo'].astype(int)
      
      

      这会产生以下数据框:

              Line    TagName                         EquipmentDescription EquipmentNo
      187877  PT_WOA  .ZS01_LA120_T05.SB.S2384_LesSwL Lifting table        1311256
      187878  PT_WOA  .ZS01_RB2202_T05.SB.S2385_FLOK  Roller bed           1311259
      187879  PT_WOA  .ZS01_LA120_T05.SB._CBAbsHy     Lifting table        1311256
      187880  PT_WOA  .ZS01_LA120_T05.SB.S3110_CBAPV  Lifting table        1311256
      187881  PT_WOA  .ZS01_LARB2204.SB.S3111_CBRelHy Roller bed           1311260
      

      如果这有帮助,请告诉我。

      【讨论】:

      • if df3.loc[index_df3, 'TagName'] 中的 df2.loc[index_df2, 'Equipment']:“'in ' 需要字符串作为左操作数,而不是浮点数”我得到了这个错误
      • 我假设您只有在使用此处未显示的观察时才会收到此错误,即因为您对 df2 中的“设备”列有非字符串观察。尝试在创建 df2 之后添加行 df2['Equipment'] = df2['Equipment'].astype(str)。这能解决您的错误吗?
      猜你喜欢
      • 2018-12-08
      • 2020-10-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-01-16
      • 2021-11-21
      相关资源
      最近更新 更多