【问题标题】:How to zip two columns into a key value pair dictionary in pandas如何将两列压缩到熊猫中的键值对字典中
【发布时间】:2023-02-14 13:50:03
【问题描述】:

我有一个包含两个相关列的数据框,需要合并到一个 dictionary 列中。

样本数据:

    skuId   coreAttributes.price    coreAttributes.amount
0   100     price                   8.84
1   102     price                   12.99
2   103     price                   9.99

预期输出:

skuId    coreAttributes
100      {'price': 8.84}
102      {'price': 12.99}
103      {'price': 9.99}

我试过的:

planProducts_T = planProducts.filter(regex = 'coreAttributes').T
planProducts_T.columns = planProducts_T.iloc[0]
planProducts_T.iloc[1:].to_dict(orient = 'records')

我得到 UserWarning: DataFrame columns are not unique, some columns will be omitted. 和这个输出:

[{'price': 9.99}]

你能帮我解决这个问题吗?

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    您可以将列表理解与 python 的 zip 一起使用:

    df['coreAttributes'] = [{k: v} for k,v in
                            zip(df['coreAttributes.price'],
                                df['coreAttributes.amount'])]
    

    输出:

       skuId coreAttributes.price  coreAttributes.amount    coreAttributes
    0    100                price                   8.84   {'price': 8.84}
    1    102                price                  12.99  {'price': 12.99}
    2    103                price                   9.99   {'price': 9.99}
    

    如果需要删除初始列,请使用pop。

    df['coreAttributes'] = [{k: v} for k,v in
                            zip(df.pop('coreAttributes.price'),
                                df.pop('coreAttributes.amount'))]
    

    输出:

       skuId    coreAttributes
    0    100   {'price': 8.84}
    1    102  {'price': 12.99}
    2    103   {'price': 9.99}
    

    【讨论】:

      【解决方案2】:

      您可以使用 apply 和 drop 进行优化计算

      df["coreAttributes"] = df.apply(lambda row: {row["coreAttributes.price"]: row["coreAttributes.amount"]}, axis=1)
      df.drop(["coreAttributes.price","coreAttributes.amount"], axis=1)
      

      输出

          skuId   coreAttributes
      0   100     {'price': 8.84}
      1   102     {'price': 12.99}
      2   103     {'price': 9.99}
      

      【讨论】:

        【解决方案3】:
        df.set_index("skuId").apply(lambda ss:{ss[0]:ss[1]},axis=1).rename("coreAttributes").reset_index()
        

        出去:

         skuId    coreAttributes
        0    100   {'price': 8.84}
        1    102  {'price': 12.99}
        2    103   {'price': 9.99}
        

        【讨论】:

          猜你喜欢
          • 2022-01-08
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2018-08-07
          • 2021-12-02
          • 2023-02-23
          • 1970-01-01
          相关资源
          最近更新 更多