【问题标题】:How to filter out Column data From Multiple rows data?如何从多行数据中过滤掉列数据?
【发布时间】:2021-01-31 08:43:09
【问题描述】:

晚上好

大家好,我从沃尔玛获得了以下关于其产品和价格的 JSON 文件。

所以我加载了 jupyter notebook,导入了 pandas,然后将其加载到具有自定义列的 Dataframe 中,如下图所示。

现在这就是我想做的:

  1. 创建名为最低价格和最高价格的新列并将数据加载到其中

我该怎么做?

这里是jupyter notebook中的代码供参考。

我也想要报价,因为有些商品没有 minprice 和 maxprice :)

编辑:这是 Python 代码:

import json
import pandas as pd


with open("walmart.json") as f:
    data = json.load(f)

walmart = data["items"]


wdf = pd.DataFrame(walmart,columns=["productId","primaryOffer"])


print(wdf.loc[0,"primaryOffer"])


pd.set_option('display.max_colwidth', None)


print(wdf)

这是 JSON 文件:

https://pastebin.com/sLGCFCDC

【问题讨论】:

  • 你能发布代码sn-p而不是截图吗?很难根据屏幕截图构建代码。示例数据集也很有用。
  • 你好,我已经编辑了我的帖子,你可以看看。谢谢你的时间:)
  • 我已将我的解决方案添加为答案,如果它解决了您的问题,请随时查看并接受答案。

标签: python pandas dataframe validation jupyter-notebook


【解决方案1】:

在您的代码之上的以下代码 sn-p 将完成所需的任务:

min_prices = []
max_prices = []
offer_prices = []
for i,row in wdf.iterrows():
    if('showMinMaxPrice' in row['primaryOffer']):
        min_prices.append(row['primaryOffer']['minPrice'])
        max_prices.append(row['primaryOffer']['maxPrice'])
        offer_prices.append('N/A')
    else:
        min_prices.append('N/A')
        max_prices.append('N/A')
        offer_prices.append(row['primaryOffer']['offerPrice'])

wdf['minPrice'] = min_prices
wdf['maxPrice'] = max_prices
wdf['offerPrice'] = offer_prices

在这里,我们从名为“primaryOffer”的列中的 json 中检查“showMinMaxPrice”元素。对于 minPrice 和 maxPrice 可用的情况,offerPrice 显示为“N/A”,反之亦然。这些首先存储在列表中,然后作为列添加到数据框中。

wdf.head() 的输出将是:

【讨论】:

  • 您好,非常感谢。但我也想分享一个我发现的不同解决方案。通过使用 json_normalize
猜你喜欢
  • 2021-12-03
  • 2022-11-22
  • 2017-09-21
  • 2021-03-21
  • 2018-02-24
  • 2020-08-13
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多