【问题标题】:Importing complex Excel with pandas使用 pandas 导入复杂的 Excel
【发布时间】:2020-09-19 15:38:57
【问题描述】:

我从复杂的 Excel 文件导入数据时遇到问题。该文件如下所示:

我要导入数据,因此需要执行两个步骤:

  1. 导入前删除前 2 行(不需要此数据)
  2. 转动“耳机”、“电视”和“收音机”

我该怎么做?

编辑:

我想在名为“产品”的“名称”后面的列中添加“耳机”、“电视”和“收音机”。

所以列应该是:“国家”“名称”“产品”“销售”“价格”“畅销书”

示例:

US | Tom | Headphones | 1200 | 100 | Headphone 1 
US | Tom | TV         | 1546 | 500 | TV 1

【问题讨论】:

  • 这里是如何删除前 x 行的示例:stackoverflow.com/questions/59725186/…(示例显示如何使用 4 行执行此操作)。在此之后,您在第 2 步中想要做什么就很不清楚了,您的“预期输出”如何?
  • 谢谢!它帮助我摆脱了第一行。现在我有一个问题,我想在名为“产品”的“名称”之后的列中出现“耳机”、“电视”和“收音机”。所以列应该是: "Country" "Name" "product" "Sales" "Price" "Bestseller" 一个例子,因为在这个论坛中似乎不可能放一张表,看起来像这样: US |汤姆 |耳机 | 1200 | 100 |耳机 1 和下一行:美国 |汤姆 |电视 |第1546章500 | TV 1 我希望这是可以理解的。
  • 我添加了您对您问题的回复......
  • 我找到了一种实现表格的方法,并在下面添加了它。我希望这能让我更清楚我需要什么。感谢您的帮助
  • 完美!谢谢卢克。我是论坛的新手,不知道该怎么做。看起来不错

标签: python excel pandas dataframe import


【解决方案1】:

Jerry,我对 python 还很陌生,但看看这个:

import pandas as pd

e=pd.ExcelFile("sales.xlsx")
df = e.parse(skiprows=2)

for index, row in df.iterrows():
  if (index==0):
      print("Country|Name|product|Sales|Price|Bestseller")
  else:
      print(row[0], row[1], "Headphones", row[2], row[3],row[4], sep='|')
      print(row[0], row[1], "TV", row[5], row[6], row[7], sep='|')
      print(row[0], row[1], "Radio", row[8], row[9], row[10],sep='|')

输出:

Country|Name|product|Sales|Price|Bestseller
US|Tom|Headphones|1200|100|Headphone 1
US|Tom|TV|1200|100|Headphone 1
US|Tom|Radio|1200|100|Headphone 1
CA|Megan|Headphones|2300|110|Headphone 2
CA|Megan|TV|2300|110|Headphone 2
CA|Megan|Radio|2300|110|Headphone 2
UK|Ryan|Headphones|1156|120|Headphone 1
UK|Ryan|TV|1156|120|Headphone 1
UK|Ryan|Radio|1156|120|Headphone 1

这段代码给出了很多改进点,因为(即)我对列名进行了“硬编码”。

编辑(因为需要数据框作为输出:

mycolumns = ['Country','Name','product','Sales','Price','Bestseller']
i = 0

d = pd.DataFrame(columns=mycolumns)
for index, row in df.iterrows():
   if (index>0):
      d.loc[i] = [row[0], row[1],"Headphones",row[2],row[3],row[4]] 
      i = i + 1
      d.loc[i] = [row[0], row[1],"TV",row[5],row[6],row[7]] 
      i = i + 1
      d.loc[i] = [row[0], row[1],"Radio",row[8],row[9],row[10]] 
      i = i + 1

更多信息DataFrame

【讨论】:

  • 嗨 Luuk,非常感谢您的帮助!我非常感谢您为此付出的努力。可悲的是,我认为我错误地描述了我需要的东西。数据最后需要在数据框中,所以我不能那样构建它。我想这与多索引和旋转它有关,但由于我是 python 新手,我不知道如何完成这个......
【解决方案2】:

非常感谢,这有助于摆脱第一行。我现在拥有的是:

Unnamed: 0 	  Unnamed: 1	Headphones 	Unnamed: 3 	Unnamed: 4    TV      ...
  0 	Country 	Name 	      Sales 	    Price       Bestseller    Sales   ...
  1 	US 	        Tom 	      1200 	      100 	Headphone 1   1546    ...
  
  # And what I want is
  
  Country Name product     Sales Price Bestseller
  US 	  Tom  Headphones  1200  100   Headphone 1
  US      Tom  TV          1546  500   TV 1

【讨论】:

    猜你喜欢
    • 2013-05-01
    • 1970-01-01
    • 2020-05-15
    • 2019-08-19
    • 1970-01-01
    • 2021-12-01
    • 1970-01-01
    • 2014-11-20
    • 2018-10-28
    相关资源
    最近更新 更多