【问题标题】:Extract sheet from many excel workbooks based on sheet title根据工作表标题从许多 Excel 工作簿中提取工作表
【发布时间】:2021-09-23 22:11:03
【问题描述】:

我需要从许多 Excel 工作簿中提取特定工作表。工作表在每个 Excel 工作簿中的标题完全相同。

提取后,我需要根据相应 Excel 工作簿标题的开头命名每个数据框(从每个提取的工作表创建)。 示例:对于标题为“Pizza”的工作表(每个 Excel 工作簿都相同)和标题为“Coke_2021”的 Excel 工作簿,数据框应自动命名为“Pizza_Coke”。 Excel工作簿的格式为:'Coke_2021'、'Sprite_2019'等,因此非常可预测。

我有以下代码,但卡在第 1 步(提取工作表)。

import tkinter as tk
from tkinter import filedialog
from tkinter import messagebox
import pyodbc
import openpyxl
from openpyxl.utils.dataframe import dataframe_to_rows
from openpyxl import load_workbook 
from openpyxl.worksheet.table import Table
import xlwings as xw
import pandas as pd
import datetime
import numpy as np
import itertools
import ntpath
import calendar

## UI - Asking user for their input and output files
root = tk.Tk()
root.withdraw()
root.databases =  filedialog.askopenfilenames(initialdir = "C:/",title = "Select the location of your Soda files (you may select multiple files)", filetypes = (("Excel Files","*.xlsx"),("all files","*.*")))
db_list = root.tk.splitlist(root.databases)

【问题讨论】:

  • “标记每个数据框”是什么意思?
  • 如果您的问题是第一步,为什么这会被标记为“tkinter”? tkinter 似乎与此问题无关。

标签: python excel pandas database extract


【解决方案1】:

我不太清楚你命名每个数据框是什么意思,但我们可以加载数据框并生成所需的名称。

Pandas 在读取 Excel 和 CSV 文件时非常有用。 pandas.read_excel() 函数可以将特定 Excel 工作表读入数据框,也可以将所有工作表作为数据框字典读入。详情请见:https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_excel.html?highlight=read_excel。

下面的代码从 Excel 文件列表中加载“Pizza”表,并将它们存储为 dataframes 的字典,命名为“sheet_file”。

import pandas as pd

sheet_name = 'Pizza'
xl_files = ['Coke_2021.xlsx', 'Sprite_2019.xlsx']

dataframes = {}
for xl_file in xl_files:
    # create new name with sheet name & first part of file name
    name = f'{sheet_name}_{xl_file.split("_")[0]}'
    dataframes[name] = pd.read_excel(xl_file, sheet_name=sheet_name)

# now can access a given dataframe by the "sheet_file" naming
# could also easily export to Excel with the new name like below
dataframes['Pizza_Coke'].to_excel('Pizza_Coke.xlsx')

在上面,我们假设您有可用的 Excel 文件名,例如分配给xl_files。从您的代码中,看起来可能存在存储在db_list 中的使用绝对路径的 Excel 文件列表。在这种情况下,我们需要稍微修改名称创建代码以仅使用文件名而不是整个路径。这样的事情应该会有所帮助:

import os

dataframes = {}
for xl_file in db_list:
    filename = os.path.basename(xl_file)  # use this in next line!
    name = f'{sheet_name}_{filename.split("_")[0]}'
    dataframes[name] = pd.read_excel(xl_file, sheet_name=sheet_name)

如果您想从一个 Excel 文件加载多张工作表,下面的代码会生成相同的结果。

import pandas as pd

xl_file = 'Coke_2021.xlsx'

# load the Excel file's sheets into a dictionary
# of form: {sheet_name: sheet_dataframe, ...}
dataframes = pd.read_excel(xl_file, None)

renamed_dataframes = {}
for sheet_name, df in dataframes.items():
    new_name = f'{sheet_name}_{xl_file.split("_")[0]}'
    renamed_dataframes[new_name] = df

renamed_dataframes['Pizza_Coke'].to_excel('Pizza_Coke.xlsx')

【讨论】:

  • 感谢您的帮助!我尝试了您的建议,但效果不佳。由于我一次导入许多 Excel 工作簿,大约 30 个,我无法在代码中具体命名它们。对于您编写的第一部分中的代码xl_files = ['Coke_2021.xlsx', 'Sprite_2019.xlsx'],我相信这必须引用导入的工作簿列表,但是当我引用 db_list(在我的代码示例中)时,它不起作用。不知道问题出在哪里
  • @Sam 我添加了一个示例来帮助使用db_list。
猜你喜欢
  • 2011-06-04
  • 1970-01-01
  • 2021-11-18
  • 1970-01-01
  • 2022-08-22
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多