【问题标题】:Reading excel file in python with pandas and multiple indices用熊猫和多个索引在python中读取excel文件
【发布时间】:2016-05-10 05:01:23
【问题描述】:

我是 python 新手,所以请原谅这个基本问题。 我的 .xlsx 文件看起来像这样

Unnamend:1    A     Unnamend:2    B
2015-01-01    10    2015-01-01    10
2015-01-02    20    2015-01-01    20
2015-01-03    30    NaT           NaN

当我在 Python 中使用 pandas.read_excel(...) 读取它时,pandas 会自动使用第一列作为时间索引。

是否有一条线告诉 pandas 注意到,每隔一列是一个时间索引,属于它旁边的时间序列?

所需的输出如下所示:

date          A     B
2015-01-01    10    10
2015-01-02    20    20
2015-01-03    30    NaN

【问题讨论】:

    标签: python excel pandas timestamp


    【解决方案1】:

    为了解析相邻的columns 块并在它们各自的datetime 索引上对齐,您可以执行以下操作:

    以df开头:

    Int64Index: 3 entries, 0 to 2
    Data columns (total 4 columns):
    Unnamed: 0    3 non-null datetime64[ns]
    A             3 non-null int64
    Unnamed: 1    2 non-null datetime64[ns]
    B             2 non-null float64
    dtypes: datetime64[ns](2), float64(1), int64(1)
    

    您可以像这样在 index 上迭代 2 列和 merge 的块:

    def chunks(l, n):
        """ Yield successive n-sized chunks from l."""
        for i in range(0, len(l), n):
            yield l[i:i + n]
    
    merged = df.loc[:, list(df)[:2]].set_index(list(df)[0])
    for cols in chunks(list(df)[2:], 2):
        merged = merged.merge(df.loc[:, cols].set_index(cols[0]).dropna(), left_index=True, right_index=True, how='outer')
    

    得到:

                 A   B
    2015-01-01  10  10
    2015-01-01  10  20
    2015-01-02  20 NaN
    2015-01-03  30 NaN
    

    pd.concat 不幸的是不起作用,因为它无法处理重复的index 条目,否则可以使用list comprehension:

    pd.concat([df.loc[:, cols].set_index(cols[0]) for cols in chunks(list(df), 2)], axis=1)
    

    【讨论】:

    • 嗨斯特凡。假设在我的示例系列 A 和 B 中切换索引,使得 B 现在是最长的系列。如果我默认选择 index_col=0 会不会导致索引值缺失(即缺失“2015-01-03”)?
    • 确实如此。人们还会认为您想合并日期上的列。如果没有必要,我们可以添加一个步骤,使最长的column 变为index。如果你想merge 而不是我们当然必须采取不同的方法。
    • 合并和采用最长索引对齐该索引上的所有其他系列之间的确切区别是什么?我来自 R,这里的神奇词确实是“合并”或 cbind...
    • 差异在于对齐 - 您希望值在日期上对齐还是仅保持行顺序?
    【解决方案2】:

    我用xlrd导入数据,用pandas显示后

    import xlrd
    import pandas as pd
    workbook = xlrd.open_workbook(xls_name)
    workbook = xlrd.open_workbook(xls_name, encoding_override="cp1252")
    worksheet = workbook.sheet_by_index(0)
    first_row = [] # The row where we stock the name of the column
    for col in range(worksheet.ncols):
        first_row.append( worksheet.cell_value(0,col) )
    data =[]
    for row in range(10, worksheet.nrows):
        elm = {}
        for col in range(worksheet.ncols):
              elm[first_row[col]]=worksheet.cell_value(row,col)
        data.append(elm)
    
    first_column=second_column=third_column=[]
    for elm in data :
        first_column.append(elm(first_row[0]))
        second_column.append(elm(first_row[1]))
        third_column.append(elm(first_row[2]))
    
    dict1={}
    dict1[first_row[0]]=first_column
    dict1[first_row[1]]=second_column
    dict1[first_row[2]]=third_column
    res=pd.DataFrame(dict1, columns=['column1', 'column2', 'column3'])
    print res
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2017-02-20
      • 2020-02-10
      • 2016-09-20
      • 2019-01-26
      • 1970-01-01
      • 1970-01-01
      • 2021-11-11
      相关资源
      最近更新 更多