【问题标题】:Problem after groupby (pandas), the grouped column is not accessiblegroupby(熊猫)之后的问题,分组列不可访问
【发布时间】:2021-09-30 12:59:43
【问题描述】:

我在 groupby 之后遇到问题并收到此错误消息:

Traceback(最近一次调用最后一次): 文件“C:\Users\User\PycharmProjects\HashTag_Curso\venv\lib\site-packages\pandas\core\indexes\base.py”,第 3080 行,在 get_loc 返回 self._engine.get_loc(casted_key) 文件“pandas_libs\index.pyx”,第 70 行,在 pandas._libs.index.IndexEngine.get_loc 文件“pandas_libs\index.pyx”,第 101 行,在 pandas._libs.index.IndexEngine.get_loc 文件“pandas_libs\hashtable_class_helper.pxi”,第 4554 行,在 pandas._libs.hashtable.PyObjectHashTable.get_item 文件“pandas_libs\hashtable_class_helper.pxi”,第 4562 行,在 pandas._libs.hashtable.PyObjectHashTable.get_item KeyError:'Ano'

上述异常是以下异常的直接原因:

Traceback(最近一次调用最后一次): 文件“C:/Users/User/PycharmProjects/Bibliotecas/Exemplo.py”,第 11 行,在 x = dfg['Ano'] getitem 中的文件“C:\Users\User\PycharmProjects\HashTag_Curso\venv\lib\site-packages\pandas\core\frame.py”,第 3024 行 索引器 = self.columns.get_loc(key) 文件“C:\Users\User\PycharmProjects\HashTag_Curso\venv\lib\site-packages\pandas\core\indexes\base.py”,第 3082 行,在 get_loc 从错误中引发 KeyError(key) KeyError:'Ano'

import pandas as pd
from matplotlib import pyplot as plt
import numpy as np
from astropy.stats import biweight_midcorrelation as bw_cor

df = pd.read_csv(r'Bases_dados\D_1_4M\Tudo/combined.csv').iloc[:100000]
df['Ano'] = df['Data decimal']//1
dfg = df.groupby(by=["Ano"]).mean()

print(dfg)
x = dfg['Ano']
y = dfg['Lances']

r = np.corrcoef(x, y)[0][1]
bwr = bw_cor(x, y)

print(bwr, r)
plt.scatter(x, y)
plt.show()

如果我使用 x = df['Ano'] y = df['Lances']

工作正常,但使用 dfg(按“Ano”分组),我会收到错误消息。

当我打印(dfg)时,“Ano”列正常显示。

【问题讨论】:

    标签: pandas pandas-groupby


    【解决方案1】:

    它已移至索引部分,因此您可以reset_index 或将as_index=False 传递给 groupby 以开始:

    dfg = df.groupby(by="Ano", as_index=False).mean()
    

    【讨论】:

    • 非常感谢!解决!
    猜你喜欢
    • 1970-01-01
    • 2018-07-19
    • 2017-01-11
    • 1970-01-01
    • 2021-03-21
    • 1970-01-01
    • 2021-03-30
    • 2014-08-02
    • 1970-01-01
    相关资源
    最近更新 更多