【问题标题】:Trouble with NaNs: set_index().reset_index() corrupts dataNaN 的问题:set_index().reset_index() 损坏数据
【发布时间】:2013-05-06 20:28:55
【问题描述】:

我读到 NaN 是有问题的,但以下导致我的数据实际损坏,而不是错误。这是一个错误吗?我错过了文档中的一些基本内容吗? 我希望第二个命令给出错误或给出与第一个命令相同的响应:

ipdb> df
    year  PRuid  QC       data
18  2007  nonQC   0  8.014261
19  2008  nonQC   0  7.859152
20  2010  nonQC   0  7.468260
21  1985     10 NaN  0.861403
22  1985     11 NaN  0.878531
23  1985     12 NaN  0.842704
24  1985     13 NaN  0.785877
25  1985     24   1  0.730625
26  1985     35 NaN  0.816686
27  1985     46 NaN  0.819271
28  1985     47 NaN  0.807050
ipdb> df.set_index(['year','PRuid','QC']).reset_index()
    year  PRuid  QC       data
0   2007  nonQC   0  8.014261
1   2008  nonQC   0  7.859152
2   2010  nonQC   0  7.468260
3   1985     10   1  0.861403
4   1985     11   1  0.878531
5   1985     12   1  0.842704
6   1985     13   1  0.785877
7   1985     24   1  0.730625
8   1985     35   1  0.816686
9   1985     46   1  0.819271
10  1985     47   1  0.807050

“QC”的值实际上是从应该是NaN的NaN变成了1。

顺便说一句,为了对称,我添加了“.reset_index()”,但数据损坏是由 set_index 引入的。

如果这很有趣,版本是:

pd.version
<module 'pandas.version' from '/usr/lib/python2.6/site-packages/pandas-0.10.1-py2.6-linux-x86_64.egg/pandas/version.pyc'>

【问题讨论】:

  • 具有 Nan 值的索引听起来有问题。我在 0.11 并且 set_index 向我显示 QC 索引级别的 NaN 值。但是查看 reset_index 源代码显示 self.index.labels 和 self.index.levels 没有返回正确的 NaN 值。我建议您向 Pandas 团队提交错误。
  • 谢谢。好的,我在github.com/pydata/pandas/issues/3586 提交了一份
  • 并由此修复:github.com/pydata/pandas/pull/3587,这是一个错误,谢谢!

标签: indexing pandas nan


【解决方案1】:

所以这是一个错误。到 2013 年 5 月末,pandas 0.11.1 应该会发布并修复错误(请参阅有关此问题的 cmets)。 同时,我避免在任何多索引中使用带有 NaN 的值,例如,对“QC”列中的 NaN 使用其他标志值 (-99)。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2010-12-13
    • 2012-03-05
    • 1970-01-01
    • 1970-01-01
    • 2013-01-12
    • 2013-12-22
    • 2016-12-05
    相关资源
    最近更新 更多