【问题标题】:Debug the "NaN" error of sklearn using pandas dataframe input使用 pandas 数据框输入调试 sklearn 的“NaN”错误
【发布时间】:2016-09-09 02:32:53
【问题描述】:

我将我的输入读取为 pandas 数据框并通过以下方式填充 NaN:

df = df.fillna(0)

之后,我分成训练集和测试集,并使用 sklearn 进行分类。

features = df.drop('class',axis=1)
labels = df['class']
features_train, features_test, labels_train, labels_test = train_test_split(features, labels, test_size=0.3, random_state=42)
clf.fit(features_train, labels_train)   

但还是有错误

“NaN 错误”:ValueError:输入包含 NaN、无穷大或对于 dtype('float32') 来说太大的值。

fillna() 似乎没有找到丢失的数据。我怎样才能找到“NaN”在哪里?

【问题讨论】:

  • "无穷大或值对于 dtype('float32') 来说太大了" 你也检查过这两种情况吗?
  • features.dtypes 显示列的类型是 int64 和 float64,对于无穷大,不,我没有
  • 您的任何功能是否包含字符串?我不认为 dropna 会考虑他们 NaN
  • 数据是由特征计算工具生成的,所以都是数值特征。但如果任何特征包含字符串,我可以使用 features.dtypes 找到它,对吧?
  • 您可以尝试使用imputer 类来估算数据,请访问文档!

标签: python pandas scikit-learn


【解决方案1】:
df.isnull().sum()

这可以告诉你数据框内是否/在哪里存在任何 NaN

【讨论】:

    【解决方案2】:

    TLDR:pip install pandas --upgrade

    我自己今天遇到了这个问题。在处理全零的稀疏数组时,似乎是 sklearn 的 train_test_split() 方法的问题。我在 scikit-learns github repo 上提出了一个错误,他们非常迅速地回应了升级 pandas 的解决方案: https://github.com/scikit-learn/scikit-learn/issues/22133

    重播的步骤/代码

    import numpy as np
    import pandas as pd
    from scipy import sparse
    from sklearn.model_selection import train_test_split
    
    X = pd.DataFrame.sparse.from_spmatrix(sparse.eye(5))
    y = pd.Series(np.zeros(5))
    
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    # output as expected (when every input column has at least one non zero value)
    print(X_train)
    
    X_train, X_test, y_train, y_test = train_test_split(X[1:], y[1:], test_size=0.2, random_state=42)
    # output column contains all NaN (when input column contains all zero's)
    print(X_train)
    

    第一个 train_test_split() 按预期输出,因为每列至少有一个非零行,但是第二个在第一列上输出 NaN,因为所有行都为零。

        0    1    2    3    4
     --------------------------
     4  0.0  0.0  0.0  0.0  1.0
     2  0.0  0.0  1.0  0.0  0.0
     0  1.0  0.0  0.0  0.0  0.0
     3  0.0  0.0  0.0  1.0  0.0
    
       0    1    2    3    4
     -------------------------
     4 NaN  0.0  0.0  0.0  1.0
     1 NaN  1.0  0.0  0.0  0.0
     3 NaN  0.0  0.0  1.0  0.0
    

    【讨论】:

      【解决方案3】:

      你问

      我怎样才能找到“NaN”在哪里

      将有问题的数据在框架中的位置可视化会有帮助吗?

      你可以试试matplotlib.pyplot.spy

      import pandas as pd
      import numpy as np
      import matplotlib.pyplot as plt
      
      # lets make some initial clean data
      df = pd.DataFrame(
          data={
              'alpha': [0, 1, 2],
              'beta': [3, 4, 5],
              'gamma': [6, 7, 8]
          },
          index=['one', 'two', 'three']
      )
      # add some problematic points
      # `NaN`s, infinities and stuff that is 
      #  just not numeric
      df.loc['one', 'beta'] = 'not a number but not NaN'
      df.loc['two', 'alpha'] = np.NaN
      df.loc['three', 'gamma'] = np.infty
      
      fig, axes = plt.subplots(1, 3)
      axes[0].spy(df.isnull())
      axes[0].set_title('NaN elements')
      axes[1].spy(df == np.infty)
      axes[1].set_title('infinite elements')
      axes[2].spy(~df.applymap(np.isreal))
      axes[2].set_title('Non numeric elements')
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2018-11-11
        • 2017-12-20
        • 2014-08-29
        • 1970-01-01
        • 2018-10-23
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多