【问题标题】:StandardScaler() python error for scaling data用于缩放数据的 StandardScaler() python 错误
【发布时间】:2021-09-09 14:49:21
【问题描述】:

如何修复此代码,是否需要将 features_trainfeatures_test 设置为 DataFrame? 任何人都知道如何修复该代码?实在看不懂问题。。。。

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.preprocessing import Normalizer
from sklearn.metrics import r2_score

admissions_data = pd.read_csv('admissions_data.csv')
labels = admissions_data.iloc[:, -1]
features = admissions_data.iloc[:, 1:8]
features_train, labels_train, features_test, labels_test = train_test_split(features, labels, test_size=0.2, random_state=13)
sc = StandardScaler()
features_train_scaled = sc.fit_transform(features_train)
features_test_scale = sc.transform(features_test)
features_train_scaled = pd.DataFrame(features_train_scaled)
features_test_scale = pd.DataFrame(features_test_scale)

错误是:

    Traceback (most recent call last):
      File "script.py", line 26, in <module>
        features_test_scale = sc.transform(features_test)
      File "/usr/local/lib/python3.6/dist-packages/sklearn/preprocessing/_data.py", line 794, in transform
        force_all_finite='allow-nan')
      File "/usr/local/lib/python3.6/dist-packages/sklearn/base.py", line 420, in _validate_data
        X = check_array(X, **check_params)
      File "/usr/local/lib/python3.6/dist-packages/sklearn/utils/validation.py", line 73, in inner_f
        return f(**kwargs)
      File "/usr/local/lib/python3.6/dist-packages/sklearn/utils/validation.py", line 624, in check_array
        "if it contains a single sample.".format(array))
    ValueError: Expected 2D array, got 1D array instead:
    array=[0.57 0.78 0.59 0.64 0.47 0.63 0.65 0.89 0.84 0.73 0.75 0.64 0.46 0.78
     0.62 0.53 0.85 0.67 0.84 0.94 0.64 0.53 0.47 0.86 0.62 0.7  0.77 0.61
     0.61 0.63 0.86 0.82 0.65 0.58 0.7  0.7  0.84 0.72 0.71 0.77 0.69 0.8
     0.52 0.62 0.79 0.71 0.9  0.84 0.6  0.86 0.67 0.61 0.71 0.52 0.62 0.37
     0.73 0.64 0.71 0.8  0.88 0.78 0.45 0.62 0.62 0.86 0.74 0.94 0.58 0.7
     0.92 0.64 0.65 0.83 0.34 0.66 0.67 0.7  0.71 0.54 0.68 0.61 0.68 0.79
     0.57 0.94 0.59 0.79 0.73 0.91 0.86 0.95 0.9  0.92 0.68 0.84 0.69 0.72
     0.94 0.53 0.45 0.77 0.77 0.91 0.61 0.78 0.77 0.82 0.9  0.92 0.54 0.92
     0.72 0.5  0.68 0.78 0.72 0.53 0.79 0.49 0.68 0.72 0.73 0.93 0.72 0.52
     0.54 0.86 0.65 0.93 0.89 0.72 0.34 0.64 0.96 0.79 0.73 0.49 0.73 0.94
     0.7  0.95 0.65 0.86 0.78 0.75 0.89 0.94 0.91 0.87 0.93 0.81 0.94 0.89
     0.57 0.77 0.39 0.46 0.78 0.64 0.76 0.58 0.56 0.53 0.79 0.9  0.92 0.96
     0.67 0.65 0.64 0.58 0.94 0.76 0.78 0.88 0.84 0.68 0.66 0.42 0.56 0.66
     0.46 0.65 0.58 0.72 0.48 0.68 0.89 0.95 0.46 0.71 0.79 0.52 0.57 0.76
     0.52 0.8  0.77 0.91 0.75 0.49 0.72 0.72 0.61 0.97 0.8  0.85 0.73 0.64
     0.87 0.63 0.97 0.72 0.82 0.54 0.71 0.45 0.8  0.49 0.77 0.93 0.89 0.93
     0.81 0.62 0.81 0.66 0.78 0.76 0.48 0.61 0.82 0.68 0.7  0.68 0.62 0.81
     0.87 0.94 0.38 0.67 0.64 0.84 0.62 0.7  0.62 0.5  0.79 0.78 0.36 0.77
     0.57 0.87 0.74 0.71 0.61 0.57 0.64 0.73 0.81 0.74 0.8  0.69 0.66 0.64
     0.93 0.64 0.59 0.71 0.82 0.69 0.69 0.89 0.93 0.74 0.64 0.84 0.91 0.97
     0.55 0.74 0.72 0.71 0.93 0.96 0.8  0.8  0.81 0.88 0.64 0.38 0.87 0.73
     0.78 0.89 0.56 0.61 0.76 0.46 0.78 0.71 0.81 0.59 0.47 0.7  0.42 0.76
     0.8  0.67 0.94 0.65 0.51 0.73 0.9  0.8  0.65 0.7  0.96 0.96 0.73 0.79
     0.86 0.89 0.85 0.76 0.76 0.71 0.83 0.76 0.42 0.9  0.58 0.66 0.86 0.71
     0.8  0.51 0.65 0.58 0.76 0.8  0.7  0.61 0.71 0.69 0.95 0.72 0.79 0.97
     0.74 0.96 0.47 0.56 0.73 0.94 0.76 0.79 0.71 0.58 0.94 0.66 0.75 0.76
     0.84 0.59 0.68 0.75 0.76 0.72 0.87 0.78 0.67 0.79 0.91 0.57 0.77 0.69
     0.73 0.43 0.93 0.68 0.82 0.67 0.74 0.82 0.85 0.62 0.54 0.71 0.92 0.85
     0.79 0.63 0.59 0.73 0.66 0.74 0.9  0.81].
    Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.

【问题讨论】:

  • 错误很明显:StandardScaler, (transform) 需要一个二维数组,但是你传递了一个一维数组。如错误消息所示,为您的数据添加另一个维度。打印features_train, labels_train, features_test, labels_test 形状以告诉您究竟需要更改什么。
  • 所以,我需要做np.array.reshape(1, -1)?
  • 它告诉我这是一个系列并且它没有任何 reshape 属性
  • 新错误:``` Traceback(最近一次调用最后):文件“script.py”,第 26 行,在 features_test_scale = sc.transform(features_test) 文件“/usr/local /lib/python3.6/dist-packages/sklearn/preprocessing/_data.py”,第 794 行,在变换 force_all_finite='allow-nan') 文件“/usr/local/lib/python3.6/dist-packages/ sklearn/base.py”,第 436 行,在 _validate_data ```
  • @ItamarCohen 请编辑您的问题以添加任何其他信息。

标签: python pandas scikit-learn


【解决方案1】:

您在拆分数据时犯了一个错误。那是因为您错误地将一维的labels_train 设置为features_test,并且由于transform 函数不期望一维数组,所以它返回错误。

train_test_split() 分别返回features_train, features_test, label_train, labels_test

所以,像这样改变你的代码:

#features_train, labels_train, features_test, labels_test = train_test_split(features, labels, test_size=0.2, random_state=13)
features_train, features_test, label_train, labels_test = train_test_split(features, labels, test_size=0.2, random_state=13) 

【讨论】:

    猜你喜欢
    • 2017-03-17
    • 2019-09-07
    • 1970-01-01
    • 2018-06-22
    • 2020-01-10
    • 2020-02-10
    • 2021-11-15
    • 2021-04-19
    • 2016-11-20
    相关资源
    最近更新 更多