【发布时间】:2019-12-29 22:12:24
【问题描述】:
我尝试使用线性回归处理其中一列中的缺失值。
列的名称是“Landsize”,我正在尝试使用其他几个变量通过线性回归预测 NaN 值。
这是林。回归代码:
# Importing the dataset
dataset = pd.read_csv('real_estate.csv')
from sklearn.linear_model import LinearRegression
linreg = LinearRegression()
data = dataset[['Price','Rooms','Distance','Landsize']]
#Step-1: Split the dataset that contains the missing values and no missing values are test and train respectively.
x_train = data[data['Landsize'].notnull()].drop(columns='Landsize')
y_train = data[data['Landsize'].notnull()]['Landsize']
x_test = data[data['Landsize'].isnull()].drop(columns='Landsize')
y_test = data[data['Landsize'].isnull()]['Landsize']
#Step-2: Train the machine learning algorithm
linreg.fit(x_train, y_train)
#Step-3: Predict the missing values in the attribute of the test data.
predicted = linreg.predict(x_test)
#Step-4: Let’s obtain the complete dataset by combining with the target attribute.
dataset.Landsize[dataset.Landsize.isnull()] = predicted
dataset.info()
当我尝试检查回归结果时,我得到了这个错误:
ValueError: Input contains NaN, infinity or a value too large for dtype('float64').
准确度:
accuracy = linreg.score(x_test, y_test)
print(accuracy*100,'%')
【问题讨论】:
-
您是否将“Nan”转换为“numeric nan 值”?
标签: python pandas scikit-learn linear-regression