【问题标题】:Loading CSV to Scikit Learn将 CSV 加载到 Scikit Learn
【发布时间】:2020-12-15 23:35:21
【问题描述】:

我是 python 新手,但我正在尝试使用一堆不同的变量进行回归。到目前为止,我已经把它归结为 Scikit。我一直在寻找几个小时,但似乎无法找到一种方法来导入数据,然后在返回每个变量的系数时对其进行线性回归。任何帮助深表感谢。我有 15 列要针对 X 运行。

X = Margin
Ys = A1, B1, C1, D1 etc. 

示例如下:

Margin,A1
-8,110.7
-10,112
-1,106.7
9,109
-9,107.5
1,108.1
-19,109.2

这是我到目前为止所得到的,我知道它并不多

import pandas as pd

data = pd.read_csv("NBA.csv")

【问题讨论】:

标签: python scikit-learn regression


【解决方案1】:

作为机器学习的惯例,我们将 X 视为特征,将 Y 视为目标。

如果要运行线性回归并提取系数,可以执行以下操作:

# import the needed libraries
import pandas as pd
from sklearn.linear_model import LinearRegression

# Import the data
data = pd.read_csv("NBA.csv")

# Specify the features and the target
target = 'Margin'
features = data.columns.tolist() # This is the column names of your data as a list
features.remove(target) # We remove the target from the list of features

# Train the model
model = LinearRegression() # Instantiate the model
model.fit(data[features].values, data[target].values) # fit the model to the data
print(features) # Returns the name of each feature
print(model.coef_) # Returns the coefficients for each feature (in the same order of your features)

【讨论】:

  • 这太棒了!谢谢你。我对 data.columns 有一个问题,我是否放置 data.A1(列名是 A1)?我需要列出每一列吗?就像我会做 data.A1、data.A2 还是如何做?对不起,我是新手,只是想学得最好。谢谢!我收到了错误: Traceback (last recent call last): line 11, in features.remove(target) # We remove the target from the list of features AttributeError: 'Index' object has no attribute 'remove'
  • 我编辑了我的答案。我们需要将 data.columns 转换为列表以使用 .remove
  • 如果你想要 p 值、R 平方等。我建议你看看另一个库,比如 statsmodels。这是一个很棒的教程:datatofish.com/statsmodels-linear-regression
  • 太棒了!谢谢 有没有办法获得类似于在 Excel 或 R 中运行回归时所做的总结报告?基本上会喜欢它的 R 平方和每个的 p 值。不确定这是否会改变一切
  • 太棒了!谢谢!超级有用的东西
猜你喜欢
  • 2017-03-05
  • 2015-08-31
  • 2013-01-29
  • 2018-05-15
  • 2015-02-24
  • 1970-01-01
  • 2017-02-16
  • 1970-01-01
  • 2015-07-18
相关资源
最近更新 更多