【问题标题】:Does sklearn use pandas index as a feature?sklearn 是否使用熊猫索引作为功能?
【发布时间】:2019-10-31 15:53:15
【问题描述】:

我将包含各种功能的 pandas DataFrame 传递给 sklearn,我不希望估算器使用数据帧索引作为功能之一。 sklearn 是否使用索引作为特征之一?

df_features = pd.DataFrame(columns=["feat1", "feat2", "target"])
# Populate the dataframe (not shown here)
y = df_features["target"]
X = df_features.drop(columns=["target"])

estimator = RandomForestClassifier()
estimator.fit(X, y)

【问题讨论】:

  • 不,它没有,它将数据框下方的数组作为您的输入。您可以通过print(X.to_numpy()) 进行检查
  • 谢谢@Erfan。你会碰巧知道这发生在 scikit-learn 源代码的什么地方吗?我搜索了“to_numpy”的代码,没有找到。

标签: pandas scikit-learn


【解决方案1】:

不,sklearn 不使用索引作为您的功能之一。它基本上发生在here,当您调用 fit 方法时,将应用check_array 函数。现在,如果您深入研究check_arrayfunction,您会发现您正在使用np.array 函数将输入转换为数组,该函数实质上是从数据帧中剥离索引,如下所示:

import pandas as pd 
import numpy as np
data = [['tom', 10], ['nick', 15], ['juli', 14]] 
df = pd.DataFrame(data, columns = ['Name', 'Age']) 
df  

    Name    Age
0   tom     10
1   nick    15
2   juli    14

np.array(df)
array([['tom', 10],
       ['nick', 15],
       ['juli', 14]], dtype=object)

希望这会有所帮助!

【讨论】:

猜你喜欢
  • 1970-01-01
  • 2018-02-16
  • 1970-01-01
  • 2023-02-07
  • 1970-01-01
  • 2014-02-10
  • 2016-11-27
相关资源
最近更新 更多