【问题标题】:Classification and random forests in Python: predictions are the same regardless of predictorsPython 中的分类和随机森林:无论预测变量如何,预测都是相同的
【发布时间】:2016-02-24 11:24:00
【问题描述】:

我正在处理一个包含 5 个变量和约 90k 观察值的小型数据集。我已经尝试拟合一个随机森林分类器来模仿来自http://blog.yhathq.com/posts/random-forests-in-python.html 的鸢尾花示例。但是,我的挑战是我的预测值都是相同的:0。我是 Python 新手,但对 R 很熟悉。不确定这是否是编码错误,或者这是否意味着我的数据是垃圾。

from sklearn.ensemble import RandomForestClassifier
data = train_df[cols_to_keep]
data = data.join(dummySubTypes.ix[:, 1:])
data = data.join(dummyLicenseTypes.ix[:, 1:])
data['is_train'] = np.random.uniform(0, 1, len(data)) <= .75
#data['type'] = pd.Categorical.from_codes(data['type'],["Type1","Type2"])
data.head()
Mytrain, Mytest = data[data['is_train']==True], data[data['is_train']==False]
Myfeatures = data.columns[1:5] # string of feature names: subtype dummy     variables
rf = RandomForestClassifier(n_jobs=2)
y, _ = pd.factorize(Mytrain['type'])
rf.fit(Mytrain[Myfeatures], y)
data.target_names = np.asarray(list(set(data['type'])))
preds = data.target_names[rf.predict(Mytest[Myfeatures])]

预测一类,Type1:

In[583]: pd.crosstab(Mytest['type'], preds, rownames=['actual'], colnames ['preds'])
Out[582]: 
preds          Type1
actual                   
Type1          17818
Type2          7247

更新: 前几行数据:

In[670]: Mytrain[Myfeatures].head()
Out[669]: 
subtype_INDUSTRIAL  subtype_INSTITUTIONAL  subtype_MULTIFAMILY  \
0                   0                      0                    0   
1                   0                      0                    0   
2                   0                      0                    0   
3                   0                      0                    0   
4                   0                      0                    0   

subtype_SINGLE FAMILY / DUPLEX  
0                               0  
1                               0  
2                               0  
3                               1  
4                               1 

当我对训练输入进行预测时,我只得到一个类别的预测:

In[675]: np.bincount(rf.predict(Mytrain[Myfeatures]))
Out[674]: array([    0, 75091])

【问题讨论】:

    标签: python-2.7 scikit-learn


    【解决方案1】:

    您的代码有几个问题,但最明显的是:

    data.target_names = np.asarray(list(set(data['type'])))
    preds = data.target_names[rf.predict(Mytest[Myfeatures])]
    

    python中的集合是固有的无序,因此没有保证在此操作后将正确标记预测。

    这是代码的清理版本:

    # build your data
    data = train_df[cols_to_keep]
    data = data.join(dummySubTypes.ix[:, 1:])
    data = data.join(dummyLicenseTypes.ix[:, 1:])
    
    # split into training/testing sets
    from sklearn.cross_validation import train_test_split
    train, test = train_test_split(data, train_size=0.75)
    
    # fit the classifier; scikit-learn factorizes labels internally
    features = data.columns[1:5]
    target = 'type'
    rf = RandomForestClassifier(n_jobs=2)
    rf.fit(train[features], train[target])
    
    # predict and compute confusion matrix
    preds = rf.predict(test[features])
    print(pd.crosstab(test[target], preds,
                      rownames=['actual'],
                      colnames=['preds']))
    

    如果结果仍然没有您的需求,我建议使用Scikit-Searn的grid_search工具在随机林上进行一些封路计优化。

    【讨论】:

    • 非常感谢您的建议。绝对是我在书中找不到的东西。我会看看你的建议。 span>
    • Re:您的评论是关于套装的无序,我试图从这里的全息答案中获取唯一值:stackoverflow.com/questions/12897374/…你会说这是一个错误的答案吗? span>
    • 这肯定是获得唯一值的有效方式(尽管np.unique会更快),但值以任意顺序返回。当您在下一行中索引到它们时,这成为问题。 span>
    猜你喜欢
    • 2014-08-07
    • 2021-02-11
    • 2014-02-28
    • 2021-03-21
    • 2019-05-04
    • 2020-08-15
    • 2014-08-17
    • 2016-04-09
    相关资源
    最近更新 更多