【发布时间】:2017-06-13 12:19:42
【问题描述】:
我正在尝试绘制随机森林模型的特征重要性并将每个特征重要性映射回原始系数。我设法创建了一个显示重要性的图,并使用原始变量名称作为标签,但现在它按照变量名称在数据集中的顺序(而不是重要性顺序)对变量名称进行排序。如何按功能重要性的顺序对它们进行排序?谢谢!
我的代码是:
importances = brf.feature_importances_
std = np.std([tree.feature_importances_ for tree in brf.estimators_],
axis=0)
indices = np.argsort(importances)[::-1]
# Print the feature ranking
print("Feature ranking:")
for f in range(x_dummies.shape[1]):
print("%d. feature %d (%f)" % (f + 1, indices[f], importances[indices[f]]))
# Plot the feature importances of the forest
plt.figure(figsize=(8,8))
plt.title("Feature importances")
plt.bar(range(x_train.shape[1]), importances[indices],
color="r", yerr=std[indices], align="center")
feature_names = x_dummies.columns
plt.xticks(range(x_dummies.shape[1]), feature_names)
plt.xticks(rotation=90)
plt.xlim([-1, x_dummies.shape[1]])
plt.show()
【问题讨论】:
-
你没有包括你目前得到的情节?
-
已编辑!我不确定该情节增加了多少价值,因为我只是想更改底部 x 标签的顺序。为小字体道歉,这是将大部分图片放入屏幕截图的唯一方法。
-
plt.bar(range(x_dummies.shape[1]), importances[indices], color="r", yerr=std[indices], align="center")? -
你说得对,x_train 应该是 x_dummies,但不幸的是,这并没有改变情节。