【发布时间】:2022-06-11 01:20:07
【问题描述】:
- 如何绘制分类器决定的概率?
- 我想表示概率,因为对于我的用例,算法选择类 1 的概率是 80% 并且是正确的,还是选择 51% 的概率是正确的,这一点很重要。
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
import seaborn as sns
import pandas as pd
X, y = make_classification(n_samples=1000, n_features=4,
n_informative=2, n_redundant=0,
random_state=0, shuffle=False)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=30)
# Train
clf = RandomForestClassifier(max_depth=2, random_state=0)
clf.fit(X_train, y_train)
# Pred
y_pred = clf.predict(X_test)
# Metric
from sklearn.metrics import confusion_matrix
from sklearn import metrics
print("Accuracy:",metrics.accuracy_score(y_test, y_pred))
print("Precison:",metrics.precision_score(y_test, y_pred))
print("Recall:",metrics.recall_score(y_test, y_pred))
print("F1 Maß:",metrics.f1_score(y_test, y_pred))
confusion_matrix(y_test, y_pred)
d = {
"y_test": y_test,
"y_pred": y_pred
}
df = pd.DataFrame(data=d)
def f(x):
if(x['y_pred'] == 1):
if(x['y_test'] == x['y_pred']):
return 'TP'
if(x['y_test'] != x['y_pred']):
return 'FP'
else:
if(x['y_test'] == x['y_pred']):
return 'TN'
if(x['y_test'] != x['y_pred']):
return 'FN'
df['cm'] = df.apply(lambda x: f(x), axis=1)
cleaned_list = []
for element in clf.predict_proba(X_test):
#print(element)
#print(max(element))
cleaned_list.append(max(element))
df['value'] = cleaned_list
sns.displot(df, x="value", hue="cm", kind="kde", fill=True)
【问题讨论】:
-
您熟悉 ROC 或 FNR 与 FPR 曲线吗? scikit-learn.org/stable/modules/generated/…
-
@Learningisamess 感谢您的提示!我听说过
ROC,但FNR vs FPR curve对我来说是新的。 -
它可以帮助您选择正确的决策阈值以匹配您在 FN 和 FP 之间的权衡
-
你想用pythonlink这个吗?
标签: python matplotlib scikit-learn seaborn classification