【问题标题】:Convert a (n_samples, n_features) ndarray to a (n_samples, 1) array of vectors to use as training labels for an sklearn SVM将 (n_samples, n_features) ndarray 转换为 (n_samples, 1) 向量数组以用作 sklearn SVM 的训练标签
【发布时间】:2019-03-18 03:48:48
【问题描述】:

我正在尝试为我正在构建的 SVM 模型计算 ROC 和 AUC。我正在关注this sklearn example 的代码。要求之一是输出标签y 需要二值化。我通过创建MultiLabelBinarizer 并编码所有标签来做到这一点,效果很好。但是,这会创建一个 (n_samples, n_features) ndarray。 classifier.fit(X, y) 函数假定为 y.shape = (n_samples)。我想基本上将y 的列“涂抹”在一起,这样 y[0][0] 将返回整个特征向量V,而不仅仅是V 的第一个值。

这是我的代码:

    enc = MultiLabelBinarizer()
    print("Encoding data...")
    # Fit the encoder onto all possible data values
    print(pandas.DataFrame(enc.fit_transform(df["present"] + df["member"].apply(str).apply(lambda x: [x])),
                           columns=enc.classes_, index=df.index))
    X, y = enc.transform(df["present"]), list(df["member"].apply(str))
    print("Training svm...")
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.5, random_state=0)
    y_train = enc.transform([[x] for x in y_train])  # Strings to 1HotVectors
    svc = svm.SVC(C=1.1, kernel="linear", probability=True, class_weight='balanced')
    svc.fit(X_train, y_train)  # y_train should be 1D but isn't

我得到的例外是:

Traceback (most recent call last):
  File "C:/Users/SawyerPC/PycharmProjects/DiscordSocialGraph/encode_and_train.py", line 129, in <module>
    enc, clf, split_data = encode_and_train(df)
  File "C:/Users/SawyerPC/PycharmProjects/DiscordSocialGraph/encode_and_train.py", line 57, in encode_and_train
    svc.fit(X_train, y_train)  # TODO y_train needs to be flattened to (n_samples,)
  File "C:\Users\SawyerPC\Anaconda3\lib\site-packages\sklearn\svm\base.py", line 149, in fit
    X, y = check_X_y(X, y, dtype=np.float64, order='C', accept_sparse='csr')
  File "C:\Users\SawyerPC\Anaconda3\lib\site-packages\sklearn\utils\validation.py", line 547, in check_X_y
    y = column_or_1d(y, warn=True)
  File "C:\Users\SawyerPC\Anaconda3\lib\site-packages\sklearn\utils\validation.py", line 583, in column_or_1d
    raise ValueError("bad input shape {0}".format(shape))
ValueError: bad input shape (5000, 10)

【问题讨论】:

  • multilabelbinarizer 的目的是将标签编码为“支持的多标签格式:表示存在类标签的(样本 x 类)二进制矩阵”。尝试运行.fit() 方法时出现什么错误?
  • @G.Anderson 我在我的问题中添加了异常的完整堆栈跟踪。
  • 我找不到关于这个问题的任何信息,但你能改用 LabelEncoder 代替 'multilabelbinarizer` 吗?
  • 我尝试使用OneVsRestClassifier 包裹 SVC,就像教程中建议的那样。这实质上是在幕后向 SVC 添加了LabelEncoder。现在唯一的问题是classifier.classes_ 现在返回编码标签而不是字符串,这与我有时比较的MTB.classes_ 冲突。 OneVsRestClassifier 有一个 LabelBinarizer_ 成员,所以我想我可以 inverse_transform 使用它,但事实并非如此。
  • 不幸的是,听起来你做的一切都是正确的,除了尝试完全不同的路线之外,恐怕我无能为力了。希望其他人能看到这一点并提供帮助!

标签: python pandas numpy encoding scikit-learn


【解决方案1】:

我最终通过使用LabelEncoder 解决了这个问题。谢谢@G.Anderson。 flat_member_list 只是标签y 和向量X 中遇到的所有唯一用户ID 的列表。

# Encode "present" users as OneHotVectors
mlb = MultiLabelBinarizer()
print("Encoding data...")
mlb.fit(df["present"] + df["member"].apply(str).apply(lambda x: [x]))

# Encode user labels as ints
enc = LabelEncoder()
flat_member_list = df["member"].apply(str).append(pandas.Series(np.concatenate(df["present"]).ravel()))
enc.fit(flat_member_list)
X, y = mlb.transform(df["present"]), enc.transform(df["member"].apply(str))
print("Training svm...")
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.5, random_state=0, stratify=y)
svc = svm.SVC(C=0.317, kernel="linear", probability=True)
svc.fit(X_train, y_train)

【讨论】:

    猜你喜欢
    • 2018-05-05
    • 2021-02-20
    • 2021-02-24
    • 2020-12-29
    • 2018-12-25
    • 2020-10-26
    • 2018-08-29
    • 1970-01-01
    相关资源
    最近更新 更多