【发布时间】:2018-12-07 08:45:21
【问题描述】:
我有以下代码:
import pandas as pd
import numpy as np
from sklearn.cross_validation import train_test_split
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.svm import LinearSVC
from sklearn.metrics import accuracy_score
from sklearn.pipeline import Pipeline, FeatureUnion
from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.grid_search import GridSearchCV
# Load ANSI file into pandas dataframe.
df = pd.read_csv(r'c:/papf.txt', encoding = 'latin1', usecols=['LAST_NAME', 'RACE'])
# Convert last name to lower case.
df['LAST_NAME'] = df['LAST_NAME'].str.lower()
# Remove the last name spaces.
df['LAST_NAME'] = df['LAST_NAME'].str.replace(' ', '')
# Remove all rows where race is NOT in African, White, Asian.
df = df.drop(df[~df['RACE'].isin(['African', 'White', 'Asian'])].index)
class AverageWordLengthExtractor(BaseEstimator, TransformerMixin):
"""Takes in dataframe, extracts last name column, outputs average word length"""
def __init__(self):
pass
def average_word_length(self, name):
"""Helper code to compute average word length of a name"""
return np.mean([len(word) for word in name.split()])
def transform(self, df, y=None):
"""The workhorse of this feature extractor"""
return df['LAST_NAME'].apply(self.average_word_length)
def fit(self, df, y=None):
"""Returns self unless something different happens in train and test"""
return self
# Split into train and test sets with 20% used for testing.
data_train, data_test, y_train, y_true = \
train_test_split(df['LAST_NAME'], df['RACE'], test_size=0.2)
# Build the pipeline.
ngram_count_pipeline = Pipeline([
('ngram', CountVectorizer(ngram_range=(1, 4), analyzer='char'))
])
pipeline = Pipeline([
('feats', FeatureUnion([
('ngram', ngram_count_pipeline), # can pass in either a pipeline
#('ngram', CountVectorizer(ngram_range=(1, 4), analyzer='char')),
('ave', AverageWordLengthExtractor()) # or a transformer
])),
('clf', LinearSVC()) # classifier
])
# Train the classifier.
classifier = LinearSVC()
model = pipeline.fit(data_train)
# Test the classifier.
y_test = model.predict(data_test)
# Print the accuracy percentage.
print(accuracy_score(y_true, y_test))
#one = ngram_counter.transform('chapman')
#print(model.predict(one))
我想出了this code based on this excellent blog post by Michelle Fullwood。
但是博文没有详细说明以下部分:
注意FeatureUnion 中的第一项是ngram_count_pipeline。这只是一个由列提取转换器创建的Pipeline 和CountVectorizer(列提取器是必需的,因为我们正在对 Pandas 数据帧进行操作,而不是直接通过管道发送道路名称列表)。
所以我的问题是如何将 n-gram CountVectorizer 添加为管道以及如何执行列提取器部分?
另外,我将如何使用该模型来预测姓氏 Chapman?
获得每个输出类的准确性和概率也很棒。
我的输入数据基本上是带有种族输出的姓氏。
我还收到以下不知道如何解决的警告:
C:\ProgramData\Anaconda3\lib\site-packages\sklearn\cross_validation.py:41: DeprecationWarning: This module was deprecated in version 0.18 in favor of the model_selection module into which all the refactored classes and functions are moved. Also note that the interface of the new CV iterators are different from that of this module. This module will be removed in 0.20.
"This module will be removed in 0.20.", DeprecationWarning)
C:\ProgramData\Anaconda3\lib\site-packages\sklearn\grid_search.py:42: DeprecationWarning: This module was deprecated in version 0.18 in favor of the model_selection module into which all the refactored classes and functions are moved. This module will be removed in 0.20. DeprecationWarning)
我已升级到最新的 Anaconda(Python 3.6.5 | 由 conda-forge 打包 | (默认,2018 年 4 月 6 日,16:13:16)[MSC v.1900 32 位(英特尔)] 在 win32 上)但是没有解决警告。
CSV 数据示例:
LAST_NAME,RACE
Ramaepadi,African
Motsamai,African
Van Rooyen,White
Khan,Asian
Du Plessis,White
Singh,Asian
Madlanga,African
Janse van Rensburg,
【问题讨论】:
-
你可能想检查this solution...
-
当然给出了一些提示,但仍然令人费解。
-
你能发布一个小的可重复的样本数据集吗?
-
添加示例数据
标签: python pandas machine-learning scikit-learn