【问题标题】:Dictvectorizer for list as one feature in Python Pandas and Scikit-learnDictvectorizer 用于列表作为 Python Pandas 和 Scikit-learn 中的一项功能
【发布时间】:2015-12-23 20:22:55
【问题描述】:

我这几天一直在尝试解决这个问题,虽然我在这里How can i vectorize list using sklearn DictVectorizer 发现了类似的问题,但解决方案过于简化了。

我想将一些特征拟合到逻辑回归模型中,以预测“中文”或“非中文”。我有一个 raw_name ,我将提取它以获得两个特征 1) 只是姓氏,2) 是姓氏的子字符串列表,例如,'Chan' 将给出 ['ch', 'ha', '一个']。但似乎 Dictvectorizer 没有将列表类型作为字典的一部分。从上面的链接中,我尝试创建一个函数list_to_dict,并成功返回一些dict元素,

{'substring=co': True, 'substring=or': True, 'substring=rn': True, 'substring=ns': True}

但我不知道如何在应用 dictvectorizer 之前将其合并到 my_dict = ... 中。

# coding=utf-8
import pandas as pd
from pandas import DataFrame, Series
import numpy as np
import nltk
import re
import random
from random import randint
import sys
reload(sys)
sys.setdefaultencoding('utf-8')

from sklearn.linear_model import LogisticRegression
from sklearn.feature_extraction import DictVectorizer

lr = LogisticRegression()
dv = DictVectorizer()

# Get csv file into data frame
data = pd.read_csv("V2-1_2000Records_Processed_SEP2015.csv", header=0, encoding="utf-8")
df = DataFrame(data)

# Pandas data frame shuffling
df_shuffled = df.iloc[np.random.permutation(len(df))]
df_shuffled.reset_index(drop=True)

# Assign X and y variables
X = df.raw_name.values
y = df.chineseScan.values

# Feature extraction functions
def feature_full_last_name(nameString):
    try:
        last_name = nameString.rsplit(None, 1)[-1]
        if len(last_name) > 1: # not accept name with only 1 character
            return last_name
        else: return None
    except: return None

def feature_twoLetters(nameString):
    placeHolder = []
    try:
        for i in range(0, len(nameString)):
            x = nameString[i:i+2]
            if len(x) == 2:
                placeHolder.append(x)
        return placeHolder
    except: return []

def list_to_dict(substring_list):
    try:
        substring_dict = {}
        for i in substring_list:
            substring_dict['substring='+str(i)] = True
        return substring_dict
    except: return None

list_example = ['co', 'or', 'rn', 'ns']
print list_to_dict(list_example)

# Transform format of X variables, and spit out a numpy array for all features
my_dict = [{'two-letter-substrings': feature_twoLetters(feature_full_last_name(i)), 
    'last-name': feature_full_last_name(i), 'dummy': 1} for i in X]

print my_dict[3]

输出:

{'substring=co': True, 'substring=or': True, 'substring=rn': True, 'substring=ns': True}
{'dummy': 1, 'two-letter-substrings': [u'co', u'or', u'rn', u'ns'], 'last-name': u'corns'}

样本数据:

Raw_name    chineseScan
Jack Anderson    non-chinese
Po Lee    chinese

【问题讨论】:

    标签: python-2.7 machine-learning scikit-learn vectorization logistic-regression


    【解决方案1】:

    如果我理解正确,您想要一种编码列表值的方法,以便拥有 DictVectorizer 可以使用的特征字典。 (迟了一年但是)可以根据情况使用这样的东西:

    my_dict_list = []
    
    for i in X:
        # create a new feature dictionary
        feat_dict = {}
        # add the features that are straight forward
        feat_dict['last-name'] = feature_full_last_name(i)
        feat_dict['dummy'] = 1
    
        # for the features that have a list of values iterate over the values and
        # create a custom feature for each value
        for two_letters in feature_twoLetters(feature_full_last_name(i)):
            # make sure the naming is unique enough so that no other feature
            # unrelated to this will have the same name/ key
            feat_dict['two-letter-substrings-' + two_letters] = True
    
        # save it to the feature dictionary list that will be used in Dict vectorizer
        my_dict_list.append(feat_dict)
    
    print my_dict_list
    
    from sklearn.feature_extraction import DictVectorizer
    dict_vect = DictVectorizer(sparse=False)
    transformed_x = dict_vect.fit_transform(my_dict_list)
    print transformed_x
    

    输出:

    [{'dummy': 1, u'two-letter-substrings-er': True, 'last-name': u'Anderson', u'two-letter-substrings-on': True, u'two-letter-substrings-de': True, u'two-letter-substrings-An': True, u'two-letter-substrings-rs': True, u'two-letter-substrings-nd': True, u'two-letter-substrings-so': True}, {'dummy': 1, u'two-letter-substrings-ee': True, u'two-letter-substrings-Le': True, 'last-name': u'Lee'}]
    [[ 1.  1.  0.  1.  0.  1.  0.  1.  1.  1.  1.  1.]
     [ 1.  0.  1.  0.  1.  0.  1.  0.  0.  0.  0.  0.]]
    

    如果您不想创建与列表中的值一样多的功能,您可以做的另一件事(但我不推荐)是这样的:

    # sorting the values would be a good idea
    feat_dict[frozenset(feature_twoLetters(feature_full_last_name(i)))] = True
    # or 
    feat_dict[" ".join(feature_twoLetters(feature_full_last_name(i)))] = True
    

    但第一个意味着您不能有任何重复的值,并且可能两者都不能产生好的特性,特别是如果您需要微调和详细的特性。此外,它们减少了两行具有相同组合的两个字母组合的可能性,因此分类可能不会很好。

    输出:

    [{'dummy': 1, 'last-name': u'Anderson', frozenset([u'on', u'rs', u'de', u'nd', u'An', u'so', u'er']): True}, {'dummy': 1, 'last-name': u'Lee', frozenset([u'ee', u'Le']): True}]
    [{'dummy': 1, 'last-name': u'Anderson', u'An nd de er rs so on': True}, {'dummy': 1, u'Le ee': True, 'last-name': u'Lee'}]
    [[ 1.  0.  1.  1.  0.]
     [ 0.  1.  1.  0.  1.]]
    

    【讨论】:

      猜你喜欢
      • 2015-06-15
      • 2015-04-08
      • 2015-02-12
      • 2015-05-09
      • 2014-08-06
      • 2016-02-01
      • 2021-10-21
      • 2014-04-16
      • 2015-08-13
      相关资源
      最近更新 更多