【发布时间】:2018-06-13 22:01:09
【问题描述】:
我是 Python 和机器学习的新手。我的实现基于 IEEE 研究论文 http://ieeexplore.ieee.org/document/7320414/(错误报告、功能请求或简单的表扬?关于自动分类应用评论)
我想将文本分类。文本是来自 google play store 或 apple app store 的用户评论。研究中使用的类别是错误、功能、用户体验、评级。鉴于这种情况,我试图在 python 中使用 sklearn 包实现决策树。我遇到了 sklearn 'IRIS' 提供的示例数据集,它使用映射到目标的特征及其值构建树模型。在本例中,它是数值数据。
我正在尝试对文本而不是数字数据进行分类。例子:
- 我非常喜欢升级到 pdf。但是,它们不再显示了 修复它,它将是完美的 [BUG]
- 我只是希望它会在我低于一定金额时通知我[FEATURE]
- 这个应用程序对我的业务非常有帮助[评分]
- 在 iTunes 中轻松查找和购买歌曲[用户体验]
鉴于这些文本和这些类别的更多用户评论,我想创建一个分类器,可以使用数据进行训练并预测任何给定用户评论的目标。
到目前为止,我已经对文本进行了预处理,并以包含预处理数据及其目标的元组列表的形式创建了训练数据。
我的预处理:
- 将多行 cmets 标记为单个句子
- 将每个句子标记为单词
- 删除分词句中的停用词
- 将分词句中的单词词形化
(['i', 'liked', 'much', 'upgrade', 'pdfs', 'however', 'displaying', 'anymore', 'fix', 'perfect'], "错误")
这是我目前所拥有的:
import json
from sklearn import tree
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from nltk.tokenize import sent_tokenize, RegexpTokenizer
# define a tokenizer to tokenize sentences and also remove punctuation
tokenizer = RegexpTokenizer(r'\w+')
# this list stores all the training data along with it's label
tagged_tokenized_comments_corpus = []
# Method: to add data to training set
# Parameter: Tuple in the format (Data, Label)
def tag_tokenized_comments_corpus(*tuple_data):
tagged_tokenized_comments_corpus.append(tuple_data)
# step 1: Load all the stop words from the nltk package
stop_words = stopwords.words("english")
stop_words.remove('not')
# creating a temporary list to copy the existing stop words
temp_stop_words = stop_words
for word in temp_stop_words:
if "n't" in word:
stop_words.remove(word)
# load the data set
files = ["Bug.txt", "Feature.txt", "Rating.txt", "UserExperience.txt"]
d = {"Bug": 0, "Feature": 1, "Rating": 2, "UserExperience": 3}
for file in files:
input_file = open(file, "r")
file_text = input_file.read()
json_content = json.loads(file_text)
# step 3: Tokenize multi sentence into single sentences from the user comments
comments_corpus = []
for i in range(len(json_content)):
comments = json_content[i]['comment']
if len(sent_tokenize(comments)) > 1:
for comment in sent_tokenize(comments):
comments_corpus.append(comment)
else:
comments_corpus.append(comments)
# step 4: Tokenize each sentence, remove stop words and lemmatize the comments corpus
lemmatizer = WordNetLemmatizer()
tokenized_comments_corpus = []
for i in range(len(comments_corpus)):
words = tokenizer.tokenize(comments_corpus[i])
tokenized_sentence = []
for w in words:
if w not in stop_words:
tokenized_sentence.append(lemmatizer.lemmatize(w.lower()))
if tokenized_sentence:
tokenized_comments_corpus.append(tokenized_sentence)
tag_tokenized_comments_corpus(tokenized_sentence, d[input_file.name.split(".")[0]])
# step 5: Create a dictionary of words from the tokenized comments corpus
unique_words = []
for sentence in tagged_tokenized_comments_corpus:
for word in sentence[0]:
unique_words.append(word)
unique_words = set(unique_words)
dictionary = {}
i = 0
for dict_word in unique_words:
dictionary.update({i, dict_word})
i = i + 1
train_target = []
train_data = []
for sentence in tagged_tokenized_comments_corpus:
train_target.append(sentence[0])
train_data.append(sentence[1])
clf = tree.DecisionTreeClassifier()
clf.fit(train_data, train_target)
test_data = "Beautiful Keep it up.. this far is the most usable app editor..
it makes my photos more beautiful and alive.."
test_words = tokenizer.tokenize(test_data)
test_tokenized_sentence = []
for test_word in test_words:
if test_word not in stop_words:
test_tokenized_sentence.append(lemmatizer.lemmatize(test_word.lower()))
#predict using the classifier
print("predicting the labels: ")
print(clf.predict(test_tokenized_sentence))
但是,这似乎不起作用,因为在我们训练算法时它会在运行时引发错误。我在想如果我可以将元组中的单词映射到字典并将文本转换为数字形式并训练算法。但我不确定这是否可行。
谁能建议我如何修复此代码?或者是否有更好的方法来实现此决策树。
Traceback (most recent call last):
File "C:/Users/venka/Documents/GitHub/RE-18/Test.py", line 87, in <module>
clf.fit(train_data, train_target)
File "C:\Users\venka\Anaconda3\lib\site-packages\sklearn\tree\tree.py", line 790, in fit
X_idx_sorted=X_idx_sorted)
File "C:\Users\venka\Anaconda3\lib\site-packages\sklearn\tree\tree.py", line 116, in fit
X = check_array(X, dtype=DTYPE, accept_sparse="csc")
File "C:\Users\venka\Anaconda3\lib\site-packages\sklearn\utils\validation.py", line 441, in check_array
"if it contains a single sample.".format(array))
ValueError: Expected 2D array, got 1D array instead:
array=[ 0. 0. 0. ..., 3. 3. 3.].
Reshape your data either using array.reshape(-1, 1) if your data has a
single feature or array.reshape(1, -1) if it contains a single sample.
【问题讨论】:
-
你有发生错误的回溯吗?在不知道实际出了什么问题的情况下,我们无法完全帮助您。也就是说,我假设您可能很难将句子直接插入 sklearn 决策树。
-
我已经用 Traceback 更新了帖子。我知道将句子直接插入决策树是非常困难的。我正在考虑创建一个字典,从语料库中获取所有唯一单词并为每个单词分配一个唯一的数值,以便我们可以将其作为数字数据元组而不是句子传递,但不确定这是否可行
-
你试过了吗?为什么你认为它行不通?这里的错误是不言自明的,您需要提交一个二维数组。是错误的数据还是目标?
-
我还没有尝试过。在我在文章中提到的 IRIS 数据集示例中,训练数据由每个数据的 4 个常量特征组成,我假设这 4 个值使算法易于理解和构建模型。但是,当我将文本转换为数字数据时,每个句子的长度都不相同。我同意这个问题即使在没有转换成数字数据的情况下也存在。关于错误,我正在再次查看堆栈跟踪。在你提到之后,现在更有意义了,我会再试一次。
-
您可以在 sklearn 中使用 CountVectorizer 或 TfidfVectorizer 将过滤后的单词转换为可供 ML 算法使用的数值数据。
标签: python machine-learning classification decision-tree sklearn-pandas