【发布时间】:2018-05-15 21:57:55
【问题描述】:
我正在使用 spaCy 和 python 尝试为 sklearn 清理一些文本。我运行循环:
for text in df.text_all:
text = str(text)
text = nlp(text)
cleaned = [token.lemma_ for token in text if token.is_punct==False and token.is_stop==False]
cleaned_text.append(' '.join(cleaned))
它工作得很好,但它留在了<br /><br /> 中的一些文本中。我认为这会被token.is_punct==False 过滤器去掉,但没有。我寻找类似html标签的东西,但找不到任何东西。有谁知道我能做什么?
【问题讨论】:
-
你总是可以在 python 之外预处理数据集,就像使用下面的命令 cat FILE_NAME | sed -r 's/\
\
//g' > NEW_FILE_NAME
标签: python scikit-learn nlp spacy