【问题标题】:spaCy and text cleaning, getting rid of '<br /><br />'spaCy和文本清理,摆脱'<br /><br />'
【发布时间】:2018-05-15 21:57:55
【问题描述】:

我正在使用 spaCy 和 python 尝试为 sklearn 清理一些文本。我运行循环:

for text in df.text_all:
    text = str(text)
    text = nlp(text)
    cleaned = [token.lemma_ for token in text if token.is_punct==False and token.is_stop==False]
    cleaned_text.append(' '.join(cleaned))

它工作得很好,但它留在了&lt;br /&gt;&lt;br /&gt; 中的一些文本中。我认为这会被token.is_punct==False 过滤器去掉,但没有。我寻找类似html标签的东西,但找不到任何东西。有谁知道我能做什么?

【问题讨论】:

  • 你总是可以在 python 之外预处理数据集,就像使用下面的命令 cat FILE_NAME | sed -r 's/\
    \
    //g' > NEW_FILE_NAME

标签: python scikit-learn nlp spacy


【解决方案1】:

你可以使用正则表达式:

import re

# ...
cleaned = [token.lemma_...

clean_regex = re.compile('<.*?>')
cleantext = re.sub(clean_regex, '', ' '.join(cleaned))

cleaned_text.append(cleantext)

注意:如果您的文本包含任何“<br /> 标记除外),此方法将不起作用

希望这会有所帮助!

【讨论】:

    猜你喜欢
    • 2015-02-17
    • 2014-01-24
    • 2010-10-15
    • 2012-12-21
    • 1970-01-01
    • 2020-07-16
    • 2010-12-29
    • 1970-01-01
    相关资源
    最近更新 更多