【问题标题】:How to tokenize text using punctuation as boundaries (Python)如何使用标点符号作为边界标记文本(Python)
【发布时间】:2017-09-15 09:09:17
【问题描述】:

我正在使用来自sklearn 的CountVectorizer 进行文本标记化(2-gram)并创建一个术语文档矩阵。如何将文本标记为以标点符号为边界的 2-gram?例如,输入句子是“this is example, with punctuation”。 我希望标记是“this is”、“is example”、“with punctuation”。 我不想用逗号隔开的“example with”。

以下是我当前的代码:

from sklearn.feature_extraction.text import CountVectorizer
df = pd.DataFrame({'title':['this is example, with punctuation'], 'page':[1]})
countvec = CountVectorizer(ngram_range=(2, 2), analyzer="word")

test_tdm = pd.DataFrame(countvec.fit_transform(df.title).toarray(), columns=countvec.get_feature_names())
print(test_tdm)

谢谢!

【问题讨论】:

    标签: python tokenize term-document-matrix


    【解决方案1】:

    一种方法是首先通过标点符号分割您想要标记的字符串。像这样的:

    import re, string
    
    patt = '[' + string.punctuation + ']'
    splitted_title = re.split(patt, df.title)
    

    然后将标记化应用于splitted_title 的每个元素

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2014-07-13
      • 2021-09-03
      • 1970-01-01
      • 1970-01-01
      • 2013-11-19
      • 2017-09-21
      • 1970-01-01
      相关资源
      最近更新 更多