【问题标题】:Regular expression that accepts tokens of three or more alphabetical characters接受三个或更多字母字符的标记的正则表达式
【发布时间】:2022-07-23 00:34:35
【问题描述】:

我正在尝试构建一个 TFIDVectorizer,它只接受使用 TFIdfVectorizer(token_pattern="(?u)\\b\\D\\D\\D+\\b")

的 3 个或更多字母字符的标记

但它的行为不正确,我知道 token_pattern="(?u)\\b\\w\\w\\w+\\b" 接受 3 个或更多 字母数字 字符的标记,所以我只是不明白为什么前者不起作用。

我错过了什么?

【问题讨论】:

  • 三个或更多字母是token_pattern="[^\W\d_]{3,}" 或token_pattern="[a-zA-Z]{3,}"

标签: python regex tfidfvectorizer


【解决方案1】:

问题在于使用\D 元字符,因为它实际上是为了匹配任何非数字 字符,而不是任何字母 字符。来自Python docs:


你可以改为:
token_pattern="(?i)[a-z]{3,}"

解释:

  • (?i) — 使匹配不区分大小写的内联标志,
  • [a-z] — 匹配任何拉丁字母,
  • {3,} — 使前一个 token 匹配三次或更多次(贪婪,即尽可能多地匹配)。

我希望这能回答您的问题。 :)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-03-10
    • 1970-01-01
    • 1970-01-01
    • 2011-01-12
    • 1970-01-01
    相关资源
    最近更新 更多