【问题标题】:RegEx: A way to handle both English and non-English characters (and my solution)RegEx:一种处理英文和非英文字符的方法(以及我的解决方案)
【发布时间】:2016-03-27 03:36:49
【问题描述】:

我想知道是否有推荐的 RegEx 模式来匹配英文和非英文字符。到目前为止,我已经根据answer provided at SO 提出了[^\x00-\x7F]+|[a-zA-Z'-]*。我的解决方案似乎有效,但由于我对 RegEx 非常友好,我想请您检查此令牌并提出一些改进建议。我知道大多数涉及此主题的解决方案,例如this,但我认为目前还没有一个好的正则表达式。

【问题讨论】:

    标签: regex autohotkey


    【解决方案1】:

    答案主要取决于语言。但一般来说,您必须启用“unicode 标志”(这通常通过在您的正则表达式前添加 (?u) 或通过附加 /u 来完成)并使用 unicode 字符串。这样\w、\s等就可以正确匹配对应的unicode字符了。

    Python 2 中的示例(Python 3 默认使用 unicode):

    >>> re.match('\w', 'è')  # byte string, no unicode flag: no match
    >>> re.match('(?u)\w', u'è')  # unicode string and unicode flag: match
    <_sre.SRE_Match object at 0x7f258bac07e8>
    >>> re.match('\w', u'è', re.UNICODE)  # another way to enable the unicode flag
    <_sre.SRE_Match object at 0x7f258bac0850>
    

    【讨论】:

    • 如何在regex101.com和AutoHotKey中使用?
    • @menteith:我不熟悉 regex101,也不知道 AutoHotKey 是什么,抱歉!尝试使用谷歌搜索“AutoHotKey unicode regex”,顺便更新您的问题,添加autohotkey 标签并明确说明您的问题是关于 AutoHotKey 的(否则您的问题可能会被关闭为题外话)
    猜你喜欢
    • 1970-01-01
    • 2012-08-19
    • 1970-01-01
    • 2020-09-26
    • 2021-06-02
    • 1970-01-01
    • 2016-12-27
    • 2016-12-13
    • 2017-09-04
    相关资源
    最近更新 更多