【问题标题】:Remove punctuation from Unicode formatted strings从 Unicode 格式的字符串中删除标点符号
【发布时间】:2017-08-21 05:55:59
【问题描述】:

我有一个从字符串列表中删除标点符号的函数:

def strip_punctuation(input):
    x = 0
    for word in input:
        input[x] = re.sub(r'[^A-Za-z0-9 ]', "", input[x])
        x += 1
    return input

我最近修改了我的脚本以使用 Unicode 字符串,这样我就可以处理其他非西方字符。这个函数在遇到这些特殊字符时会中断,只返回空的 Unicode 字符串。如何可靠地从 Unicode 格式的字符串中删除标点符号?

【问题讨论】:

  • strip_punctuation() 应该接受字符串而不是字符串列表,然后如果您需要它可以list_of_strings = map(strip_punctuation, list_of_strings)
  • 这实际上可能是一个更好的方法。我喜欢你和 F.C. 使用 unicode 类别的实现。

标签: python unicode


【解决方案1】:

你可以使用unicode.translate()方法:

import unicodedata
import sys

tbl = dict.fromkeys(i for i in xrange(sys.maxunicode)
                      if unicodedata.category(unichr(i)).startswith('P'))
def remove_punctuation(text):
    return text.translate(tbl)

你也可以使用regex module支持的r'\p{P}':

import regex as re

def remove_punctuation(text):
    return re.sub(ur"\p{P}+", "", text)

【讨论】:

  • +1 用于建议正则表达式 - 这是 的方式。值得注意的是,它是非标准的(尚未),必须单独安装。此外,在 py2 中,您需要模式为 unicode (ur"..") 才能切换 unicode 匹配模式。
  • @thg435:我已经添加了 regex 模块的链接并使模式 unicode
  • @acpigeon:我已将tbl 移至全局范围,以明确它只需要生成一次
  • re 模块(不是regex)似乎不支持\p{P},是吗?
  • @posdef 它是 Python 2 代码(阅读第一条评论)。在 Python 3 上将 u'' 前缀放在 r'' 之前或使用 u"\\p{P}+" (在这种情况下,您必须手动转义)。
【解决方案2】:

如果你想在 Python 3 中使用 J.F. Sebastian 的解决方案:

import unicodedata
import sys

tbl = dict.fromkeys(i for i in range(sys.maxunicode)
                      if unicodedata.category(chr(i)).startswith('P'))
def remove_punctuation(text):
    return text.translate(tbl)

【讨论】:

    【解决方案3】:

    您可以使用unicodedata 模块的category 函数遍历字符串以确定字符是否为标点符号。

    有关category 的可能输出,请参阅General Category Values 上的unicode.org 文档

    import unicodedata.category as cat
    def strip_punctuation(word):
        return "".join(char for char in word if cat(char).startswith('P'))
    filtered = [strip_punctuation(word) for word in input]
    

    此外,请确保您正确处理编码和类型。此演示文稿是一个很好的起点:http://bit.ly/unipain

    【讨论】:

    • +1 用于 unipain 链接。我正在尝试实现这一点,但在 result[i] 行上出现“IndexError:list assignment index out of range”。我会继续胡闹的。
    • @acpigeon:出于某种原因,我认为您可以在不预先填充的情况下以稀疏的方式分配给列表。用更好的方法编辑。
    • 这个答案中有一个小但重要的错误:strip_punctuation 实际上与您的意图相反,并且会返回 only 标点符号,因为您忘记了 not你的理解。我会编辑答案来修复它,除了“编辑必须至少有 6 个字符。”
    【解决方案4】:

    基于Daenyth answer 的更短版本

    import unicodedata
    
    def strip_punctuation(text):
        """
        >>> strip_punctuation(u'something')
        u'something'
    
        >>> strip_punctuation(u'something.,:else really')
        u'somethingelse really'
        """
        punctutation_cats = set(['Pc', 'Pd', 'Ps', 'Pe', 'Pi', 'Pf', 'Po'])
        return ''.join(x for x in text
                       if unicodedata.category(x) not in punctutation_cats)
    
    input_data = [u'somehting', u'something, else', u'nothing.']
    without_punctuation = map(strip_punctuation, input_data)
    

    【讨论】:

    • OP 说input_data 是一个字符串列表,而不仅仅是一个字符串。 (当然,你可以只映射你的版本)
    猜你喜欢
    • 2015-07-07
    • 2014-08-29
    • 1970-01-01
    • 1970-01-01
    • 2013-10-08
    • 2016-02-20
    • 1970-01-01
    相关资源
    最近更新 更多