【问题标题】:Handling special characters (decoding) in a Popen generator在 Popen 生成器中处理特殊字符(解码)
【发布时间】:2019-11-20 06:37:32
【问题描述】:

上下文

我有一个生成器,它不断输出特定命令的每一行(参见下面的代码 sn-p,代码取自 here)。

def execute(cmd):
    popen = subprocess.Popen(cmd, stdout=subprocess.PIPE, shell=True, universal_newlines=True)

    for stdoutLine in iter(popen.stdout.readline, ""):
        yield stdoutLine.rstrip('\r|\n')

问题

问题是,stdout 行可能包含 cp1252 无法处理的特殊字符。 (请参阅下面的多条错误消息,每条都来自不同的测试)

UnicodeDecodeError: 'charmap' codec can't decode byte 0x8d in position 6210: character maps to <undefined>
UnicodeDecodeError: 'charmap' codec can't decode byte 0x8d in position 3691: character maps to <undefined>
UnicodeDecodeError: 'charmap' codec can't decode byte 0x8d in position 6228: character maps to <undefined>

问题

如何处理这些特殊字符?

【问题讨论】:

    标签: python-3.x encoding special-characters decode popen


    【解决方案1】:

    解决方案很简单:如果没有必要,不要解码标准输出。

    我的解决方案是在执行函数中添加一个参数,该参数决定生成器是生成解码后的字符串还是未触及的字节。

    def execute(cmd, decode=False):
        popen = subprocess.Popen(cmd, stdout=subprocess.PIPE, shell=True, universal_newlines=decode)
    
        for stdoutLine in iter(popen.stdout.readline, ""):
            if decode:
                yield stdoutLine.rstrip('\r|\n')
            else:
                yield stdoutLine.rstrip(b'\r|\n')
    

    因此,当我知道我正在执行的命令将返回 ASCII 字符并且需要解码后的字符串时,我会传递 decode=True 参数。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-09-13
      • 2011-11-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多