【问题标题】:JSON escape double quotesJSON转义双引号
【发布时间】:2017-01-14 22:21:32
【问题描述】:

我知道这个标题在这里似乎很受欢迎,但是快速浏览它们通常会涉及到提问者有一个单独的 JSON 部分的情况。

在某些情况下," 用于表示英寸,或者它包装了一个短语来表示某种昵称,无论它出现在已经用双引号括起来的 JS 对象的值字符串中。

这是一个我遇到问题的 JS 对象字符串的示例(我使用正则表达式来双引号键并删除额外的空格,但这是所有荣耀中的刮擦字符串):

'{\n\t\t\n\t\t\t\t\t\n\t\n\n\t\n\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"16241885":{title: "Nosefrida Fridababy Windi Gas & Colic Relief", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
\n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"8650356":{title: "Babyganics Face- Hand & Baby Wipes- Fragrance Free- 100 Count", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"16249889":{title: "Nosefrida Nasal Aspirator Replacement Filters", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"8650355":{title: "Babyganics Face- Hand & Baby Wipes- Fragrance Free- 40 Count", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"15490928":{title: "BabyGanics Newborn Ultra Absorbent Jumbo Size Diapers - 36 Count", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"14712536":{title: "Marvel Superhero Bandages", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"16263505":{title: "Nosefrida "The Snotsucker" Nasal Aspirator", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t\t\t, \n\t\t\n\t\n\n\t\t\n\n\t\t \n\n\n\n 
"14848093":{title: "Zarbee\'s Children\'s Cough Syrup - Grape", isIneligible: false, isDiscontinued: false, isLowInventory: false, isAllowed: true}
    \n\n\t\n\n\t\t\n\t \n\n\t\t\n\t}'

我已经尝试过,json.dumps 首先在字符串上,但这只是双重转义,需要双重json.loads,这让我回到了第一方。我试过这样的正则表达式:

double_quotes_in_json = re.compile(r'(?<=:)(\s*"[^"]*)(")([^"]*)(")?(?=[^"]*",|"\s*\})')


def escape_double_quotes(jsn_string, pattern=double_quotes_in_json):
    for match in pattern.finditer(jsn_string):
        # current pattern only matches 1 instance of either one double quote in JSON value string
        # (presumably signifying inches) or 1 instance of phrase wrapped in double quotes
        # for something like nicknames
        # matches will have either 3 or 4 groups, representing one of the 2 match types described above
        groups_matched = len(match.groups())
        entire_match = match.group()
        if groups_matched == 3:
            # we only matched one double quote
            subbed_match = pattern.sub('$1\\$2$3', entire_match)
            jsn_string = re.sub(entire_match, subbed_match, jsn_string)
        elif groups_matched == 4:
            # we matched a phrase wrapped in double quotes
            subbed_match = pattern.sub('$1\\$2$3\\$4', entire_match)
            jsn_string = re.sub(entire_match, subbed_match, jsn_string)
    return jsn_string

虽然这似乎是最有希望的,但它似乎重新插入了双引号,而没有我在子中的转义字符,同时也没有在第一组中重新插入。(我尝试过使用和不使用原始字符串在子函数r) 所以对于上面的问题部分(下面是一个子字符串):

 "16263505":{title: "Nosefrida "The Snotsucker" Nasal Aspirator"

该模式不会将 1 重新分组,并且出于某种原因在单引号中进行分组(以下是失败的正则表达式处理的子字符串):

"16263505":{title: "The Snotsucker"' Nasal Aspirator"

无论如何json.loads 抱怨未转义的"。

编辑 1: 我的正则表达式可以提取未转义的引号,但将其重新插入并没有按预期进行,我可能在这里做了一些愚蠢的事情并且可以使用新的眼睛。

我的函数带有打印语句的示例输出:

low_inventory = response.xpath(
                '//script[contains(., "islistEligibility") or contains(., "ishlistEligibility")]/text()'
                ).re_first(r'(?s)(?<=registryWislistEligibilityMap)(?:\s*=\s*)(\{.+\})')

In [453]: for m in double_quotes_in_json.finditer(low_inventory):
     ...:     groups_matched = len(m.groups())
     ...:     print('groups: ', m.groups())
     ...:     entire_match = m.group()
     ...:     print('entire match: ', m.group())
     ...:     if groups_matched == 3:
     ...:             # we only matched a single double quote
     ...:             subbed_match = double_quotes_in_json.sub(r'$1\\$2$3', entire_match)
     ...:             print('subbed3: ', subbed_match)
     ...:             jsn_string = re.sub(entire_match, subbed_match, jsn_string)
     ...:     elif groups_matched == 4:
     ...:             subbed_match = double_quotes_in_json.sub(r'$1\\$2$3\\\$4', entire_match)
     ...:             print('subbed4: ', subbed_match)
     ...:             jsn_string = re.sub(entire_match, subbed_match, jsn_string)
     ...: print(jsn_string)
     ...: 
groups:  (' "Nosefrida ', '"', 'The Snotsucker', '"')
entire match:   "Nosefrida "The Snotsucker"
subbed4:   "Nosefrida "The Snotsucker"
{  "16241885":{"title": "Nosefrida Fridababy Windi Gas &amp; Colic Relief", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "8650356":{"title": "Babyganics Face- Hand &amp; Baby Wipes- Fragrance Free- 100 Count", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "16249889":{"title": "Nosefrida Nasal Aspirator Replacement Filters", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "8650355":{"title": "Babyganics Face- Hand &amp; Baby Wipes- Fragrance Free- 40 Count", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "15490928":{"title": "BabyGanics Newborn Ultra Absorbent Jumbo Size Diapers - 36 Count", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "14712536":{"title": "Marvel Superhero Bandages", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "16263505":{"title": "The Snotsucker"' Nasal Aspirator", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true},   "14848093":{"title": "Zarbee's Children's Cough Syrup - Grape", "isIneligible": false, "isDiscontinued": false, "isLowInventory": false, "isAllowed": true} }

【问题讨论】:

  • 这既不是有效的 JSON 也不是 Javascript。它在我能想到的任何语言中都是无效的。您正在尝试解析垃圾。这些垃圾是从哪里来的?能不能从源头上解决?
  • 当有人试图执行该脚本时,它应该会因语法错误而崩溃。这肯定是原始数据吗?有没有可能在抓取过程中删除一些\?
  • 您在开始使用您讨论的正则表达式处理数据之前或之后呈现的数据?
  • @jwpfox 示例数据在任何操作之前
  • 如果你知道数据有一个定义的结构,你应该相应地解析它,例如title: "(.+?)" isIneligible: …。再次,字符串定界符是垃圾,显然不能依赖于解析。如果没有这样的东西你可以依赖......好吧,就任何理智的解析器而言,这个字符串没有正确的答案。

标签: python json regex escaping


【解决方案1】:

由于某种原因,使用 pythons 内置替换功能达到了预期的结果,而 re.sub 没有正确转义双引号。 (这是在带有单转义的原始字符串或带有双转义的常规字符串中使用组引用)。无论哪种方式,这是工作功能。如果有人对为什么使用替换比 re.sub 有效,我会很感兴趣。

(旧代码注释掉)

double_quotes_in_json = re.compile(r'(?<=:)(\s*")([^"]*)(")([^"]*)(")?(?=[^"]*",|"\s*\})')


def escape_double_quotes(jsn_string, pattern=double_quotes_in_json):
    for match in pattern.finditer(jsn_string):
        # current pattern only matches 1 instance of either one double quote in JSON value string
        # (presumably signifying inches) or 1 instance of phrase wrapped in double quotes
        # for something like nicknames
        # matches will have either 3 or 4 groups, representing one of the 2 match types described above
        num_groups_matched = len(match.groups())
        groups = match.groups()
        entire_match = match.group()
        print('groups: ', match.groups())
        print('entire: ', entire_match)
        if num_groups_matched == 4:
            # we only matched one double quote
            # subbed_match = pattern.sub('$1$2\\$3$4', entire_match)
            # jsn_string = re.sub(entire_match, subbed_match, jsn_string)
            target = ''.join(groups[1:4])
            replaced = target.replace('"', '\\"')
            print(replaced)
            jsn_string = jsn_string.replace(target, replaced)
        elif num_groups_matched == 5:
            # we matched a phrase wrapped in double quotes
            # subbed_match = pattern.sub('$1$2\\$3$4\\$5', entire_match)
            # jsn_string = re.sub(entire_match, subbed_match, jsn_string)
            target = ''.join(groups[1:])
            replaced = target.replace('"', '\\"')
            print(replaced)
            jsn_string = jsn_string.replace(target, replaced)
    return jsn_string

编辑#1(又名:经过一些睡眠方法):

double_quotes_in_title_attr = re.compile(
    r'(?<="title":)(?:\s*")(?P<value>.+?)(?=",\s*"\w+":|"\s*\})'
)


def escape_double_quotes_in_title(jsn_string, pattern=double_quotes_in_title_attr):
    for match in pattern.finditer(jsn_string):
        target = match.group('value')
        replaced = target.replace('"', '\\"')
        jsn_string = jsn_string.replace(target, replaced)
    return jsn_string

# use this first to properly quote keys so the above pattern will match
unquoted_key_pattern = re.compile(r'(?!")(\'?(?P<key>\w+)\'?)(?=:\s*(?:"|false|true|\d|\[|\{))')

def fix_json_keys(jsn, pattern=unquoted_key_pattern):
    return pattern.sub(r'"\g<key>"', jsn)

感谢@deceze 的帮助。

【讨论】:

    猜你喜欢
    • 2020-11-26
    • 1970-01-01
    • 2013-06-17
    • 2013-03-16
    • 1970-01-01
    • 1970-01-01
    • 2013-08-09
    • 1970-01-01
    相关资源
    最近更新 更多