【问题标题】:Extract dict from string从字符串中提取字典
【发布时间】:2020-12-30 03:21:35
【问题描述】:

我正在调用一个函数,该函数返回一个包含字典的字符串。我如何提取这个字典,记住第一行和最后一行可以包含'{'和'}'。

This is a {testing string} example
This {is} a testing {string} example
{"website": "stackoverflow",
"type": "question",
"date": "10-09-2020"
}
This is a {testing string} example
This {is} a testing {string} example

我需要将此值提取为 dict 变量。

{"website": "stackoverflow",
"type": "question",
"date": "10-09-2020"
}

【问题讨论】:

  • 是否应该更新函数以返回正确的字典? (本质上是从根本上解决数据问题,而不是编写额外的代码来处理它。)
  • 不幸的是,我无法控制函数的输出。我正在使用 subprocess 调用 shell 命令,并且我希望输出始终采用该格式。
  • dict 是否保证始终位于字符串中的同一位置?
  • 所以我需要提取里面的字典,这样我就可以创建一个正确的输出,即删除某些键/值
  • 由于您的输入可以包含带有 { 和 } 字符的内容,因此您必须找到每一对并查看它是否包含有效的字典(并希望不会发生任何事情看起来像一本有效的字典)。

标签: python string dictionary


【解决方案1】:

更新答案


使用来自@martineau 和@ekhumoro 的cmet,以下编辑后的代码包含一个搜索字符串并提取所有有效dicts 的函数。这是对我之前回答的一种更稳健的方法,因为现实世界 dict 的内容可能会有所不同,而这个逻辑(希望)可以解释这一点。

示例代码:

import json
import re

def extract_dict(s) -> list:
    """Extract all valid dicts from a string.
    
    Args:
        s (str): A string possibly containing dicts.
    
    Returns:
        A list containing all valid dicts.
    
    """
    results = []
    s_ = ' '.join(s.split('\n')).strip()
    exp = re.compile(r'(\{.*?\})')
    for i in exp.findall(s_):
        try:
            results.append(json.loads(i))        
        except json.JSONDecodeError:
            pass    
    return results

测试字符串:

OP 的原始字符串已更新为添加多个dicts、一个数值作为最后一个字段和一个list 值。

s = """
This is a {testing string} example
This {is} a testing {string} example
{"website": "stackoverflow",
"type": "question",
"date": 5
}
{"website": "stackoverflow",
"type": "question",
"date": "2020-09-11"
}
{"website": "stackoverflow",
"type": "question",
"dates": ["2020-09-11", "2020-09-12"]
}
This is a {testing string} example
This {is} a testing {string} example
"""

输出:

正如 OP 所述,字符串中通常只有一个 dict,因此(显然)可以使用 results[0] 访问。

>>> results = extract_dict(s)

[{'website': 'stackoverflow', 'type': 'question', 'date': 5},
 {'website': 'stackoverflow', 'type': 'question', 'date': '2020-09-11'},
 {'website': 'stackoverflow', 'type': 'question', 'dates': ['2020-09-11', '2020-09-12']}]

原答案:


忽略此部分。虽然代码有效,但它特别适合 OP 的要求,并且不适合其他用途。

此示例使用正则表达式来识别 dict start {" 和 dict end "} 并提取中间,然后将字符串转换为正确的 dict。随着新行的出现和正则表达式的复杂化,我只是将字符串展平以开始。

根据@jizhihaoSAMA 的评论,我已更新为使用json.loads 将字符串转换为dict,因为它更简洁。如果您不想额外导入,eval 也可以,但不推荐。

示例代码:

import json
import re

s = """
This is a {testing string} example
This {is} a testing {string} example
{"website": "stackoverflow",
"type": "question",
"date": "10-09-2020"
}
This is a {testing string} example
This {is} a testing {string} example
"""

s_ = ' '.join(s.split('\n')).strip()
d = json.loads(re.findall(r'(\{\".*\"\s?\})', s_)[0])

>>> d
>>> d['website']

输出:

{"website": "stackoverflow", "type": "question", "date": "10-09-2020"}

'stackoverflow'

【讨论】:

  • eval() 不推荐,尝试使用更安全的函数,如ast.literal_eval() 或json.loads。
  • OP 的真实输入不太可能像测试示例那样简单,因此这种方法在实践中发挥作用的可能性很小。
  • @ekhumoro - 虽然这在现实世界中是一个公平的假设,但我觉得这个评论是不公平的 - 因为这种观点几乎可以用于 SO 的所有答案,并使它们无效。该解决方案符合 OP 的要求并提供了示例。超出范围的任何内容都超出了范围。
  • 这根本不是真的。通常可以从测试用例中进行概括,以考虑大多数现实世界的可能性。如果您的解决方案可以通过对输入的微小更改而失效,那么它实际上并没有以一种有用的方式满足要求。 (例如,如果 dict 中的最后一个值恰好是数字而不是字符串,您的解决方案将中断)。
  • @ekhumoro - 点得好,答案已更新。感谢您的扩展思路。
猜你喜欢
  • 2023-01-13
  • 2017-02-09
  • 2016-06-24
  • 2022-06-10
  • 2017-12-15
  • 1970-01-01
  • 2021-09-28
  • 1970-01-01
相关资源
最近更新 更多