【问题标题】:Exporting and Loading nested Pydantic models导出和加载嵌套的 Pydantic 模型
【发布时间】:2021-11-13 17:23:20
【问题描述】:

我有一个带有嵌套数据结构的简单 pydantic 模型。 我希望能够简单地将这个模型的实例保存和加载为 .json 文件。

所有模型都继承自 Base 类,配置简单。

class Base(pydantic.BaseModel):
    class Config:
        extra = 'forbid'   # forbid use of extra kwargs

有一些带有继承的简单数据模型

class Thing(Base):
    thing_id: int

class SubThing(Thing):
    name: str

还有一个Container 类,其中包含一个Thing

class Container(Base):
    thing: Thing

我可以创建一个Container 实例并将其保存为.json

# make instance of container
c = Container(
    thing = SubThing(
        thing_id=1,
        name='my_thing')
)

json_string = c.json(indent=2)
print(json_string)

"""
{
  "thing": {
    "thing_id": 1,
    "name": "my_thing"
  }
}
"""

但 json 字符串未指定 thing 字段是使用 SubThing 构造的。因此,当我尝试将此字符串加载到新的 Container 实例中时,我收到错误消息。

print(c)
"""
Traceback (most recent call last):
  File "...", line 36, in <module>
    c = Container.parse_raw(json_string)
  File "pydantic/main.py", line 601, in pydantic.main.BaseModel.parse_raw
  File "pydantic/main.py", line 578, in pydantic.main.BaseModel.parse_obj
  File "pydantic/main.py", line 406, in pydantic.main.BaseModel.__init__
pydantic.error_wrappers.ValidationError: 1 validation error for Container
thing -> name
  extra fields not permitted (type=value_error.extra)
"""

有没有一种简单的方法来保存Container 实例,同时保留有关thing 类类型的信息,以便我可以可靠地重建初始Container 实例?如果可能的话,我想避免腌制物体。

一种可能的解决方案是手动序列化,例如使用


def serialize(attr_name, attr_value, dictionary=None):
    if dictionary is None:
        dictionary = {}
    if not isinstance(attr_value, pydantic.BaseModel):
        dictionary[attr_name] = attr_value
    else:
        sub_dictionary = {}
        for (sub_name, sub_value) in attr_value:
            serialize(sub_name, sub_value, dictionary=sub_dictionary)
        dictionary[attr_name] = {type(attr_value).__name__: sub_dictionary}
    return dictionary


c1 = Container(
    container_name='my_container',
    thing=SubThing(
        thing_id=1,
        name='my_thing')
)

from pprint import pprint as print
print(serialize('Container', c1))

{'Container': {'Container': {'container_name': 'my_container',
                             'thing': {'SubThing': {'name': 'my_thing',
                                                    'thing_id': 1}}}}}

但这消除了利用包进行序列化的大部分好处。

【问题讨论】:

  • 你为什么要使用pydantic——就像你从它提供的验证中受益一样?只是好奇
  • 是的,我主要将它用于验证,但原则上我可以使用其他东西。这是我实际应用程序的一个极其简化的版本。
  • 只在网上粗略看了一下,看起来这是一个已知问题,pydantic 不支持将嵌套的 json 加载到模型类中,但有计划在未来提供支持用例。实际上,我很惊讶 pydantic 没有将 dict 解析为嵌套模型 - 对我来说似乎是一个足够常见的用例。
  • 使用here 提到的根验证器也可能有效
  • 我已经使用数据类测试了序列化,并且在大多数情况下都可以完美运行。我确实注意到某些字段类型存在问题,例如 defaultdict 字段。看起来数据类没有按预期处理此类字段类型的序列化(我猜它会将其视为普通字典)。您可以使用dataclasses.asdict() 帮助函数来序列化数据类实例,这也适用于嵌套数据类。唯一的问题是从字典中反序列化它,不幸的是,这似乎是数据类中缺少的链接。

标签: python json pydantic


【解决方案1】:

试试这个解决方案,我能够让它与pydantic一起工作。它有点丑陋,有点骇人听闻,但至少它可以按预期工作。

import pydantic


class Base(pydantic.BaseModel):
    class Config:
        extra = 'forbid'   # forbid use of extra kwargs


class Thing(Base):
    thing_id: int


class SubThing(Thing):
    name: str


class Container(Base):
    thing: Thing

    def __init__(self, **kwargs):
        # This answer helped steer me towards this solution:
        #   https://stackoverflow.com/a/66582140/10237506
        if not isinstance(kwargs['thing'], SubThing):
            kwargs['thing'] = SubThing(**kwargs['thing'])
        super().__init__(**kwargs)


def main():
    # make instance of container
    c1 = Container(
        thing=SubThing(
            thing_id=1,
            name='my_thing')
    )

    d = c1.dict()
    print(d)
    # {'thing': {'thing_id': 1, 'name': 'my_thing'}}

    # Now it works!
    c2 = Container(**d)

    print(c2)
    # thing=SubThing(thing_id=1, name='my_thing')
    
    # assert that the values for the de-serialized instance is the same
    assert c1 == c2


if __name__ == '__main__':
    main()

如果您不需要pydantic 提供的某些功能(例如数据验证),则可以轻松使用普通数据类。您可以将其与dataclass-wizard 之类的(反)序列化库配对,该库提供自动大小写转换和类型转换(例如字符串到带注释的int),其工作方式与pydantic 的工作方式大致相同。下面是一个非常简单的用法:

from dataclasses import dataclass

from dataclass_wizard import asdict, fromdict


@dataclass
class Thing:
    thing_id: int


@dataclass
class SubThing(Thing):
    name: str


@dataclass
class Container:
    # Note: I had to update the annotation to `SubThing`. otherwise
    # when de-serializing, it creates a `Thing` instance which is not
    # what we want.
    thing: SubThing


def main():
    # make instance of container
    c1 = Container(
        thing=SubThing(
            thing_id=1,
            name='my_thing')
    )

    d = asdict(c1)
    print(d)
    # {'thing': {'thingId': 1, 'name': 'my_thing'}}

    # De-serialize a dict object in a new `Container` instance
    c2 = fromdict(Container, d)

    print(c2)
    # Container(thing=SubThing(thing_id=1, name='my_thing'))

    # assert that the values for the de-serialized instance is the same
    assert c1 == c2


if __name__ == '__main__':
    main()

【讨论】:

  • 是的,你是对的。我不得不将注释更改为thing: SubThing,否则它会尝试将字典加载到Thing 类型中。我会更新答案以澄清。
  • 实际上,我现在看到了问题。我想我没有仔细阅读上面的问题。
  • 不,但是Things 的其他几个子类应该能够处理这些。
  • 我写了一个快速递归函数来手动序列化它。我将使用代码编辑我的原始问题。
  • 是的,没问题,我明白你现在要问的了。我想如果您想将字段注释为更通用的类,例如Thing,但稍后用子类(例如SubThing)填充它,那只会使反序列化回容器变得有点困难(因为它会根据注释创建一个Thing)。一个带有 pydantic 的自定义序列化程序有望适用于这个用例。
猜你喜欢
  • 2022-11-26
  • 2022-01-15
  • 2021-09-06
  • 2022-07-28
  • 1970-01-01
  • 2017-12-16
  • 2021-03-04
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多