【问题标题】:What is the best alternative for the following situation?以下情况的最佳选择是什么?
【发布时间】:2017-05-08 02:53:12
【问题描述】:

我在我的 Postgresql 数据库中使用 JSONB 字段来存储以下文档。我拥有数千份文件。我需要使用这些数据创建报告,但搜索速度很慢。

如果我需要创建一个报告来说明一个月的新用户,我需要查看整个文档来比较用户是否在一个月内而不是另一个月内。

消息文档:

[{"recipient":1,"user":4,"created_at":"2016-11-10","content":"Duis aliquam convallis nunc.","is_sender_user":true},
{"recipient":1,"user":18,"created_at":"2016-12-10","content":"Proin eu mi.","is_sender_user":false},
{"recipient":1,"user":4,"created_at":"2016-11-20","content":"In hac habitasse platea dictumstm.","is_sender_user":true},
{"recipient":1,"user":20,"created_at":"2016-12-14","content":"Donec ut dolor.","is_sender_user":true},
{"recipient":1,"user":13,"created_at":"2016-12-06","content":"Nulla mollis molestie lorem. Quisque ut erat. Curabitur gravida nisi at nibh.","is_sender_user":true}]

最好创建一个用户表并创建一个 JSONB 消息字段来存储您的消息。或者我可以使用 JSONB 查询创建我的报告?

【问题讨论】:

  • 考虑到您的所有字段都在顶层,您可以在整个列上创建 GIN 类型索引 - 这应该会加快速度。 Postgre 还允许您在 jsonb 中的字段上创建索引(例如 field_name -> 'created_at')

标签: postgresql jsonb


【解决方案1】:

您的消息文档描述了用户之间的关系:发送者将内容传输给接收者。发件人可能会发送许多消息,收件人可能会收到许多消息。这最好用关系结构来表示,用户表和消息表对发送者和接收者具有外键约束。

可以像您正在做的那样将所有内容放入 JSONB 字段,但有一些主要缺点:查询性能受到影响,尽管正如 Samuil Petrov 提到的,这可以通过索引来改善;但更重要的是,没有什么可以阻止消息具有无效的用户或收件人 ID。使用无模式 JSONB 字段可以简化开发,同时您仍在计算需要存储的内容,但是一旦您知道自己需要什么,就应该由您的模式强制执行。

【讨论】:

    【解决方案2】:

    正如 Samuil Petrov 提到的,您可以在 jsonb 字段上创建索引,我建议在 created_at 和 user 的月份部分创建索引

    create INDEX td002_si3 ON testData002 (substring(doc->>'created',0,8),(doc->>'user'));
    

    用这个查询

    SELECT 
          substring(doc ->> 'created', 0, 8) AS m,
          ARRAY_AGG(DISTINCT doc ->> 'user')          AS users
        FROM testData002
        GROUP BY substring(doc ->> 'created', 0, 8)
    

    将通过索引扫描为您提供每月用户数

    GroupAggregate  (cost=0.28..381.52 rows=3485 width=50)
      Group Key: ""substring""((doc ->> 'created'::text), 0, 8)
      ->  Index Scan using td002_si3 on testdata002  (cost=0.28..294.28 rows=3500 width=50)
    

    生成的测试数据

    create table testData002 as 
         select row_number() OVER () as id
               ,jsonb_build_object('created',dt::DATE
                                  ,'user',(random()*1000)::INT) as doc 
           from generate_series(1,10),generate_series('2016-01-01'::TIMESTAMP,'2016-12-15'::TIMESTAMP,'1 day'::INTERVAL) as dt;
    

    【讨论】:

    • "GroupAggregate" 是 Postgresql 函数吗?我在 pgadmin 3 上运行,但它无法识别。
    • @cske 感谢您的回复。我的 JSON 的结构有点不同。我的表中有以下内容: doc 字段接收数组 [{"user": 1, "created": "2016-01-01"}, { "user": 14, "created": "2016-02 - 04 "}, ...]
    • @JoãoPedro 可以轻松转换为答案中使用的结构
    猜你喜欢
    • 2020-12-05
    • 2017-09-28
    • 1970-01-01
    • 1970-01-01
    • 2015-08-30
    • 2012-08-02
    • 1970-01-01
    • 1970-01-01
    • 2021-11-13
    相关资源
    最近更新 更多