【问题标题】:CosmosDB Distinct query - slow performanceCosmosDB Distinct 查询 - 性能缓慢
【发布时间】:2022-01-09 16:15:14
【问题描述】:

我们正在使用 CosmosDB,我正在运行 Distinct 查询,如下所示

Select Distinct c.SomeType, c.SomeName
From c
Where c.ptkey = 'WHATEVERPTKEY'
And c.SomeCategory = 'WhateverCategory'

...其中 ptkey 是保存分区键的字段。上述工作但需要大约 1-1.5 分钟才能完成(我假设因为一些/许多文档非常大) - 我尝试使用“分组依据”过滤分区的唯一键(id),并使用“Order By”(当您组合这两个和/或 Order By 中只允许一个字段时应用限制,除非您有复合键),但变化不大。

确实产生重大影响的一件事是创建如下索引策略

"compositeIndexes": [
    [
        {
            "path": "/ptkey"
        },
        {
            "path": "/SomeCategory"
        },
        {
            "path": "/SomeType"
        },
        {
            "path": "/SomeName"
        }
    ]
]

...但是,我的问题是如何将这个复合键定义限制为仅适用于上述查询所针对的特定分区键('WHATEVERPTKEY' - 因为我们的数据库/集合中有大约十几个分区键),其次是否有任何替代/更好的选择(除了重新建模我们的数据)

在 Azure CosmosDB 数据资源管理器中运行查询时,请注意我的查询统计信息没有索引如下

Query Statistics
746.96 RUs
1 - 12
Retrieved document count 0
Retrieved document size 0 bytes
Output document count 0
Output document size 1789 bytes
Index hit document count 0
Index lookup time 1.1900000000000002 ms
Document load time 366.48990000000003 ms
Query engine execution time 29.0101 ms
System function execution time 0.76 ms
User defined function execution time 0 ms
Document write time 0.02 ms
1

更新

  • 集合的完整索引策略如下
{
    "indexingMode": "consistent",
    "automatic": true,
    "includedPaths": [
        {
            "path": "/*"
        }
    ],
    "excludedPaths": [
        {
            "path": "/\"_etag\"/?"
        }
    ],
    "compositeIndexes": [
        [
            {
                "path": "/ptkey",
                "order": "ascending"
            },
            {
                "path": "/SomeCategory",
                "order": "ascending"
            },
            {
                "path": "/SomeType",
                "order": "ascending"
            },
            {
                "path": "/SomeName",
                "order": "ascending"
            }
        ]
    ]
}

【问题讨论】:

  • 复合索引对此查询没有影响。您能否包含此容器的完整索引策略?您肯定没有使用基于索引命中数的索引。
  • @MarkBrown - 我已经用索引策略更新了帖子 - 我正在运行的查询在索引策略到位后看起来性能有所提高,如下所示 SELECT uniq.ptkey, uniq.SomeCategory,uniq.SomeType,uniq.SomeName FROM ( SELECT c.ptkey,c.SomeCategory,c.SomeType,c.SomeName from c where c.ptkey = 'WHATEVERPTKEY' and c.SomeCategory='WhateverCategory' group by c .ptkey,c.SomeCategory,c.SomeType, c.SomeName ) as uniq
  • 上述查询带来以下查询统计信息:7809.127 RUs, 1 - 80, Retrieved document count 21000, Retrieved document size 1893070755 bytes, Output document count 80, Output document size 7095 bytes, Index hit document count 30.01,索引查找时间1.1400000000000001 ms,文档加载时间3801.42 ms,查询引擎执行时间316.4699 ms,系统函数执行时间0 ms,用户定义函数执行时间0 ms,文档写入时间0.17 ms,1

标签: performance query-optimization azure-cosmosdb


【解决方案1】:

因此,当存在大量多页结果时,门户中的查询指标有时无法显示准确的数据,但它通常确实有效。

使用 DISTINCT,成本可能取决于您处理的结果数量。如果您只期待一些结果,那么对成本的影响很小。如果你成千,它可能会变得非常昂贵。正在开展工作以降低成本。在它发布之前还有一段路要走。

您可以将这个作为您的索引策略再试一次吗?

{
"indexingMode": "consistent",
"automatic": true,
"includedPaths": [
    {
        "path": "/*"
    }
],
"excludedPaths": [
    {
        "path": "/\"_etag\"/?"
    }
],
"compositeIndexes": [
    [
        {
            "path": "/ptkey",
            "order": "ascending"
        },
        {
            "path": "/SomeCategory",
            "order": "ascending"
        }
    ]
]

}

【讨论】:

  • 会试一试 - 最初的问题是关于如何将特定索引策略隔离到单个分区键 - 你知道这是如何完成的吗?
  • 这是不可能的。不知道你为什么要这样做,因为这与拥有一个可以扩展的集合有点矛盾。如果容器中有不同的实体对分区键值是唯一的,只需为它们定义索引。鉴于您只谈论 20GB 的数据,您会注意到很大的不同。另外,鉴于它是按 PK 计算的,只要您在 WHERE 中指定 PK,性能就可以了。
猜你喜欢
  • 2015-04-11
  • 2016-03-09
  • 1970-01-01
  • 1970-01-01
  • 2019-06-25
  • 2022-01-04
  • 2016-05-10
  • 1970-01-01
  • 2023-03-30
相关资源
最近更新 更多