【问题标题】:Django query with annotation and conditional count too slow带有注释和条件计数的 Django 查询太慢
【发布时间】:2017-03-25 02:33:33
【问题描述】:

我有这个带有注释、计数和条件表达式的查询,运行速度非常慢,需要很长时间。

我有两个模型,一个存储 instagram 出版物,另一个存储 twitter 出版物。每个出版物还具有另一个模型的 FK,该模型代表城市内的六边形地理区域。

出版物 [FK] -> HexCityArea

TwitterPublication [FK] -> HexCityArea

我正在尝试计算每个六边形的出版物,但出版物已被日期等其他字段预先过滤,因此代码为:

instagram_publications_ids = list(instagram_publications.values_list('id', flat=True))
twitter_publications_ids = list(twitter_publications.values_list('id', flat=True))

print "\n[HEXAGONS QUERY]> List of publications ids insta\n %s \n" % instagram_publications.query
print instagram_publications.explain()
print "\n[HEXAGONS QUERY]> List of publications ids twitter\n %s \n" % twitter_publications.query
print twitter_publications.explain()

# Get count of publications by hexagon
resultant_hexagons = HexagonalCityArea.objects.filter(city=city).annotate(
    instagram_count=Count(Case(
        When(publication__id__in=instagram_publications_ids, then=1),
        output_field=IntegerField(),
    ))
).annotate(
    twitter_count=Count(Case(
        When(twitterpublication__id__in=twitter_publications_ids, then=1),
        output_field=IntegerField(),
    ))
)#filter(instagram_count__gt=0).filter(twitter_count__gt=0) # Discard empty hexagons

# For debug only
print "\n[HEXAGONS QUERY]> Count of publications\n %s \n" % resultant_hexagons.query
print resultant_hexagons.explain()

resultant_hexagons_list = list(resultant_hexagons)
# Iterate remaining hexagons
city_hexagons = [h for h in resultant_hexagons_list if h.instagram_count > 0 or h.twitter_count > 0]

如您所见,首先我获取所选出版物的 ID 列表,然后使用它们仅计算这些出版物。

我看到的一个问题是 ID 列表非常长,大约有 28000 个元素,但是如果我不使用 ID 列表,我不会得到想要的结果,计数条件不能正常工作并且该市的所有出版物都被计算在内。

我已尝试这样做以避免使用 ID 列表:

        resultant_hexagons = HexagonalCityArea.objects.filter(city=city).annotate(
            instagram_count=Count(Case(

                When(publication__in=instagram_publications, then=1),
                output_field=IntegerField(),
            ))
        ).annotate(
            twitter_count=Count(Case(

                When(twitterpublication__in=twitter_publications, then=1),
                output_field=IntegerField(),
            ))
        ).filter(instagram_count__gt=0).filter(twitter_count__gt=0) # Discard empty hexagons

        # For debug only
        print "\n[HEXAGONS QUERY]> Count of publications\n %s \n" % resultant_hexagons.query
        print resultant_hexagons.explain()

这是生成的 SQL:

SELECT
   "instanalysis_hexagonalcityarea"."id",
   "instanalysis_hexagonalcityarea"."created",
   "instanalysis_hexagonalcityarea"."modified",
   "instanalysis_hexagonalcityarea"."geom",
   "instanalysis_hexagonalcityarea"."city_id",
   COUNT(
   CASE
      WHEN
         "instanalysis_publication"."id" IN 
         (
            SELECT
               U0."id" 
            FROM
               "instanalysis_publication" U0 
               INNER JOIN
                  "instanalysis_instagramlocation" U1 
                  ON (U0."location_id" = U1."id") 
               INNER JOIN
                  "instanalysis_spot" U2 
                  ON (U1."spot_id" = U2."id") 
               INNER JOIN
                  "instanalysis_city" U3 
                  ON (U2."city_id" = U3."id") 
            WHERE
               (
                  U3."name" = Durban 
                  AND U0."publication_date" >= 2016 - 12 - 01 00:00:00 + 01:00 
                  AND U0."publication_date" <= 2016 - 12 - 11 00:00:00 + 01:00
               )
         )
      THEN
         1 
      ELSE
         NULL 
   END
) AS "instagram_count", COUNT(
   CASE
      WHEN
         "instanalysis_twitterpublication"."id" IN 
         (
            SELECT
               U0."id" 
            FROM
               "instanalysis_twitterpublication" U0 
               INNER JOIN
                  "instanalysis_twitterlocation" U1 
                  ON (U0."location_id" = U1."id") 
               INNER JOIN
                  "instanalysis_spot" U2 
                  ON (U1."spot_id" = U2."id") 
               INNER JOIN
                  "instanalysis_city" U3 
                  ON (U2."city_id" = U3."id") 
            WHERE
               (
                  U3."name" = Durban 
                  AND U0."publication_date" >= 2016 - 12 - 01 00:00:00 + 01:00 
                  AND U0."publication_date" <= 2016 - 12 - 11 00:00:00 + 01:00
               )
         )
      THEN
         1 
      ELSE
         NULL 
   END
) AS "twitter_count" 
FROM
   "instanalysis_hexagonalcityarea" 
   LEFT OUTER JOIN
      "instanalysis_publication" 
      ON ("instanalysis_hexagonalcityarea"."id" = "instanalysis_publication"."hexagon_id") 
   LEFT OUTER JOIN
      "instanalysis_twitterpublication" 
      ON ("instanalysis_hexagonalcityarea"."id" = "instanalysis_twitterpublication"."hexagon_id") 
WHERE
   "instanalysis_hexagonalcityarea"."city_id" = 7 
GROUP BY
   "instanalysis_hexagonalcityarea"."id" 
HAVING
(COUNT(
   CASE
      WHEN
         "instanalysis_publication"."id" IN 
         (
            SELECT
               U0."id" 
            FROM
               "instanalysis_publication" U0 
               INNER JOIN
                  "instanalysis_instagramlocation" U1 
                  ON (U0."location_id" = U1."id") 
               INNER JOIN
                  "instanalysis_spot" U2 
                  ON (U1."spot_id" = U2."id") 
               INNER JOIN
                  "instanalysis_city" U3 
                  ON (U2."city_id" = U3."id") 
            WHERE
               (
                  U3."name" = Durban 
                  AND U0."publication_date" >= 2016 - 12 - 01 00:00:00 + 01:00 
                  AND U0."publication_date" <= 2016 - 12 - 11 00:00:00 + 01:00
               )
         )
      THEN
         1 
      ELSE
         NULL 
   END
) > 0 
   AND COUNT(
   CASE
      WHEN
         "instanalysis_twitterpublication"."id" IN 
         (
            SELECT
               U0."id" 
            FROM
               "instanalysis_twitterpublication" U0 
               INNER JOIN
                  "instanalysis_twitterlocation" U1 
                  ON (U0."location_id" = U1."id") 
               INNER JOIN
                  "instanalysis_spot" U2 
                  ON (U1."spot_id" = U2."id") 
               INNER JOIN
                  "instanalysis_city" U3 
                  ON (U2."city_id" = U3."id") 
            WHERE
               (
                  U3."name" = Durban 
                  AND U0."publication_date" >= 2016 - 12 - 01 00:00:00 + 01:00 
                  AND U0."publication_date" <= 2016 - 12 - 11 00:00:00 + 01:00
               )
         )
      THEN
         1 
      ELSE
         NULL 
   END
) > 0)

这要快得多,见解释分析:

GroupAggregate  (cost=1.14..743590.08 rows=3300 width=184) (actual time=5186.606..46907.530 rows=334 loops=1)
  Group Key: instanalysis_hexagonalcityarea.id
  Filter: ((count(CASE WHEN (hashed SubPlan 3) THEN 1 ELSE NULL::integer END) > 0) AND (count(CASE WHEN (hashed SubPlan 4) THEN 1 ELSE NULL::integer END) > 0))
  Rows Removed by Filter: 2966
  ->  Merge Left Join  (cost=1.14..320194.96 rows=7166797 width=184) (actual time=4851.792..17369.232 rows=70436610 loops=1)
        Merge Cond: (instanalysis_hexagonalcityarea.id = instanalysis_publication.hexagon_id)
        ->  Merge Left Join  (cost=0.71..21686.40 rows=49328 width=180) (actual time=109.033..164.451 rows=30857 loops=1)
              Merge Cond: (instanalysis_hexagonalcityarea.id = instanalysis_twitterpublication.hexagon_id)
              ->  Index Scan using instanalysis_hexagonalcityarea_pkey on instanalysis_hexagonalcityarea  (cost=0.29..591.47 rows=3300 width=176) (actual time=22.783..23.878 rows=3300 loops=1)
                    Filter: (city_id = 7)
                    Rows Removed by Filter: 7282
              ->  Index Scan using instanalysis_twitterpublication_5c78aecb on instanalysis_twitterpublication  (cost=0.42..64392.25 rows=504291 width=8) (actual time=0.018..111.677 rows=170305 loops=1)
        ->  Materialize  (cost=0.43..501402.61 rows=3754731 width=8) (actual time=0.011..6788.670 rows=71922153 loops=1)
              ->  Index Scan using instanalysis_publication_5c78aecb on instanalysis_publication  (cost=0.43..492015.78 rows=3754731 width=8) (actual time=0.005..4034.838 rows=1778030 loops=1)
  SubPlan 1
    ->  Nested Loop  (cost=0.72..105061.24 rows=27624 width=4) (actual time=0.326..74.024 rows=21824 loops=1)
          ->  Nested Loop  (cost=0.29..620.11 rows=2767 width=4) (actual time=0.024..2.915 rows=3374 loops=1)
                ->  Nested Loop  (cost=0.00..143.13 rows=504 width=4) (actual time=0.016..0.618 rows=829 loops=1)
                      Join Filter: (u2.city_id = u3.id)
                      Rows Removed by Join Filter: 3350
                      ->  Seq Scan on instanalysis_city u3  (cost=0.00..1.10 rows=1 width=4) (actual time=0.004..0.006 rows=1 loops=1)
                            Filter: ((name)::text = 'Durban'::text)
                            Rows Removed by Filter: 7
                      ->  Seq Scan on instanalysis_spot u2  (cost=0.00..89.79 rows=4179 width=8) (actual time=0.001..0.242 rows=4179 loops=1)
                ->  Index Scan using instanalysis_instagramlocation_e72b53d4 on instanalysis_instagramlocation u1  (cost=0.29..0.89 rows=6 width=8) (actual time=0.001..0.002 rows=4 loops=829)
                      Index Cond: (spot_id = u2.id)
          ->  Index Scan using instanalysis_publication_e274a5da on instanalysis_publication u0  (cost=0.43..37.45 rows=30 width=8) (actual time=0.006..0.021 rows=6 loops=3374)
                Index Cond: (location_id = u1.id)
                Filter: ((publication_date >= '2016-11-30 23:00:00+00'::timestamp with time zone) AND (publication_date <= '2016-12-10 23:00:00+00'::timestamp with time zone))
                Rows Removed by Filter: 80
  SubPlan 2
    ->  Hash Join  (cost=2595.62..25893.51 rows=9013 width=4) (actual time=22.511..73.141 rows=6220 loops=1)
          Hash Cond: (u0_1.location_id = u1_1.id)
          ->  Seq Scan on instanalysis_twitterpublication u0_1  (cost=0.00..22927.36 rows=74772 width=8) (actual time=15.212..59.628 rows=75775 loops=1)
                Filter: ((publication_date >= '2016-11-30 23:00:00+00'::timestamp with time zone) AND (publication_date <= '2016-12-10 23:00:00+00'::timestamp with time zone))
                Rows Removed by Filter: 428516
          ->  Hash  (cost=2348.24..2348.24 rows=19790 width=4) (actual time=6.538..6.538 rows=15589 loops=1)
                Buckets: 32768  Batches: 1  Memory Usage: 805kB
                ->  Nested Loop  (cost=0.70..2348.24 rows=19790 width=4) (actual time=0.023..5.052 rows=15589 loops=1)
                      ->  Nested Loop  (cost=0.28..39.28 rows=504 width=4) (actual time=0.015..0.186 rows=829 loops=1)
                            ->  Seq Scan on instanalysis_city u3_1  (cost=0.00..1.10 rows=1 width=4) (actual time=0.003..0.004 rows=1 loops=1)
                                  Filter: ((name)::text = 'Durban'::text)
                                  Rows Removed by Filter: 7
                            ->  Index Scan using instanalysis_spot_c7141997 on instanalysis_spot u2_1  (cost=0.28..33.14 rows=504 width=8) (actual time=0.010..0.124 rows=829 loops=1)
                                  Index Cond: (city_id = u3_1.id)
                      ->  Index Scan using instanalysis_twitterlocation_e72b53d4 on instanalysis_twitterlocation u1_1  (cost=0.42..3.93 rows=65 width=8) (actual time=0.001..0.004 rows=19 loops=829)
                            Index Cond: (spot_id = u2_1.id)
  SubPlan 3
    ->  Nested Loop  (cost=0.72..105061.24 rows=27624 width=4) (actual time=0.348..80.863 rows=21824 loops=1)
          ->  Nested Loop  (cost=0.29..620.11 rows=2767 width=4) (actual time=0.028..3.507 rows=3374 loops=1)
                ->  Nested Loop  (cost=0.00..143.13 rows=504 width=4) (actual time=0.016..0.646 rows=829 loops=1)
                      Join Filter: (u2_2.city_id = u3_2.id)
                      Rows Removed by Join Filter: 3350
                      ->  Seq Scan on instanalysis_city u3_2  (cost=0.00..1.10 rows=1 width=4) (actual time=0.003..0.004 rows=1 loops=1)
                            Filter: ((name)::text = 'Durban'::text)
                            Rows Removed by Filter: 7
                      ->  Seq Scan on instanalysis_spot u2_2  (cost=0.00..89.79 rows=4179 width=8) (actual time=0.001..0.276 rows=4179 loops=1)
                ->  Index Scan using instanalysis_instagramlocation_e72b53d4 on instanalysis_instagramlocation u1_2  (cost=0.29..0.89 rows=6 width=8) (actual time=0.001..0.003 rows=4 loops=829)
                      Index Cond: (spot_id = u2_2.id)
          ->  Index Scan using instanalysis_publication_e274a5da on instanalysis_publication u0_2  (cost=0.43..37.45 rows=30 width=8) (actual time=0.007..0.022 rows=6 loops=3374)
                Index Cond: (location_id = u1_2.id)
                Filter: ((publication_date >= '2016-11-30 23:00:00+00'::timestamp with time zone) AND (publication_date <= '2016-12-10 23:00:00+00'::timestamp with time zone))
                Rows Removed by Filter: 80
  SubPlan 4
    ->  Hash Join  (cost=2595.62..25893.51 rows=9013 width=4) (actual time=41.392..92.680 rows=6220 loops=1)
          Hash Cond: (u0_3.location_id = u1_3.id)
          ->  Seq Scan on instanalysis_twitterpublication u0_3  (cost=0.00..22927.36 rows=74772 width=8) (actual time=32.641..78.020 rows=75775 loops=1)
                Filter: ((publication_date >= '2016-11-30 23:00:00+00'::timestamp with time zone) AND (publication_date <= '2016-12-10 23:00:00+00'::timestamp with time zone))
                Rows Removed by Filter: 428516
          ->  Hash  (cost=2348.24..2348.24 rows=19790 width=4) (actual time=7.907..7.907 rows=15589 loops=1)
                Buckets: 32768  Batches: 1  Memory Usage: 805kB
                ->  Nested Loop  (cost=0.70..2348.24 rows=19790 width=4) (actual time=0.044..6.136 rows=15589 loops=1)
                      ->  Nested Loop  (cost=0.28..39.28 rows=504 width=4) (actual time=0.026..0.220 rows=829 loops=1)
                            ->  Seq Scan on instanalysis_city u3_3  (cost=0.00..1.10 rows=1 width=4) (actual time=0.006..0.008 rows=1 loops=1)
                                  Filter: ((name)::text = 'Durban'::text)
                                  Rows Removed by Filter: 7
                            ->  Index Scan using instanalysis_spot_c7141997 on instanalysis_spot u2_3  (cost=0.28..33.14 rows=504 width=8) (actual time=0.016..0.135 rows=829 loops=1)
                                  Index Cond: (city_id = u3_3.id)
                      ->  Index Scan using instanalysis_twitterlocation_e72b53d4 on instanalysis_twitterlocation u1_3  (cost=0.42..3.93 rows=65 width=8) (actual time=0.001..0.005 rows=19 loops=829)
                            Index Cond: (spot_id = u2_3.id)
Planning time: 50.735 ms
Execution time: 46908.482 ms

问题是我没有得到我想要的,它似乎在计算更多的出版物。出版物之前按日期过滤,我只想计算每个六边形中有多少过滤过的出版物,但它似乎是按六边形计算所有出版物,就像当子句不起作用时一样。

感谢您的帮助。

【问题讨论】:

  • 为什么count aggregate 不是一个选项?理论上,使用count 的两个聚合查询应该比使用 IN 子句的联合查询更有效
  • 感谢您的 cmets @Marat。是的,你的方式要快得多,但问题是我得到了错误的结果。我已经用 SQL 更新了帖子并解释了分析。

标签: python sql django postgresql database-performance


【解决方案1】:

它变慢的最大原因是子查询,即对于 hexarea 数据库服务器中的每条记录,都会发出另一个查询来计算与其 id 匹配的 instagram/twitter 记录。即使在更新之后,它仍然在本质上做同样的事情。

如何解决:运行聚合查询。这样,DB 服务器可以只在 id 列表中线性运行一次,这可能效率要高一个数量级。示例:

from django.db.models import Count

# assuming "instagram_publications" is the related_name
# of the correspondent Instagram/Twitter post model 
instacounts = HexagonalCityArea.objects.filter(city=city
             ).filter(instagram_publications__publicat‌​ion_date__lte=end_da‌​te
             ).filter(instagram_publications__publicat‌​ion_date__gte=start_da‌​te
             ).aggregate(Count('instagram_publications')))

【讨论】:

  • 再次感谢@Marat。我现在看到聚合比注释快得多,但我怎么能使用聚合条件来只计算某些 instagram 出版物而不是全部?在我的情况下,instagram_publications 不是相关模型的related_name,它是一个查询集,其中包含先前过滤的出版物,例如:instagram_publications = Publication.objects.filter(location__spot__city__name=location).filter(publication_date__gte=start_date).filter(publication_date__lte=end_date)跨度>
  • @MartinezMariano 我更新了这个例子。您不需要按位置过滤帖子,因为 hexarea 已经由确切位置标识
  • 再次感谢@Marat,这种方法有两个问题。第一个是聚合返回单个元组,其中所有六边形的总数加起来,我需要六边形的计数,儿子我应该使用注释。我尝试注释时的第二个问题是,在对 Hexagon 模型进行查询时应用于发布模型的那些过滤器似乎工作错误,因为我为每个六边形获得的最终计数是错误的,返回的结果是太大了。
猜你喜欢
  • 2022-01-19
  • 1970-01-01
  • 2018-04-16
  • 2020-05-14
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-05-13
  • 1970-01-01
相关资源
最近更新 更多