【问题标题】:Rails get count of association through join tableRails 通过连接表获取关联计数
【发布时间】:2015-07-19 22:32:26
【问题描述】:

这个问题是HABTM associations in Rails : collecting and counting the categories of a model's children的一个分支。

给定:

class Category < ActiveRecord::Base
  has_and_belongs_to_many :books
  validates_uniqueness_of :name
end

class Book < ActiveRecord::Base
  has_and_belongs_to_many :categories
end

class Store < ActiveRecord::Base
  has_many :books
  has_many :categories, through: :books
end

任务:

给定一家商店,列出每个类别的图书数量。

Store.first.books_per_category

想要的输出:

[ { name: 'mystery', count: 5 }, { name: 'fantasy', count: 6 } ]

但是,每个商店可能都有大量的书籍和类别。

我正在尝试创建一个单一的高性能查询,它只获取与商店关联的每个不同类别的名称列和图书计数,而不将图书加载到内存中。

到目前为止我已经尝试过:

class Store < ActiveRecord::Base

  # Will load each book into memory
  def books_per_category
    categories.eager_load(:books).map do |c|
      {
          name: c.name,
          count: c.books.size # Using size instead of count is important since count will always query the DB
      }
    end
  end

  # will query books count for each category.
  def books_per_category2
    categories.distinct.map do |c|
      {
          name: c.name,
          count: c.books.count
      }
    end
  end
end

数据库架构:

ActiveRecord::Schema.define(version: 20150508184514) do

  create_table "books", force: true do |t|
    t.string   "title"
    t.datetime "created_at"
    t.datetime "updated_at"
    t.integer  "store_id"
  end

  add_index "books", ["store_id"], name: "index_books_on_store_id"

  create_table "books_categories", id: false, force: true do |t|
    t.integer "book_id",     null: false
    t.integer "category_id", null: false
  end

  add_index "books_categories", ["book_id", "category_id"], name: "index_books_categories_on_book_id_and_category_id"
  add_index "books_categories", ["category_id", "book_id"], name: "index_books_categories_on_category_id_and_book_id"

  create_table "categories", force: true do |t|
    t.string   "name"
    t.datetime "created_at"
    t.datetime "updated_at"
  end

  create_table "stores", force: true do |t|
    t.string   "name"
    t.datetime "created_at"
    t.datetime "updated_at"
  end
end

【问题讨论】:

标签: ruby-on-rails performance postgresql


【解决方案1】:

您可以使用链式select 和group 来汇总每个类别的图书数量。您的 books_per_category 方法可能如下所示:

def books_per_category
  categories.select('categories.id, categories.name, count(books.id) as count')
            .group('categories.id, categories.name').map do |c|
    {
      name: c.name,
      count: c.count
    }
  end
end

这将产生以下 SQL 查询:

SELECT categories.id, categories.name, count(books.id) as count 
  FROM "categories" 
  INNER JOIN "books_categories" ON "categories"."id" = "books_categories"."category_id" 
  INNER JOIN "books" ON "books_categories"."book_id" = "books"."id" 
  WHERE "books"."store_id" = 1 
  GROUP BY categories.id, categories.name

【讨论】:

  • 工作得非常好 - 在我的 MacBook Air 上运行大约 5-10 毫秒的 ca 500 本书,目前听起来像是受到了 astma 攻击。
【解决方案2】:

您将希望在 Categories 对象上创建一个方法(或范围),例如。

Category.joins(:books)
        .select('categories.*, COUNT(books.id) as book_count')
        .group('categories.id')

生成的对象现在将具有类别实例的每个属性,并响应一个方法,book_count,该方法返回具有该实例类别 ID 的书籍的数量。

值得注意的是,这将省略任何没有与之关联的书籍的类别。如果您想包含这些,则查询需要更新为以下内容:

Category.left_outer_joins(:books)
        .select('categories.*, COUNT(books_categories.book_id) as book_count')
        .group('categories.id')

【讨论】:

  • 谢谢,但两者都比@jakub-kosiński 提供的解决方案慢得多(815.9ms/807.9ms vs 10.5ms)。
  • select categories.name 与 categories.* 的速度有多快
  • select categories.name 稍微快一点 - 大约 10-100 毫秒。
  • 如果你想要repo,你可以自己尝试一下,你的回答有点“需要一些组装”,所以我希望我没有搞砸实现。我通过播种 Postgres 数据库并从控制台运行 Store.first.books_per_category4 对其进行了测试。
  • 哈哈,我相信你,我只是好奇。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2012-08-29
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多