【问题标题】:Slow Berlin sparql benchmark queries in Neo4jNeo4j 中缓慢的柏林 sparql 基准查询
【发布时间】:2014-06-21 23:19:57
【问题描述】:

我正在 neo4j 中尝试柏林基准 SPARQL 查询。我使用http://michaelbloggs.blogspot.de/2013/05/importing-ttl-turtle-ontologies-in-neo4j.html从三元组创建了 Neo4j 图

总结数据加载,我的图表有以下结构,

Subject   => Node
Predicate => Relationship
Object    => Node 

如果谓词是日期、字符串、整数(原始),则创建属性而不是关系并将其存储在节点中。

现在,我正在尝试跟踪在 Noe4j 中非常慢的查询,

Query 4: Feature with the highest ratio between price with that feature and price without that feature. 

    corresponding SPARQL query for this, 

            prefix bsbm: <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/vocabulary/>
            prefix bsbm-inst: <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/>
            prefix xsd: <http://www.w3.org/2001/XMLSchema#>

            Select ?feature ((?sumF*(?countTotal-?countF))/(?countF*(?sumTotal-?sumF)) As ?priceRatio)
            {
              { Select (count(?price) As ?countTotal) (sum(xsd:float(str(?price))) As ?sumTotal)
                {
                  ?product a <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/ProductType294> .
                  ?offer bsbm:product ?product ;
                         bsbm:price ?price .
                }
              }
              { Select ?feature (count(?price2) As ?countF) (sum(xsd:float(str(?price2))) As ?sumF)
                {
                  ?product2 a <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/ProductType294> ;
                           bsbm:productFeature ?feature .
                  ?offer2 bsbm:product ?product2 ;
                         bsbm:price ?price2 .
                }
                Group By ?feature
              }
            }
           Order By desc(?priceRatio) ?feature
           Limit 100
 Cypher query I created for this,

    MATCH p1 = (offer1:Offer)-[r1:`product`]->(products1:ProductType294)
    MATCH p2 = (offer2:Offer)-[r2:`product`]->products2:ProductType294)-[:`productFeature`]->features
    return (sum( DISTINCT offer2.price) * ( count( DISTINCT offer1.price) - count( DISTINCT offer2.price)) /(count(DISTINCT offer2.price)*(sum( DISTINCT offer1.price) - sum(DISTINCT offer2.price)))) AS cnt,features.__URI__ AS frui
    ORDER BY cnt DESC,frui 

这个查询真的很慢,请让我知道我是否以错误的方式制定查询。

Another query is Query 5: Show the most popular products of a specific product type for each country - by review count ,

      prefix bsbm: <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/vocabulary/>
      prefix bsbm-inst: <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/>
      prefix rev: <http://purl.org/stuff/rev#>
      prefix xsd: <http://www.w3.org/2001/XMLSchema#>

      Select ?country ?product ?nrOfReviews ?avgPrice
      {
        { Select ?country (max(?nrOfReviews) As ?maxReviews)
          {
            { Select ?country ?product (count(?review) As ?nrOfReviews)
              {
                ?product a <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/ProductType403> .
                ?review bsbm:reviewFor ?product ;
                        rev:reviewer ?reviewer .
                ?reviewer bsbm:country ?country .
              }
              Group By ?country ?product
            }
          }
          Group By ?country
        }
        { Select ?product (avg(xsd:float(str(?price))) As ?avgPrice)
          {
            ?product a <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/ProductType403> .
            ?offer bsbm:product ?product .
            ?offer bsbm:price ?price .
          }
          Group By ?product
        }
        { Select ?country ?product (count(?review) As ?nrOfReviews)
          {
            ?product a <http://www4.wiwiss.fu-berlin.de/bizer/bsbm/v01/instances/ProductType403> .
            ?review bsbm:reviewFor ?product .
            ?review rev:reviewer ?reviewer .
            ?reviewer bsbm:country ?country .
          }
          Group By ?country ?product
        }
        FILTER(?nrOfReviews=?maxReviews)
      }
      Order By desc(?nrOfReviews) ?country ?product

Cypher query I created for this is following,

    MATCH (products2:ProductType403)<-[:`reviewFor`]-(reviews:Review)-[:`reviewer`]->(rvrs)-[:`country`]->(countries)
    with count(reviews) AS reviewcount,products2.__URI__ AS pruis, countries.__URI__ AS cntrs
    MATCH (products1:ProductType403)<-[:`product`]-(offer:Offer)
    with AVG(offer.price) AS avgPrice, MAX(reviewcount) AS maxrevs, cntrs
    MATCH (products2:ProductType403)<-[:`reviewFor`]-(reviews:Review)-[:`reviewer`]->(rvrs)-[:`country`]->(countries)
    with avgPrice, maxrevs,countries, count(reviews) AS rvs, countries.__URI__ AS curis, products2.__URI__ AS puris
    where maxrevs=rvs
    RETURN curis,puris,rvs,avgPrice

即使是这个查询也很慢。我是否以正确的方式制定查询?

  • 我有 10M 三元组(柏林基准数据集)
  • 每个类型谓词都转换为标签。
  • (对于查询 4)我想要得到的是与价格之间的比率最高的功能
  • 该功能和没有该功能的价格。这是正确的方法吗 制定查询?
  • (对于查询 4)我得到了此查询的正确结果。
  • 如果我不计算总和和计数,则查询会很快执行。

在此先感谢 :) SPARQL 查询和信息可在以下位置找到:http://wifo5-03.informatik.uni-mannheim.de/bizer/berlinsparqlbenchmark/spec/BusinessIntelligenceUseCase/index.html#queries

【问题讨论】:

  • 也许您可以共享您导入的数据库,并添加您正在使用的图形模型的模型图片。然后我们可以尝试帮助你。只是从描述中,我真的不明白你在做什么。除了您的查询似乎是图全局查询。

标签: neo4j cypher graph-databases


【解决方案1】:

对我来说,这些看起来像是全局图形查询? 你的数据集有多大?

您在两条路径之间创建了笛卡尔积? 这两条路径不应该以某种方式连接吗?

ProductType 标签上不应该有属性type 吗? (:ProductType {type:"294"}) 如果你有一个关于 :ProductType(type) 的索引,可能还有 :Order(orderNo)

我真的不懂计算?

不同价格计数的增量乘以报价 2 的不同价格之和 经过 报价 2 的不同价格的计数,乘以两个订单价格之和的 delta?

MATCH (offer1:Offer)-[r1:`product`]->(products1:ProductType294)
MATCH (offer2:Offer)-[r2:`product`]->(products2:ProductType294)-[:`productFeature`]->features

RETURN (sum( DISTINCT offer2.price) * 
       ( count( DISTINCT offer1.price) - count( DISTINCT offer2.price)) 
       / (count(DISTINCT offer2.price)*
       (sum( DISTINCT offer1.price) - sum(DISTINCT offer2.price)))) 
       AS cnt,features.__URI__ AS frui
ORDER BY cnt DESC,frui 

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-04-10
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多