【问题标题】:How to fix error: Script evaluation exceeded the configured 'scriptEvaluationTimeout' threshold of 30000如何修复错误:脚本评估超出了配置的“scriptEvaluationTimeout”阈值 30000
【发布时间】:2020-06-16 01:08:18
【问题描述】:

我们有连接到 Gremlin 服务器的主要 API(从 link 构建的 EC2 实例)。它使用 DynamoDB 来存储数据。

一个查询有效,但是当我们同时尝试多个查询时,问题就开始出现了。它正在获取“scriptEvaluationTimeout”。

查询是这样的: g.V().has("users","email","user@example.com").inE().otherV().where(inE().otherV().has('users','email','user@example.com')).valueMap()

尝试了以下方法:

  1. 将 EC2 实例从 t2.micro 升级到 t2.medium。我们发现实例内存不足,无法运行 Java Runtime Environment,因此我们进行了升级。
  2. 已将 DynamoDB 容量升级到 Auto Scaling,因为它超出了预置的静态容量。
  3. 尝试将超时更新为 60000 并尝试更新其他配置,但没有任何效果。
  4. 尝试添加索引,但不确定这是否是添加索引的正确方法。

这是我第一次在这里发布问题,也是我第一次使用 Gremlin,JanusGraph。最初创建这个的开发人员不再有联系,我正在通过这个平台寻求帮助。有没有人经历过这个?请帮忙。谢谢。

更新/补充:

  1. 我在上面发布的查询只是遇到超时问题的查询之一。这个更容易受到该错误的影响。
WARN  org.apache.tinkerpop.gremlin.server.op.AbstractEvalOpProcessor  - Script evaluation exceeded the configured threshold for request [RequestMessage{, requestId=d60d8cf0-1b74-11ea-9299-d76eb09532d9, op='eval', processor='', args={gremlin=g.V().or(__.has('users','phone',within('+918329086936','112','18003001947')),__.has('users','email',within('user1@example.com','user2@yahoo.com','user3@example.com','user4@fastmail.com','user5@yahoo.com','user6@example.com','user7@gmail.com','user8@gmail.com','user9@gmail.com','user10@hotmail.com'))).not(has('users','email',within('dev@example.com','user5@yahoo.com','user1@example.com','user3@example.com','user9@gmail.com'))).not(has('users','email','dev@example.com')).order().by('name').valueMap(), bindings={}, accept=application/json, language=gremlin-groovy}}]
java.util.concurrent.TimeoutException: Script evaluation exceeded the configured 'scriptEvaluationTimeout' threshold of 30000 ms or evaluation was otherwise cancelled directly for request [g.V().or(__.has('users','phone',within('+918329086936','112','18003001947')),__.has('users','email',within('user1@example.com','user2@yahoo.com','user3@example.com','user4@fastmail.com','user5@yahoo.com','user6@example.com','user7@gmail.com','user8@gmail.com','user9@gmail.com','user10@hotmail.com'))).not(has('users','email',within('dev@example.com','user5@yahoo.com','user1@example.com','user3@example.com','user9@gmail.com'))).not(has('users','email','dev@example.com')).order().by('name').valueMap()]
    at org.apache.tinkerpop.gremlin.groovy.engine.GremlinExecutor.lambda$eval$1(GremlinExecutor.java:337)
    at io.netty.util.concurrent.PromiseTask$RunnableAdapter.call(PromiseTask.java:38)
    at io.netty.util.concurrent.ScheduledFutureTask.run(ScheduledFutureTask.java:120)
    at io.netty.util.concurrent.SingleThreadEventExecutor.runAllTasks(SingleThreadEventExecutor.java:399)
    at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:464)
    at io.netty.util.concurrent.SingleThreadEventExecutor$2.run(SingleThreadEventExecutor.java:131)
    at java.lang.Thread.run(Thread.java:748)
  1. 我也在尝试添加索引。我上周发现了这个answer,并且能够创建没有错误的索引,但没有改进查询。我还尝试使用来自JanusGraph documentation 的此引用添加复合索引,但出现错误:
groovy.lang.MissingPropertyException: No such property: graph for class: groovysh_evaluate

【问题讨论】:

    标签: amazon-web-services amazon-dynamodb gremlin tinkerpop3 gremlin-server


    【解决方案1】:

    很难说到底哪里出了问题,因为要考虑的因素很多。您提到您添加了一个索引,但没有指定什么样的索引。鉴于您的遍历,我认为最重要的是“电子邮件”属性上的复合索引,以便您快速初始查找起始顶点(我假设您期望单个起始顶点)。我想你会想确保这部分遍历是“快速的”:

    g.V().has("users","email","user@example.com")
    

    如果不是,那么您的索引可能设置不正确。之后,您需要了解图表所遍历的结构:

    g.V().has("users","email","user@example.com").
      inE().otherV().
      where(inE().otherV().has('users','email','user@example.com')).
      valueMap()
    

    首先请注意,当您确定遍历方向时,通常应避免使用otherV(),因此:

    g.V().has("users","email","user@example.com").
      in().
      where(__.in().has('users','email','user@example.com')).
      valueMap()
    

    然后问题就变成了,您通过该遍历分析了多少条路径,是否有机会将它们过滤一下(您可以向in() 提供边缘标签或将has() 属性过滤器应用于边缘,以便您可以例如使用以顶点为中心的索引)。

    否则,这是一个相当简单的遍历,所以除非遍历in().in() 会强制分析数十万或数百万条边,否则我不认为这是一个非常耗时的查询,其中包含前面讨论的索引。如果您还没有这样做,您可能会考虑为 Gremlin Server 提供更多内存 - 您可以在 gremlin-server.sh(或者在您的情况下可能是 janusgraph.sh)中看到 -Xmx 设置。

    除此之外,我想不出你还可以尝试什么。我不确定 JanusGraph 的 DynamoDB 后端是否最受欢迎。您可以尝试在JanusGraph User Mailing List 上询问更多相关信息,看看人们对此有何看法。

    【讨论】:

    • 感谢您回答@stephenmallette。我在上面发布的查询只是遇到超时问题的查询之一。我有一个更容易受到该错误影响的查询。我会在此处的某处发布日志(此评论部分的字符太长)。
    • 我已经更新了上面的问题,并从日志中添加了更多详细信息和错误。谢谢。
    • 根据您的其他信息,我假设您没有创建索引。您链接到的答案是针对 TinkerGraph,而不是 JanusGraph,所以我不确定它会产生什么影响。您应该按照指向 JanusGraph 文档的链接中的说明创建索引。您收到的错误意味着“图形”作为变量名称尚未在 Gremlin 服务器中正确配置 - 也许您将其命名为其他名称。
    • 你说的遍历更容易超时可能只是缺少索引的问题,但我也想知道 JanusGraph 如何优化该查询。考虑到那里的复杂性,即使您成功创建了索引 JanusGraph 也可能无法正确解决您在那里使用它的条件。我会 explain() 和 profile() 遍历以查看它是如何编译/执行的。
    • 最后,我假设您没有通过 Gremlin 控制台提交这些请求,并且您正在使用驱动程序将脚本提交到 Gremlin 服务器。如果是这样,您似乎没有对这些脚本进行参数化。请使用基于字节码的遍历或至少参数化您的脚本:tinkerpop.apache.org/docs/current/reference/…,这将对资源使用和性能产生非常大的影响。
    猜你喜欢
    • 1970-01-01
    • 2015-08-20
    • 2020-11-04
    • 2021-03-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多