【问题标题】:How can I use graphframes with pyspark on AWS EMR?如何在 AWS EMR 上使用带有 pyspark 的图框?
【发布时间】:2019-10-20 03:00:10
【问题描述】:

我正在尝试在 AWS EMR 上的 Jupyter Notebook(使用 Sagemaker 和 sparkmagic)的 pyspark 中使用 graphframes 包。我尝试在 AWS 控制台中创建 EMR 集群时添加配置选项:

[{"classification":"spark-defaults", "properties":{"spark.jars.packages":"graphframes:graphframes:0.7.0-spark2.4-s_2.11"}, "configurations":[]}]

但是当我尝试在 jupyter notebook 的 pyspark 代码中使用 graphframes 包时,我仍然遇到错误。

这是我的代码(来自 graphframes 示例):

# Create a Vertex DataFrame with unique ID column "id"
v = spark.createDataFrame([
  ("a", "Alice", 34),
  ("b", "Bob", 36),
  ("c", "Charlie", 30),
], ["id", "name", "age"])
# Create an Edge DataFrame with "src" and "dst" columns
e = spark.createDataFrame([
  ("a", "b", "friend"),
  ("b", "c", "follow"),
  ("c", "b", "follow"),
], ["src", "dst", "relationship"])
# Create a GraphFrame
from graphframes import *
g = GraphFrame(v, e)

# Query: Get in-degree of each vertex.
g.inDegrees.show()

# Query: Count the number of "follow" connections in the graph.
g.edges.filter("relationship = 'follow'").count()

# Run PageRank algorithm, and show results.
results = g.pageRank(resetProbability=0.01, maxIter=20)
results.vertices.select("id", "pagerank").show()

这是输出/错误:

ImportError: No module named graphframes

我通读了this git thread,但所有潜在的解决方法似乎都非常复杂,需要通过 ssh 连接到 EMR 集群的主节点。

【问题讨论】:

    标签: apache-spark pyspark jupyter-notebook amazon-emr graphframes


    【解决方案1】:

    我终于知道有一个PyPi package for graphframes。我用它来创建一个引导操作,详细说明 here,尽管我做了一些改动。

    这是我为使图形框架在 EMR 上工作所做的工作:

    1. 首先,我创建了一个 shell 脚本并将其保存为 s3,命名为“install_jupyter_libraries_emr.sh”:
    #!/bin/bash
    
    sudo pip install graphframes
    
    1. 然后,我在 AWS 控制台中完成了高级选项 EMR 创建过程。
      • 在第 1 步中,我在编辑软件设置文本框中添加了 graphframes 包的 maven 坐标:
      [{"classification":"spark-defaults","properties":{"spark.jars.packages":"graphframes:graphframes:0.7.0-spark2.4-s_2.11"}}]
      
      • 在第 3 步:常规集群设置中,我进入了引导操作部分
      • 在引导操作部分中,我添加了一个新的自定义引导操作:
        • 任意名称
        • 我的“install_jupyter_libraries_emr.sh”脚本的 s3 位置
        • 没有可选参数
      • 然后我开始创建集群
    2. 集群启动后,我进入 Jupyter 并运行我的代码:
    # Create a Vertex DataFrame with unique ID column "id"
    v = spark.createDataFrame([
      ("a", "Alice", 34),
      ("b", "Bob", 36),
      ("c", "Charlie", 30),
    ], ["id", "name", "age"])
    # Create an Edge DataFrame with "src" and "dst" columns
    e = spark.createDataFrame([
      ("a", "b", "friend"),
      ("b", "c", "follow"),
      ("c", "b", "follow"),
    ], ["src", "dst", "relationship"])
    # Create a GraphFrame
    from graphframes import *
    g = GraphFrame(v, e)
    
    # Query: Get in-degree of each vertex.
    g.inDegrees.show()
    
    # Query: Count the number of "follow" connections in the graph.
    g.edges.filter("relationship = 'follow'").count()
    
    # Run PageRank algorithm, and show results.
    results = g.pageRank(resetProbability=0.01, maxIter=20)
    results.vertices.select("id", "pagerank").show()
    

    这一次,我终于得到了正确的输出:

    +---+--------+
    | id|inDegree|
    +---+--------+
    |  c|       1|
    |  b|       2|
    +---+--------+
    
    +---+------------------+
    | id|          pagerank|
    +---+------------------+
    |  b|1.0905890109440908|
    |  a|              0.01|
    |  c|1.8994109890559092|
    +---+------------------+
    

    【讨论】:

    • 精彩的答案,我很感激你回来了你的解决方案。如果由我决定,我会给你所有虚假的互联网积分。非常感谢
    • 使用最新的 AWS EMR 集群,我必须使用“sudo pip-3.6 install graphframes”才能使其工作(而不是简单的 pip)
    • 真的需要第 1 步吗? graphframes documentation 表示“我们使用 --packages 参数自动下载 graphframes 包和任何依赖项。”
    • @panc 可能不再需要了,我没有测试过。当时,如果我只是做了 spark 包,而不是 pip install,当我尝试在 jupyter notebook 中导入 graphframes 时,我会收到一个错误,它找不到 graphframes 包。
    【解决方案2】:

    @Bob Swain 的回答很好,但现在图形框架的存储库位于 https://repos.spark-packages.org/。因此,为了使其正常工作,分类应更改为:

    [
     {
      "classification":"spark-defaults",
      "properties":{
        "spark.jars.packages":"graphframes:graphframes:0.8.0-spark2.4-s_2.11",
        "spark.jars.repositories":"https://repos.spark-packages.org/"
      }
     }
    ]
    

    【讨论】:

    • 我正在使用 sagemaker 的 sparkmagic 内核 (pyspark) 和带有 Livy 的 EMR 来运行 pyspark 代码。通过在 Sparkmagic 的 %%config 中添加这两个配置,我可以成功运行 pyspark graphframes 示例。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2021-10-29
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-04-15
    相关资源
    最近更新 更多