【问题标题】:Execute queries on redshift using Pyspark使用 Pyspark 在 redshift 上执行查询
【发布时间】:2021-07-29 05:36:10
【问题描述】:

你们中的任何人都可以建议使用 pyspark 对 redshift 表执行查询的方法吗?

【问题讨论】:

标签: pyspark amazon-redshift amazon-connect


【解决方案1】:

使用 pyspark 数据帧选项是读取/写入 Redshift 表的一种选择。在数据框的帮助下执行查询仅限于使用preactions/postactions 方法。
如果需要执行多个查询,一种方法是使用psycopg2 模块。 首先,您需要在您的服务器上安装 psycopg2 模块:
sudo python -m pip install psycopg2
然后打开 pyspark shell 并执行以下命令:

import psycopg2

conn={'db_name': {'hostname':'redshift_host_url','database':'database_on_redshift','username':'redshift_username','password':'p', 'port':your_port} }

db='db_name'

hostname, username, password, database, portnumber = conn[db]['hostname'], conn[db]['username'], conn[db]['password'], conn[db]['database'], conn[db]['port']

con = psycopg2.connect( host=hostname, port=portnumber, user=username, password=password, dbname=database)

query="insert into sample_table select * from table1"
cur = con.cursor()
rs = cur.execute(query)
con.commit()
con.close()

您也可以参考:https://www.psycopg.org/docs/usage.html

【讨论】:

    猜你喜欢
    • 2017-04-07
    • 2022-08-03
    • 1970-01-01
    • 2020-11-12
    • 1970-01-01
    • 1970-01-01
    • 2017-01-12
    • 1970-01-01
    • 2018-07-10
    相关资源
    最近更新 更多