【发布时间】:2019-01-08 14:55:34
【问题描述】:
我想使用 AWS 粘合作业从 Mysql 实例中读取过滤后的数据。由于胶水 jdbc 连接不允许我下推谓词,因此我试图在我的代码中显式创建 jdbc 连接。
我想使用如下所示的 jdbc 连接对 Mysql 数据库运行带有 where 子句的选择查询
import com.amazonaws.services.glue.GlueContext
import org.apache.spark.SparkContext
import org.apache.spark.sql.SparkSession
object TryMe {
def main(args: Array[String]): Unit = {
val sc: SparkContext = new SparkContext()
val glueContext: GlueContext = new GlueContext(sc)
val spark: SparkSession = glueContext.getSparkSession
// Read data into a DynamicFrame using the Data Catalog metadata
val t = glueContext.read.format("jdbc").option("url","jdbc:mysql://serverIP:port/database").option("user","username").option("password","password").option("dbtable","select * from table1 where 1=1").option("driver","com.mysql.jdbc.Driver").load()
}
}
失败并出现错误
com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException 你有一个 SQL 语法错误;检查与您对应的手册 MySQL 服务器版本,用于在 'select * from 附近使用正确的语法 table1 where 1=1 WHERE 1=0' at line 1
这不应该吗?如何在不将整个表读入数据框的情况下使用 JDBC 连接检索过滤后的数据?
【问题讨论】:
标签: aws-glue mssql-jdbc