【问题标题】:How to read CSV in pyspark with "," delimiter but not ", "如何使用 \",\" 分隔符而不是 \", \" 在 pyspark 中读取 CSV
【发布时间】:2022-10-06 05:07:30
【问题描述】:

我正在使用以下代码读取 PySpark 中的 CSV 文件

cb_sdf = sqlContext.read.format(\"csv\") \\
                        .options(header=\'true\', 
                                 multiLine = \'True\', 
                                 inferschema=\'true\', 
                                 treatEmptyValuesAsNulls=\'true\') \\
                        .load(cb_file)

行数是正确的。但是对于某些行,列的分隔不正确。我认为是因为当前的分隔符是\",\",但是有些单元格在文本中也包含\", \"。

例如,pandas 数据框中的以下行(我使用 pd.read_csv 进行调试)

name industry country 111 package/freight delivery russia
name industry country 111 tourism\"\"\" package/freight delivery

如何将分隔符设置为完全 \",\" 而没有任何空格?

更新:

我检查了 CSV 文件,原行是:

111,\"cjsc \"\"transport, customs, tourism\"\"\",ttt-w.ru,package/freight delivery,\"vyborg, leningrad, russia\",russia,1 - 10

那么它仍然是分隔符的问题,还是引号的问题?

  • 请将示例数据发布为文本,而不是图像;见How to Ask。如果 csv 中的字段包含逗号,则该字段需要用引号引起来。如果您的 csv 字段没有被引用,请与损坏输出的生产者联系。
  • trimming那些专栏看完后怎么样?

标签: python pyspark delimiter


【解决方案1】:

我认为分离我们将有:

col1:111 col2: "cjsc ""运输、海关、旅游""" col3: ttt-w.ru,包裹/货运 col4: "维堡,列宁格勒,俄罗斯" col5:俄罗斯 col6: 1 - 10

【讨论】:

  • 使用 cb_sdf = sqlContext.read.format("csv") \ .options(header='true', sep=',', multiLine = 'True', inferschema='true',treatEmptyValuesAsNulls='true') \ .load (cb_file)
猜你喜欢
  • 2021-12-15
  • 2017-04-12
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2016-03-19
  • 1970-01-01
  • 2020-05-06
相关资源
最近更新 更多