【发布时间】:2018-10-08 12:43:48
【问题描述】:
我正在使用 PySpark 并加载一个 csv 文件。我有一列包含欧洲格式的数字,这意味着逗号替换了点,反之亦然。
例如:我有2.416,67 而不是2,416.67。
My data in .csv file looks like this -
ID; Revenue
21; 2.645,45
23; 31.147,05
.
.
55; 1.009,11
在 pandas 中,通过在 pd.read_csv() 中指定 decimal=',' 和 thousands='.' 选项以读取欧洲格式,可以轻松读取此类文件。
熊猫代码:
import pandas as pd
df=pd.read_csv("filepath/revenues.csv",sep=';',decimal=',',thousands='.')
我不知道如何在 PySpark 中做到这一点。
PySpark 代码:
from pyspark.sql.types import StructType, StructField, FloatType, StringType
schema = StructType([
StructField("ID", StringType(), True),
StructField("Revenue", FloatType(), True)
])
df=spark.read.csv("filepath/revenues.csv",sep=';',encoding='UTF-8', schema=schema, header=True)
谁能建议我们如何使用上面提到的.csv() 函数在 PySpark 中加载这样的文件?
【问题讨论】:
-
为什么要指定分号分隔符?您提供的示例文件看起来是以空格分隔的,可能带有制表符?
-
这只是一个例子。好的,让我改一下。