【问题标题】:Error when converting from spark dataframe with dates to pandas dataframe从带有日期的 spark 数据框转换为 pandas 数据框时出错
【发布时间】:2019-07-21 04:38:57
【问题描述】:

我有一个具有此架构的 spark 数据框:

root
 |-- product_id: integer (nullable = true)
 |-- stock: integer (nullable = true)
 |-- start_date: date (nullable = true)
 |-- end_date: date (nullable = true)

当尝试将其传递给 pandas_udf 或转换为 pandas 数据框时:

pandas_df = spark_df.toPandas()

它返回此错误:

AttributeError        Traceback (most recent call last)
<ipython-input-86-4bccc6e8422d> in <module>()
     10 # spark_df.printSchema()
     11 
---> 12 pandas_df = spark_df.toPandas()

/home/.../lib/python2.7/site-packages/pyspark/sql/dataframe.pyc in toPandas(self)
   2123                         table = pyarrow.Table.from_batches(batches)
   2124                         pdf = table.to_pandas()
-> 2125                         pdf = _check_dataframe_convert_date(pdf, self.schema)
   2126                         return _check_dataframe_localize_timestamps(pdf, timezone)
   2127                     else:

/home.../lib/python2.7/site-packages/pyspark/sql/types.pyc in _check_dataframe_convert_date(pdf, schema)
   1705     """
   1706     for field in schema:
-> 1707         pdf[field.name] = _check_series_convert_date(pdf[field.name], field.dataType)
   1708     return pdf
   1709 

/home/.../lib/python2.7/site-packages/pyspark/sql/types.pyc in _check_series_convert_date(series, data_type)
   1690     """
   1691     if type(data_type) == DateType:
-> 1692         return series.dt.date
   1693     else:
   1694         return series

/home/.../lib/python2.7/site-packages/pandas/core/generic.pyc in __getattr__(self, name)
   5061         if (name in self._internal_names_set or name in self._metadata or
   5062                 name in self._accessors):
-> 5063             return object.__getattribute__(self, name)
   5064         else:
   5065             if self._info_axis._can_hold_identifiers_and_holds_name(name):

/home/.../lib/python2.7/site-packages/pandas/core/accessor.pyc in __get__(self, obj, cls)
    169             # we're accessing the attribute of the class, i.e., Dataset.geo
    170             return self._accessor
--> 171         accessor_obj = self._accessor(obj)
    172         # Replace the property with the accessor object. Inspired by:
    173         # http://www.pydanny.com/cached-property.html

/home/.../lib/python2.7/site-packages/pandas/core/indexes/accessors.pyc in __new__(cls, data)
    322             pass  # we raise an attribute error anyway
    323 
--> 324         raise AttributeError("Can only use .dt accessor with datetimelike "
    325                              "values")

AttributeError: Can only use .dt accessor with datetimelike values

如果从 spark 数据框中删除日期字段,则转换不会出现问题。

我检查了数据不包含任何空值,但如果知道如何处理这些也很好。

我正在使用python2.7:

  • pyspark==2.4.0
  • pyarrow==0.12.1
  • 熊猫==0.24.1

【问题讨论】:

  • 如果可以,请尝试将日期文件转换为 DateTypeTimestamp
  • @VictorValente 你所说的 DateType 是什么意思?架构中显示的字段不是已经在这种类型中了吗?

标签: pandas apache-spark dataframe pyspark


【解决方案1】:

看起来像一个错误。 pyarrow==0.12.1 和 pyarrow==0.12.0 也有同样的问题。将 spark 数据框列转换为 TIMESTAMP 对我有用。

spark.sql('SELECT CAST(date_column as TIMESTAMP) FROM foo')

还回滚到 pyarrow==0.11.0 解决了这个问题。 (我的python是3.7.1和pandas 0.24.2)

【讨论】:

  • 仍然发生在 pyarrow==0.14.0 和 spark 2.4.0
  • pyspark 2.4.3 和 pyarrow 0.17.0 存在同样的问题
【解决方案2】:

根据他们在 Spark 3 中修复的Jira。作为一种解决方法,您可以考虑将日期列转换为时间戳(这更符合 pandas 日期时间类型)。

import pyspark.sql.functions as func
df = df.select(func.to_timestamp(func.col('session_date'), 'yyyy-MM-dd').alias('session_date')
df.toPandas()

在 Pyspark 2.4.4 中测试

【讨论】:

    【解决方案3】:

    这对我有用:

    import pyspark.sql.functions as f
    
    spark_df = spark_df.withColumn('start_date', f.to_timestamp(f.col('start_date')))
    spark_df = spark_df.withColumn('end_date',   f.to_timestamp(f.col('end_date')))
    pandas_df = spark_df.toPandas()
    

    【讨论】:

      猜你喜欢
      • 2020-09-20
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-12-25
      • 1970-01-01
      • 2018-02-09
      相关资源
      最近更新 更多