【问题标题】:How to .dot in pyspark (AttributeError: 'DataFrame' object has no attribute 'dot')如何在 pyspark 中添加 .dot(AttributeError: \'DataFrame\' 对象没有属性 \'dot\')
【发布时间】:2022-08-09 10:10:41
【问题描述】:

在熊猫中,我们知道df1.dot(df2.T) 用于点积,但是当我在 pySpark 中运行时

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
<ipython-input-7-2219b97587ee> in <module>
----> 1 df1.dot(df2.T)

/opt/cloudera/parcels/CDH-7.1.3-1.cdh7.1.3.p0.4992530/lib/spark/python/pyspark/sql/dataframe.py in __getattr__(self, name)
   1302         if name not in self.columns:
   1303             raise AttributeError(
-> 1304                 \"\'%s\' object has no attribute \'%s\'\" % (self.__class__.__name__, name))
   1305         jc = self._jdf.apply(name)
   1306         return Column(jc)

AttributeError: \'DataFrame\' object has no attribute \'dot\'

    标签: python pandas pyspark


    【解决方案1】:

    你试过pandas-on-spark吗?

    import pyspark.pandas as ps
    ps.set_option('compute.ops_on_diff_frames', True)
    
    df = ps.DataFrame([[0, 1, -2, -1], [1, 1, 1, 1]])
    s = ps.Series([1, 1, 2, 1])
    
    print(df @ s)
    
    0   -4
    1    5
    dtype: int64
    
    

    笔记

    • pyspark.pandas 需要 pyspark &gt;= 3.2
    • 点积仅适用于 Dataframe (dot) Series 之间的操作。所以你不能用它来做类似的事情:df @ df
    • 要从spark RDD 转换为pandas-on-spark,您可以使用类似:pdf = df.to_pandas_on_spark()

    【讨论】:

      猜你喜欢
      • 2016-11-30
      • 2023-02-10
      • 1970-01-01
      • 1970-01-01
      • 2013-10-23
      • 2017-01-24
      • 2018-10-10
      • 2019-08-18
      • 2021-01-20
      相关资源
      最近更新 更多