【问题标题】:Unresolved reference while trying to import col from pyspark.sql.functions in python 3.5尝试从 python 3.5 中的 pyspark.sql.functions 导入 col 时未解析的引用
【发布时间】:2018-01-04 04:37:01
【问题描述】:

请参阅此处的帖子: Spark structured streaming with python 我想在 python 3.5 中导入“col”

from pyspark.sql.functions import col

但是我收到一个错误,提示未解决对 col 的引用。我已经安装了 pyspark 库,所以只是想知道 'col' 是否已从 pyspark 库中删除?那么我该如何导入'col'。

【问题讨论】:

标签: python apache-spark pyspark pyspark-sql spark-structured-streaming


【解决方案1】:

像col这样的函数不是python代码中定义的显式函数,而是动态生成的。

还会被pylint等静态分析工具报错

所以最简单的使用方法应该是这样的

from pyspark.sql import functions as F

F.col("colname")

python/pyspark/sql/functions.py中的如下代码

_functions = {
    'lit': _lit_doc,
    'col': 'Returns a :class:`Column` based on the given column name.',
    'column': 'Returns a :class:`Column` based on the given column name.',
    'asc': 'Returns a sort expression based on the ascending order of the given column name.',
    'desc': 'Returns a sort expression based on the descending order of the given column name.',

    'upper': 'Converts a string expression to upper case.',
    'lower': 'Converts a string expression to upper case.',
    'sqrt': 'Computes the square root of the specified float value.',
    'abs': 'Computes the absolute value.',

    'max': 'Aggregate function: returns the maximum value of the expression in a group.',
    'min': 'Aggregate function: returns the minimum value of the expression in a group.',
    'count': 'Aggregate function: returns the number of items in a group.',
    'sum': 'Aggregate function: returns the sum of all values in the expression.',
    'avg': 'Aggregate function: returns the average of the values in a group.',
    'mean': 'Aggregate function: returns the average of the values in a group.',
    'sumDistinct': 'Aggregate function: returns the sum of distinct values in the expression.',
}

def _create_function(name, doc=""):
    """ Create a function for aggregator by name"""
    def _(col):
        sc = SparkContext._active_spark_context
        jc = getattr(sc._jvm.functions, name)(col._jc if isinstance(col, Column) else col)
        return Column(jc)
    _.__name__ = name
    _.__doc__ = doc
    return _

for _name, _doc in _functions.items():
    globals()[_name] = since(1.3)(_create_function(_name, _doc))

【讨论】:

    【解决方案2】:

    尝试安装'pyspark-stubs',我在 PyCharm 中遇到了同样的问题,通过这样做我解决了。

    【讨论】:

    • conda install -c conda-forge pyspark-stubs 为我工作
    • 谢谢,它帮助我摆脱了这个烦人的问题。
    【解决方案3】:

    这似乎是 PyCharm 编辑器的问题,我也可以通过 Python 控制台使用 trim() 运行程序。

    【讨论】:

      【解决方案4】:

      原来是 IntelliJ IDEA 的问题。即使它显示未解析的引用,我的程序在命令行中仍然可以正常运行。

      【讨论】:

      • 这个answer很好地解释了这种行为。
      • 原因是对的,但没有提供实际的解决方案
      猜你喜欢
      • 2018-02-25
      • 2019-06-28
      • 2021-07-24
      • 2015-06-09
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-06-17
      • 2023-03-26
      相关资源
      最近更新 更多