【问题标题】:How to merge columns into one on top of each other in pyspark?如何在pyspark中将列合并为一个?
【发布时间】:2021-12-29 19:51:03
【问题描述】:

我有一个看起来像这样的 pyspark 数据框,

data = [("James","Joyce"),
    ("Michael","Doglus"),
    ("Robert","Connings"),
    ("Maria","XYZ"),
    ("Jen","PQR")
  ]

df2 = spark.createDataFrame(data, ["Name", "Lots_of_names"])
df2


    Name    Lots_of_names
0   James   Joyce
1   Michael     Doglus
2   Robert  Connings
3   Maria   XYZ
4   Jen     PQR

我想将两列合并为一个长列(可能在一个新的数据框中),它将有 10 行。有没有办法到达那里?提前致谢。

【问题讨论】:

    标签: dataframe pyspark concatenation


    【解决方案1】:

    你可能想要做这样的事情

    import pyspark.sql.functions as F
    
    df_out = df2.select(F.explode(F.array("Name", "Lots_of_names")).alias("one_col"))
    

    产生 df_out 如下

    # one_col
    #------
    # James
    # Joyce
    # Michael
    # Doglus
    # Robert
    # Connings
    # Maria
    # XYZ
    # Jen
    # PQR
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2020-02-21
      • 2020-06-11
      • 1970-01-01
      • 1970-01-01
      • 2017-08-05
      • 1970-01-01
      • 2013-01-23
      相关资源
      最近更新 更多