【问题标题】:Creating multiple pyspark dataframes from a single dataframe [duplicate]从单个数据帧创建多个 pyspark 数据帧 [重复]
【发布时间】:2023-03-24 00:15:02
【问题描述】:

我需要根据 python 列表中可用的值在 pyspark 中动态创建多个数据帧

我的数据框(df)有数据:

date        gender balance
2018-01-01   M     100
2018-02-01   F     100
2018-03-01   M     100

my_list = [2018-01-01, 2018-02-01, 2018-03-01]
for i in my_list:
  df_i = df.select("*").filter("date=i").limit(1000)

你能帮忙吗?

【问题讨论】:

  • 我想根据日期将现有的 pyspark 数据帧拆分为 3 个数据帧,因此我创建了一个具有不同值的列表,我尝试通过传递日期使用过滤器函数进行过滤。
  • 是的,我现在明白了。对不起这是我的错。 Mohan,list 是关键字,不应用作变量名。我按照规定进行了编辑。

标签: python pandas apache-spark pyspark


【解决方案1】:

我不确定您是否可以在PySpark 中动态创建数据框的名称。在 Python 中,你甚至不能 dynamically 分配变量的名称,更不用说 dataframes。

一种方法是创建dataframes 的字典,其中key 对应于每个date,而该字典的value 对应于数据框。

对于 Python: 请参阅此 link,其中有人问过关于名称动态的类似问题。

这是一个小的PySpark 实现 -

from pyspark.sql.functions import col
values = [('2018-01-01','M',100),('2018-02-01','F',100),('2018-03-01','M',100)]
df = sqlContext.createDataFrame(values,['date','gender','balance'])
df.show()
+----------+------+-------+
|      date|gender|balance|
+----------+------+-------+
|2018-01-01|     M|    100|
|2018-02-01|     F|    100|
|2018-03-01|     M|    100|
+----------+------+-------+

# Creating a dictionary to store the dataframes.
# Key: It contains the date from my_list.
# Value: Contains the corresponding dataframe.
dictionary_df = {}  

my_list = ['2018-01-01', '2018-02-01', '2018-03-01']
for i in my_list:
    dictionary_df[i] = df.filter(col('date')==i)

for i in my_list:
    print('DF: '+i)
    dictionary_df[i].show() 

DF: 2018-01-01
+----------+------+-------+
|      date|gender|balance|
+----------+------+-------+
|2018-01-01|     M|    100|
+----------+------+-------+

DF: 2018-02-01
+----------+------+-------+
|      date|gender|balance|
+----------+------+-------+
|2018-02-01|     F|    100|
+----------+------+-------+

DF: 2018-03-01
+----------+------+-------+
|      date|gender|balance|
+----------+------+-------+
|2018-03-01|     M|    100|
+----------+------+-------+

print(dictionary_df)
    {'2018-01-01': DataFrame[date: string, gender: string, balance: bigint], '2018-02-01': DataFrame[date: string, gender: string, balance: bigint], '2018-03-01': DataFrame[date: string, gender: string, balance: bigint]}

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-08-12
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-06-13
    相关资源
    最近更新 更多