【问题标题】:How to loop through elements from a python pandas dataframe to a new nested dictionary?如何将元素从 python pandas 数据帧循环到新的嵌套字典?
【发布时间】:2020-12-09 22:33:46
【问题描述】:

我目前正在使用 pandas 库从 CSV 文件中读取数据。数据包括一个由 1 和 0 组成的“数据”列,以及一个具有唯一时间和日期戳的“published_at”列(我已将其转换为数据帧的索引)。 Click here to see picture of the Dataframe from CSV(我删除了core_id数据,因为它无关紧要)。

在数据中,“1”表示是,“0”表示否。我想通过从某个开始日期到结束日期(即 2020 年 11 月 26 日到 2020 年 11 月 27 日)循环遍历数据帧来分析数据,并计算“1”(yes_data)出现的次数,以及如何每天多次出现“0”(no_data)。从那里,我想创建一个包含该数据的新 CSV 文件或数据框,以便我可以从那里进行分析。

我尝试解决此问题的方法是创建一个嵌套字典,并尝试通过循环遍历主数据框并计算每天出现“是”和“否”的次数来填充它。

我希望得到一个包含 3 列的字典(或数据框、csv 文件等):日期(即 2020-11-26)、“是”计数和“否”计数。

下面是我想出的代码:

yes_data = 0
no_data = 0
date_id = '2020-11-26'

# Create a dictionary to populate a new dataframe
#new_data = {
#  date_id: {'yes': yes_data, 'no': no_data},
#  "2020-11-27": {'yes': 2, 'no': 2}}
new_data = {"":{}}

# I tried to convert the csv data to a dictionary but I don't know
# if this is necessary so I commented it out
# csv_dict = csv_data.to_dict

csv_dict = csv_data

for day in csv_dict['2020-11-26':'2020-11-27']:
    new_data[day] = csv_dict[day]
    for state in csv_dict['2020-11-26':'2020-11-27'].data:
        if state == 1:
            yes_data += 1
            new_data[day][state] == yes_data
        elif state == 0:
            no_data += 1
            new_data[day][state] == no_data

但是代码不起作用(我到处都遇到错误......)。我怎样才能修复它来做我想做的事情?任何帮助表示赞赏。谢谢!

附:我对 Python 还很陌生,在这里尽我所能!

【问题讨论】:

  • 您遇到的错误是什么?你能提供一个最小的例子吗? (另外,尽量避免数据帧上的循环。如果要计算系列中的值,请使用.value_counts()

标签: python python-3.x pandas dataframe for-loop


【解决方案1】:

希望你一切顺利。 这个 sn-p 将帮助您完成这项工作!

result = {}
for index, row in df.iterrows(): # Iterates over the row
    date = row['published_at'].split(' ')[0]  # This line takes only the date of the row ( not hour and minute ...)
    ans = row['data']  # Finds if the data is zero or one
    if date not in result:  # Creates an entery for this date if it hasn't created yet
        result[date] = {'yes':0, 'no':0}
    if ans: # Increases the number of yes if ans == 1
        result[date]['yes']+=1 
    else:  # Increases the number of yes if ans == 0
        result[date]['no'] +=1

请记住,df 是您的数据框,而 row['data'] 表示 0 或 1。因此,如果您有不同的名称,请更改它。 在这段代码的末尾,您将拥有一个结构如下所述的字典。

result = {'2020-12-01': {'yes': 2, 'no': 4}, '2020-12-4': {'yes':n, 'no':m} }

祝你有美好的一天

【讨论】:

  • 非常感谢您!!!我通过创建一个新列 csv_data['date'] 并复制 csv_data.index.date 对其进行了调整以适应我的需要。这样做允许我从日期变量中删除 .split(' ')[0] ,这由于某种原因给了我错误。其余的工作就像一个魅力!再次感谢
  • @encrypted_shadow。别客气!快乐编码:D
猜你喜欢
  • 2021-03-20
  • 2021-11-06
  • 2023-01-25
  • 2019-10-03
  • 2020-12-21
  • 2020-10-09
  • 2018-07-13
  • 2021-05-05
  • 2023-01-22
相关资源
最近更新 更多