【问题标题】:Python - For loop through dictionary with value as findall regular expressionPython - For 循环遍历字典,值为 findall 正则表达式
【发布时间】:2018-04-26 21:33:54
【问题描述】:

我有一个 csv 文件,其中包含每年的天气信息。我创建了一个字典,其中键是年份,值是正则表达式,用于收集一年内的所有日期,就像这样

import csv, re

with open('weather_data.csv') as csvfile:
    readCSV = csv.reader(csvfile, delimiter=',')

csvfile = csvfile.read()

years = {'year_92': re.findall(r'\d+/\d+/1992', csvfile), 'year_93': re.findall(r'\d+/\d+/1993', csvfile),
     'year_94': re.findall(r'\d+/\d+/1994', csvfile), 'year_95': re.findall(r'\d+/\d+/1995', csvfile)}

CSV 的第一列是 mm/dd/yyyy 格式的日期,第二列是温度。我想做的是用最好的方法把一年内的所有温度都拿下来,然后找到它们的平均值。

目前,我正在尝试遍历字典以将 1992 年的每个温度附加(即)到一个列表中,以便我可以平均该列表。

temps = csvfile[1]
temp_92 = []
 for line in years.items(), temps:
     temp_92.append(line)
     print(temp_92)

但是,这显然有问题。代码确实运行,除了它返回 mm/dd/yy。我尝试切换 csvfile[] 但没有结果。

这是我的输出的样子

[dict_items([('year_95', ['1/1/1995', '1/2/1995', '1/3/1995', '1/4/1995', '1/5/1995', '1/6/1995', '1/7/1995', '1/8/1995', '1/9/1995', '1/10/1995', '1/11/1995', '1/12/1995', '1/13/1995', '1/14/1995', '1/15/1995', '1/16/1995', '1/17/1995', '1/18/1995', '1/19/1995', '1/20/1995', '1/21/1995', '1/22/1995', '1/23/1995', '1/24/1995', '1/25/1995', '1/26/1995', '1/27/1995', '1/28/1995', '1/30/1995', '1/31/1995' ...and so on

编辑:这是每个请求的 CSV 中的一些示例数据!尽量格式化。

A1:日期 B1:温度
A2:10/1/1992 B2:53
A3:1992 年 10 月 2 日 B3:58
A4:1992 年 10 月 3 日 B4强>: 62

【问题讨论】:

  • 你能在你的 csv 文件中包含几行吗?
  • 添加 CSV 行
  • 我已经更新了我的答案以反映你的 csv

标签: python regex csv dictionary for-loop


【解决方案1】:

我认为您的正则表达式和 csv 阅读器对此过于矫枉过正。我对您的输入格式做了一个假设,但以下使用所有标准库。

输入文件示例

A1: Date B1: Temp
A2: 10/1/1992 B2: 53
A3: 10/2/1992 B3: 58
A4: 10/3/1992 B4: 62

每年平均计算器示例:

from itertools import groupby
from functools import reduce

csv_file = '/tmp/temp_example.txt'

with open(csv_file, 'r') as csv_fh:
    next(csv_fh)  # skip header
    split_lines = [line.strip('\n').split(' ')[1::2] for line in csv_fh]
    split_lines.sort(key=lambda x: x[0])  # sort by date first
    # group by year
    year_temps = {}
    for key, group in groupby(split_lines, lambda x: x[0].split('/')[-1]):
        year_temps[key] = [int(row[1]) for row in group]
    yearly_averages = {year: (reduce(lambda x, y: x + y, temps) / len(temps))
                       for year, temps in year_temps.items()}
    print(yearly_averages)

这里的策略是在你的行中读取的,并将它们分成各自的列,然后使用 groupby 创建你的年份字典:[temps]。可以根据您最喜欢的方式计算平均值。

您的数据一开始就非常结构化,实际上不需要正则表达式的复杂性来提取年份。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2015-10-30
    • 1970-01-01
    • 2013-11-29
    • 2013-02-20
    • 2011-03-18
    • 2023-03-18
    相关资源
    最近更新 更多