【问题标题】:How to create a min-max plot by month with fill_between如何使用fill_between按月创建最小-最大图
【发布时间】:2021-10-12 08:03:22
【问题描述】:

我必须将月份名称显示为 xticks,当我绘制图形并将 x 作为月份名称传递时,它绘制错误。 我还必须在折线图上叠加一个散点图。

我无法在此处粘贴完整代码,因为这是一个 MOOC 作业,我只是在寻找我在这里做错了什么。

plt.figure(figsize=(8,5))

plt.plot(mint['Mean'],linewidth= 1, label = 'Minumum')
plt.plot(maxt['Mean'],linewidth = 1, label = 'Maximum')

plt.scatter(broken_low,mint15.iloc[broken_low]['Mean'],alpha = 0.75)
plt.scatter(broken_high,maxt15.iloc[broken_high]['Mean'],alpha = .75)

数据集链接在这里:

https://drive.google.com/file/d/1qJnnHDK_0ghmHQl4OuyKDr-0K5ETo7Td/view?usp=sharing

它应该看起来像这样,填充线之间的区域和 x 轴为月份,y 轴为摄氏度

【问题讨论】:

    标签: python pandas matplotlib


    【解决方案1】:

    使用来自 OP 的数据进行更新

    • 第一种方法的问题在于它要求 x 轴为日期时间格式。
    • 正在对问题中的数据进行分组并根据字符串绘制,该字符串是月份和日期的组合
    • x 轴代表 365 天,闰年已被删除。
      • 在每个月初的适当位置放置刻度
      • 为刻度添加标签
    • 这似乎来自coursera: Applied Data Science with Python Specialization
    import pandas as pd
    import matplotlib.pyplot as plot
    import calendar
    
    # load the data
    url = 'https://raw.githubusercontent.com/trenton3983/stack_overflow/master/data/so_data/2020-07-17_62929123/temperature.csv'
    df = pd.read_csv(url, parse_dates=['Date'])
    
    # remove leap day
    df = df[~((df.Date.dt.month == 2) & (df.Date.dt.day == 29))]
    
    # add a year column
    df['Year'] = df.Date.dt.year
    
    # add a month-day column to use for groupby
    df['Month-Day'] = df.Date.dt.month.astype('str') + '-' + df.Date.dt.day.astype('str')
    
    # select 2015 data
    df_15 = df[df.Year == 2015].copy().reset_index()
    
    # select data before 2015
    df_14 = df[df.Year < 2015].copy().reset_index()
    
    # filter data to either max or min and groupby month-day
    max_14 = df_14[df_14.Element == 'TMAX'].groupby(['Month-Day']).agg({'Data_Value': max}).reset_index().rename(columns={'Data_Value': 'Daily_Max'})
    min_14 = df_14[df_14.Element == 'TMIN'].groupby(['Month-Day']).agg({'Data_Value': min}).reset_index().rename(columns={'Data_Value': 'Daily_Min'})
    max_15 = df_15[df_15.Element == 'TMAX'].groupby(['Month-Day']).agg({'Data_Value': max}).reset_index().rename(columns={'Data_Value': 'Daily_Max'})
    min_15 = df_15[df_15.Element == 'TMIN'].groupby(['Month-Day']).agg({'Data_Value': max}).reset_index().rename(columns={'Data_Value': 'Daily_Min'})
    
    # select max values from 2015 that are greater than the recorded max
    higher_14 = max_15[max_15 > max_14]
    
    # select min values from 2015 that are less than the recorded min
    lower_14 = min_15[min_15 < min_14]
    
    # plot the min and max lines
    ax = max_14.plot(label='Max Recorded', color='tab:red', figsize=(12, 8))
    min_14.plot(ax=ax, label='Min Recorded', color='tab:blue')
    
    # add the fill, between min and max
    plt.fill_between(max_14.index, max_14.Daily_Max, min_14.Daily_Min, alpha=0.10, color='tab:orange')
    
    # add points greater than max or less than min
    plt.scatter(higher_14.index, higher_14.Daily_Max, label='2015 Max > Record', alpha=0.75, color='tab:red')
    plt.scatter(lower_14.index, lower_14.Daily_Min, label='2015 Min < Record', alpha=0.75, color='tab:blue')
    
    # set plot xlim
    plt.xlim(-5, 370)
    
    # tick locations
    ticks=[-5, 0, 31, 59, 90, 120, 151, 181, 212, 243, 273, 304, 334, 365, 370]
    
    # tick labels
    labels = list(calendar.month_abbr)  # list of months
    labels.extend(['Jan', ''])
    
    # add the custom ticks and labels
    plt.xticks(ticks=ticks, labels=labels)
    
    # plot cosmetics
    plt.legend()
    plt.xlabel('Day of Year: 0-365 Displaying Start of Month')
    plt.ylabel('Temperature °C')
    plt.title('Daily Max and Min: 2009 - 2014\nRecorded Max < 2015 Temperatures < Recorded Min')
    plt.tight_layout()
    plt.show()
    

    原答案

    import pandas as pd
    import matplotlib.pyplot as plt
    import matplotlib.dates as mdates
    
    # plot styling parameters
    plt.style.use('seaborn')
    plt.rcParams['figure.figsize'] = (16.0, 10.0)
    plt.rcParams["patch.force_edgecolor"] = True
    
    # locate the Month and format the label
    months = mdates.MonthLocator()  # every month
    months_fmt = mdates.DateFormatter('%b')
    
    # plot the data
    fig, ax = plt.subplots()
    ax.plot(max_15.index, 'rolling', data=max_15, label='max rolling mean')
    ax.scatter(x=max_15.index, y='v', data=max_15, alpha=0.75, label='MAX')
    
    ax.plot(min_15.index, 'rolling', data=min_15, label='min rolling mean')
    ax.scatter(x=min_15.index, y='v', data=min_15, alpha=0.75, label='MIN')
    ax.legend()
    
    # format the ticks
    ax.xaxis.set_major_locator(months)
    ax.xaxis.set_major_formatter(months_fmt)
    

    可重现的数据

    import pandas as pd
    
    # download data into dataframe, it's in a wide format
    pdx_19 = pd.read_csv('http://www.weather.gov/source/pqr/climate/webdata/Portland_dailyclimatedata.csv', header=6)
    
    # clean and label data
    pdx_19.drop(columns=['AVG or Total'], inplace=True)
    pdx_19.columns = list(pdx_19.columns[:3]) + [f'v_{day}' for day in pdx_19.columns[3:]]
    pdx_19.rename(columns={'Unnamed: 2': 'TYPE'}, inplace=True)
    pdx_19 = pdx_19[pdx_19.TYPE.isin(['TX', 'TN', 'PR'])]
    
    # convert to long format
    pdx = pd.wide_to_long(pdx_19, stubnames='v', sep='_', i=['YR', 'MO', 'TYPE'], j='day').reset_index()
    
    # additional cleaning
    pdx.TYPE = pdx.TYPE.map({'TX': 'MAX', 'TN': 'MIN', 'PR': 'PRE'})
    pdx.rename(columns={'YR': 'year', 'MO': 'month'}, inplace=True)
    pdx = pdx[pdx.v != '-'].copy()
    pdx['date'] = pd.to_datetime(pdx[['year', 'month', 'day']])
    pdx.drop(columns=['year', 'month', 'day'], inplace=True)
    pdx.v.replace({'M': np.nan, 'T': np.nan}, inplace=True)
    pdx.v = pdx.v.astype('float')
    
    # select on 2015
    pdx_2015 = pdx[pdx.date.dt.year == 2015].reset_index(drop=True).set_index('date')
    
    # select only MAX temps
    max_15 = pdx_2015[pdx_2015.TYPE == 'MAX'].copy()
    
    # select only MIN temps
    min_15 = pdx_2015[pdx_2015.TYPE == 'MIN'].copy()
    
    # calculate rolling mean
    max_15['rolling'] = max_15.v.rolling(7).mean()
    min_15['rolling'] = min_15.v.rolling(7).mean()
    

    max_15

               TYPE     v    rolling
    date                            
    2015-01-01  MAX  39.0        NaN
    2015-01-02  MAX  41.0        NaN
    2015-01-03  MAX  41.0        NaN
    2015-01-04  MAX  53.0        NaN
    2015-01-05  MAX  57.0        NaN
    2015-01-06  MAX  47.0        NaN
    2015-01-07  MAX  51.0  47.000000
    2015-01-08  MAX  45.0  47.857143
    2015-01-09  MAX  50.0  49.142857
    2015-01-10  MAX  42.0  49.285714
    

    min_15

               TYPE     v    rolling
    date                            
    2015-01-01  MIN  24.0        NaN
    2015-01-02  MIN  26.0        NaN
    2015-01-03  MIN  35.0        NaN
    2015-01-04  MIN  38.0        NaN
    2015-01-05  MIN  42.0        NaN
    2015-01-06  MIN  38.0        NaN
    2015-01-07  MIN  34.0  33.857143
    2015-01-08  MIN  35.0  35.428571
    2015-01-09  MIN  37.0  37.000000
    2015-01-10  MIN  36.0  37.142857
    

    【讨论】:

    • 感谢您的帮助。您的代码非常清晰,因此易于理解。我正在使用字符串数据并对该数据进行数据处理,而不是首先将其正确转换为to_datetime。如果您知道我做错了什么,请指出,再次感谢。
    猜你喜欢
    • 2021-01-11
    • 2020-09-15
    • 1970-01-01
    • 2012-05-26
    • 1970-01-01
    • 1970-01-01
    • 2014-06-14
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多