【问题标题】:resampling with origin='end_day'使用 origin='end_day' 重新采样
【发布时间】:2022-01-21 23:29:30
【问题描述】:

我不明白origin='end_day' 做了什么。

docs 举个例子:

>>> start, end = '2000-10-01 23:30:00', '2000-10-02 00:30:00'
>>> rng = pd.date_range(start, end, freq='7min')
>>> ts = pd.Series(np.arange(len(rng)) * 3, index=rng)
>>> ts 
2000-10-01 23:30:00     0
2000-10-01 23:37:00     3
2000-10-01 23:44:00     6
2000-10-01 23:51:00     9
2000-10-01 23:58:00    12
2000-10-02 00:05:00    15
2000-10-02 00:12:00    18
2000-10-02 00:19:00    21
2000-10-02 00:26:00    24
Freq: 7T, dtype: int32
>>> ts.resample('17min', origin='end_day').sum()
2000-10-01 23:38:00     3
2000-10-01 23:55:00    15
2000-10-02 00:12:00    45
2000-10-02 00:29:00    45
Freq: 17T, dtype: int32

文档这样解释origin='end_day':

‘end_day’:原点是最后一天的天花板午夜

据我所知,这条线

ts.resample('17min', origin='end_day').sum()

应该等价于

ts.resample('17min', origin=ts.index.max().ceil('1d')).sum()

但是,传递时间戳ts.index.max().ceil('1d') 会产生不同的结果:

>>> ts.resample('17min', origin=ts.index.max().ceil('1d')).sum() 
2000-10-01 23:21:00     3
2000-10-01 23:38:00    15
2000-10-01 23:55:00    27
2000-10-02 00:12:00    63

我正在寻找对这种差异的解释,也许是对 'end_day' 参数的一般描述比文档提供的更好。

编辑:我正在使用pandas 1.3.5

【问题讨论】:

    标签: python pandas pandas-resample


    【解决方案1】:

    origin='end_day' 的真正等价物是:

    >>> ts.resample('17min', origin=ts.index.max().ceil('D'), 
                    closed='right', label='right').sum()
    
    2000-10-01 23:38:00     3
    2000-10-01 23:55:00    15
    2000-10-02 00:12:00    45
    2000-10-02 00:29:00    45
    Freq: 17T, dtype: int64
    

    更新 1:

    1. 如果我使用 origin='end_day' 但还明确传入已关闭且标签不是“正确”,该怎么办?为此定义的行为在哪里?

    来自resample的source code:

                # The backward resample sets ``closed`` to ``'right'`` by default
                # since the last value should be considered as the edge point for
                # the last bin. When origin in "end" or "end_day", the value for a
                # specific ``Timestamp`` index stands for the resample result from
                # the current ``Timestamp`` minus ``freq`` to the current
                # ``Timestamp`` with a right close.
                if origin in ["end", "end_day"]:
                    if closed is None:
                        closed = "right"
                    if label is None:
                        label = "right"
                else:
                    if closed is None:
                        closed = "left"
                    if label is None:
                        label = "left"
    

    更新 2a:

    1. 考虑df = pd.DataFrame(index=pd.date_range(start='2021-04-22 01:00:00', end='2021-04-28 01:00', freq='1d'), data=range(7))。现在 df.resample(rule='7d', origin='end_day') 崩溃并出现 ValueError。

    如果您没有明确设置closed 参数,则将resample 设置为right,因为origin='end_day'(见上文)。所以origin 现在是 '2021-04-29' 并且第一个 bin 值是 '2021-04-22' 被排除在外。你有一种情况Values falls before first bin:

    df = pd.DataFrame(index=pd.date_range(start='2021-04-22 01:00:00', end='2021-04-28 01:00', freq='1d'), data=range(7))
    df.resample(rule='7d', origin='end_day', closed='left')  # <- HERE
    

    更新 2b:

    如果“2021-04-22”是第一个 bin,那么哪个时间戳不在其中? '2021-04-22 01:00:00' 更晚了,对吧?

    df = pd.DataFrame(index=pd.date_range(start='2021-04-21 01:00:00', end='2021-04-28 01:00', freq='1d'), data=range(8))
    print(df)
    
    # Output:
                         0
    2021-04-21 01:00:00  0
    2021-04-22 01:00:00  1
    2021-04-23 01:00:00  2
    2021-04-24 01:00:00  3
    2021-04-25 01:00:00  4
    2021-04-26 01:00:00  5
    2021-04-27 01:00:00  6
    2021-04-28 01:00:00  7
    

    有了这个样本,我想你应该更清楚:

    # closed='right' (default)
    >>> df.resample(rule='7d', origin='end_day').sum()
                 0
    2021-04-22   1  # ('2021-04-15', '2021-04-22']
    2021-04-29  27  # ('2021-04-22', '2021-04-29']
    
    # closed='left'
    >>> df.resample(rule='7d', origin='end_day', closed='left').sum()
                 0
    2021-04-22   0  # ['2021-04-15', '2021-04-22')
    2021-04-29  28  # ['2021-04-22', '2021-04-29')
    
    bin_edges
    

    bin_edges 的值为:

    # closed='right' (default)
    >>> bin_edges
    [1618531199999999999 1619135999999999999 1619740799999999999]
    
    # after conversion
    DatetimeIndex(['2021-04-15 23:59:59.999999999',
                   '2021-04-22 23:59:59.999999999',
                   '2021-04-29 23:59:59.999999999'],
                  dtype='datetime64[ns]', freq=None)
    
    
    # closed='left'
    >>> bin_edges
    [1618444800000000000 1619049600000000000 1619654400000000000]
    
    # after conversion
    DatetimeIndex(['2021-04-15',
                   '2021-04-22',
                   '2021-04-29'],
                  dtype='datetime64[ns]', freq=None)
    

    【讨论】:

    • 谢谢。我仍然对两点感到困惑。我将把它们分成两个 cmets。 1.如果我使用origin='end_day',但也明确传入closed和label不是'right'怎么办?为此定义的行为在哪里?
    • 2.考虑df = pd.DataFrame(index=pd.date_range(start='2021-04-22 01:00:00', end='2021-04-28 01:00', freq='1d'), data=range(7))。现在df.resample(rule='7d', origin='end_day') 与ValueError 一起崩溃。知道为什么吗?
    • 您的编辑回答了我的第一个问题,谢谢。
    • @actual_panda。我更新了第 2 点的答案。你现在清楚了吗?
    • 谢谢。并不真地。如果“2021-04-22”是第一个 bin,那么哪个时间戳不在其中? '2021-04-22 01:00:00' 更晚了,对吧?即使任何时间戳从第一个 bin 中掉出来,为什么 resample 不添加 bin,直到所有时间戳都被 bin 之后,就像它应该做的那样?
    猜你喜欢
    • 1970-01-01
    • 2018-12-27
    • 2014-01-20
    • 2014-05-07
    • 1970-01-01
    • 1970-01-01
    • 2021-01-07
    • 2013-10-04
    • 1970-01-01
    相关资源
    最近更新 更多