【问题标题】:Counting the number of consecutive values that meets a condition (Pandas Dataframe)计算满足条件的连续值的数量(Pandas Dataframe)
【发布时间】:2019-03-11 07:07:59
【问题描述】:

因此,我在 2 天前就我的问题创建了 this 帖子,谢天谢地得到了答复。

我有一个由 20 行和 2500 列组成的数据。每列是一个独特的产品,行是时间序列,测量结果。因此每个产品测量 20 次,有 2500 个产品。

这次我想知道我的测量结果可以连续多少行保持在特定阈值之上。 又名:我想计算大于某个值的连续值的数量,比如说 5。

A = [1, 2, 6, 8, 7, 3, 2, 3, 6, 10, 2, 1, 0, 2] 我们有这些粗体值,根据我上面定义的,我应该得到 NumofConsFeature = 3 作为结果。 (满足条件的系列超过1个,取最大值)

我想过使用 .gt 进行过滤,然后获取索引并在之后使用循环以检测连续的索引号,但无法使其工作。

在第二阶段,我想知道我的连续系列的第一个值的索引。对于上面的示例,这将是 3。 但我不知道这个怎么做。

提前致谢。

【问题讨论】:

  • 考虑接受您之前帖子中的答案 - 您可以通过单击答案旁边的复选标记来做到这一点。
  • 应该6, 8, 7 是6, 7, 8?
  • 你的意思是第一个值的索引应该是2而不是3?对于A=[1,2,6,7,8...],索引 6 是 2。
  • 从0开始,你是对的,应该是2。不,6, 8, 7没有理由从小到大排序
  • 好的 - 但6,8,7 不是一个连续的系列。您如何确定对子序列进行排序的窗口?

标签: python pandas numpy dataframe series


【解决方案1】:

这是仅使用 Pandas 函数的另一个答案:

A = [1, 2, 6, 8, 7, 3, 2, 3, 6, 10, 2, 1, 0, 2]
a = pd.DataFrame(A, columns = ['foo'])
a['is_large'] = (a.foo > 5)
a['crossing'] = (a.is_large != a.is_large.shift()).cumsum()
a['count'] = a.groupby(['is_large', 'crossing']).cumcount(ascending=False) + 1
a.loc[a.is_large == False, 'count'] = 0

给了

    foo  is_large  crossing  count
0     1     False         1      0
1     2     False         1      0
2     6      True         2      3
3     8      True         2      2
4     7      True         2      1
5     3     False         3      0
6     2     False         3      0
7     3     False         3      0
8     6      True         4      2
9    10      True         4      1
10    2     False         5      0
11    1     False         5      0
12    0     False         5      0
13    2     False         5      0

从那里您可以轻松找到最大值及其索引。

【讨论】:

    【解决方案2】:

    有一种简单的方法可以做到这一点。
    假设您的列表如下: A = [1, 2, 6, 8, 7, 6, 8, 3, 2, 3, 6, 10,6,7,8, 2, 1, 0, 2]
    并且您想找出有多少个连续系列的值大于 6 且长度为 5。例如,您的答案是 2。有两个系列的值大于 6 且长度为系列是 5。在 python 和 pandas 中,我们这样做如下:

     condition = (df.wanted_row > 6) & \
                (df.wanted_row.shift(-1) > 6) & \
                (df.wanted_row.shift(-2) > 6) & \
                (df.wanted_row.shift(-3) > 6) & \
                (df.wanted_row.shift(-4) > 6)
    
    consecutive_count = df[condition].count().head(1)[0]
    

    【讨论】:

    • 在这里,我用很好的语法写了它:condition = eval(' & '.join([f'(pos.shift({x})>6)' for x in range(num_consecutive)])) = 1
    【解决方案3】:

    这是一个maxisland_start_len_mask -

    # https://stackoverflow.com/a/52718782/ @Divakar
    def maxisland_start_len_mask(a, fillna_index = -1, fillna_len = 0):
        # a is a boolean array
    
        pad = np.zeros(a.shape[1],dtype=bool)
        mask = np.vstack((pad, a, pad))
    
        mask_step = mask[1:] != mask[:-1]
        idx = np.flatnonzero(mask_step.T)
        island_starts = idx[::2]
        island_lens = idx[1::2] - idx[::2]
        n_islands_percol = mask_step.sum(0)//2
    
        bins = np.repeat(np.arange(a.shape[1]),n_islands_percol)
        scale = island_lens.max()+1
    
        scaled_idx = np.argsort(scale*bins + island_lens)
        grp_shift_idx = np.r_[0,n_islands_percol.cumsum()]
        max_island_starts = island_starts[scaled_idx[grp_shift_idx[1:]-1]]
    
        max_island_percol_start = max_island_starts%(a.shape[0]+1)
    
        valid = n_islands_percol!=0
        cut_idx = grp_shift_idx[:-1][valid]
        max_island_percol_len = np.maximum.reduceat(island_lens, cut_idx)
    
        out_len = np.full(a.shape[1], fillna_len, dtype=int)
        out_len[valid] = max_island_percol_len
        out_index = np.where(valid,max_island_percol_start,fillna_index)
        return out_index, out_len
    
    def maxisland_start_len(a, trigger_val, comp_func=np.greater):
        # a is 2D array as the data
        mask = comp_func(a,trigger_val)
        return maxisland_start_len_mask(mask, fillna_index = -1, fillna_len = 0)
    

    示例运行 -

    In [169]: a
    Out[169]: 
    array([[ 1,  0,  3],
           [ 2,  7,  3],
           [ 6,  8,  4],
           [ 8,  6,  8],
           [ 7,  1,  6],
           [ 3,  7,  8],
           [ 2,  5,  8],
           [ 3,  3,  0],
           [ 6,  5,  0],
           [10,  3,  8],
           [ 2,  3,  3],
           [ 1,  7,  0],
           [ 0,  0,  4],
           [ 2,  3,  2]])
    
    # Per column results
    In [170]: row_index, length = maxisland_start_len(a, 5)
    
    In [172]: row_index
    Out[172]: array([2, 1, 3])
    
    In [173]: length
    Out[173]: array([3, 3, 4])
    

    【讨论】:

    • 感谢您的精彩回复。尽管我很难理解它,但它在示例数组 a 上效果很好。但是,当我尝试使用数据时,它给了我一个错误,上面写着File "<string>", line 74, in maxisland_start_len IndexError: index 126 out-of-bounds in maximum.reduceat [0, 126),第 74 行是max_island_percol_len = np.maximum.reduceat(island_lens, grp_shift_idx[:-1]) 我是新手,所以语法对我来说有点陌生。我试图理解,但没有 cmets 这不是很容易。你能详细说明一下吗?提前谢谢你。
    • @crinix 输入是 NumPy 数组还是其他什么?
    • 它是 Pandas DataFrame,但是被转换为带有 na = df.values 的 Numpy 数组
    • @crinix na 的形状和数据类型是什么?
    • 形状:(34, 288) Dtype:float64
    【解决方案4】:

    您可以在您的系列上应用diff(),然后只计算差值为 1 且实际值高于您的截止值的连续条目数。最大计数是连续值的最大数量。

    第一次计算diff():

    df = pd.DataFrame({"a":[1, 2, 6, 7, 8, 3, 2, 3, 6, 10, 2, 1, 0, 2]})
    df['b'] = df.a.diff()
    
    df
         a    b
    0    1  NaN
    1    2  1.0
    2    6  4.0
    3    7  1.0
    4    8  1.0
    5    3 -5.0
    6    2 -1.0
    7    3  1.0
    8    6  3.0
    9   10  4.0
    10   2 -8.0
    11   1 -1.0
    12   0 -1.0
    13   2  2.0
    

    现在计算连续序列:

    above = 5
    n_consec = 1
    max_n_consec = 1
    
    for a, b in df.values[1:]:
        if (a > above) & (b == 1):
            n_consec += 1
        else: # check for new max, then start again from 1
            max_n_consec = max(n_consec, max_n_consec)
            n_consec = 1
    
    max_n_consec
    3
    

    【讨论】:

    • 感谢您的回复,但是我有很多专栏,而且它们只会随着时间的推移而增加,因为它们是独特的产品。因此我不能为每个系列/产品编写代码。所以我需要一个更通用的解决方案,用于无限数量的列。
    【解决方案5】:

    这是我使用 numpy 的方法:

    import pandas as pd
    import numpy as np
    
    
    df = pd.DataFrame({"a":[1, 2, 6, 7, 8, 3, 2, 3, 6, 10, 2, 1, 0, 2]})
    
    
    consecutive_steps = 2
    marginal_price = 5
    
    assertions = [(df.loc[:, "a"].shift(-i) < marginal_price) for i in range(consecutive_steps)]
    condition = np.all(assertions, axis=0)
    
    consecutive_count = df.loc[condition, :].count()
    print(consecutive_count)
    

    产生6。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-11-14
      • 1970-01-01
      • 2019-07-13
      • 2019-05-10
      相关资源
      最近更新 更多