【问题标题】:Convert a list of numbers to ranges将数字列表转换为范围
【发布时间】:2017-10-02 22:35:48
【问题描述】:

我有一堆数字,说如下:

1 2 3 4  6 7 8  20 24 28 32

那里提供的信息可以在 Python 中表示为范围:

[range(1, 5), range(6, 9), range(20, 33, 4)]

在我的输出中,我会写 1..4, 6..8, 20..32..4,但这只是表示的问题。

Another answer 展示了如何为连续范围执行此操作。我不知道如何轻松地为上面的跨步范围做到这一点。有没有类似的技巧?

【问题讨论】:

    标签: python


    【解决方案1】:

    这是解决问题的直接方法。

    def get_ranges(ls):
        N = len(ls)
        while ls:
            # single element remains, yield the trivial range
            if N == 1:
                yield range(ls[0], ls[0] + 1)
                break
    
            diff = ls[1] - ls[0]
            # find the last index that satisfies the determined difference
            i = next(i for i in range(1, N) if i + 1 == N or ls[i+1] - ls[i] != diff)
    
            yield range(ls[0], ls[i] + 1, diff)
    
            # update variables
            ls = ls[i+1:]
            N -= i + 1
    

    【讨论】:

    • get_ranges([1,2,4,5,7,9] 在末尾给出 [7, 9] 的范围。
    • @George 你会期待什么?上述算法将按预期产生 [1,2]、[4,5]、[7,9],因为它贪婪地填充范围。如果您想要一个非贪心算法,则需要一种完全不同的方法,并且这个问题中没有任何内容表明它是。
    • 哎呀,我误解了这个问题。没关系:)
    【解决方案2】:

    它可能不是超级短或优雅,但它似乎工作:

    def ranges(ls):
        li = iter(ls)
        first = next(li)
        while True:
            try:
                element = next(li)
            except StopIteration:
                yield range(first, first+1)
                return
            step = element - first
            last = element
            while True:
                try:
                    element = next(li)
                except StopIteration:
                    yield range(first, last+step, step)
                    return
                if element - last != step:
                    yield range(first, last+step, step)
                    first = element
                    break
                last = element
    

    这遍历列表的迭代器,并产生范围对象:

    >>> list(ranges([1, 2, 3, 4, 6, 7, 8, 20, 24, 28, 32]))
    [range(1, 5), range(6, 9), range(20, 33, 4)]
    

    它还处理负范围,以及只有一个元素的范围:

    >>> list(ranges([9,8,7, 1,3,5, 99])
    [range(9, 6, -1), range(1, 7, 2), range(99, 100)]
    

    【讨论】:

      【解决方案3】:

      您可以使用来自itertools 模块的groupbycount 以及来自collections 模块的Counter,如下例所示:

      更新:查看 cmets 以了解此解决方案背后的逻辑及其局限性。

      from itertools import groupby, count
      from collections import Counter
      
      def ranges_list(data=list, func=range, min_condition=1):
          # Sort in place the ranges list
          data.sort()
      
          # Find all the steps between the ranges's elements
          steps = [v-k for k,v in zip(data, data[1:])]
      
          # Find the repeated items's steps based on condition. 
          # Default: repeated more than once (min_condition = 1)
          repeated = [item for item, count in Counter(steps).items() if count > min_condition]
      
          # Group the items in to a dict based on the repeated steps
          groups = {k:[list(v) for _,v in groupby(data, lambda n, c = count(step = k): n-next(c))] for k in repeated}
      
          # Create a dict:
          # - keys are the steps
          # - values are the grouped elements
          sub = {k:[j for j in v if len(j) > 1] for k,v in groups.items()}
      
          # Those two lines are for pretty printing purpose:
          # They are meant to have a sorted output.
          # You can replace them by:
          # return [func(j[0], j[-1]+1,k) for k,v in sub.items() for j in v]
          # Otherwise:
          final = [(j[0], j[-1]+1,k) for k,v in sub.items() for j in v]
          return [func(*k) for k in sorted(final, key = lambda x: x[0])]
      
      ranges1 = [1, 2, 3, 4, 6, 7, 8, 20, 24, 28, 32]
      ranges2 = [1, 2, 3, 4, 6, 7, 10, 20, 24, 28, 50,51,59,60]
      
      print(ranges_list(ranges1))
      print(ranges_list(ranges2))
      

      输出:

      [range(1, 5), range(6, 9), range(20, 33, 4)]
      [range(1, 5), range(6, 8), range(20, 29, 4), range(50, 52), range(59, 61)]
      

      限制:

      使用这种输入:

      ranges3 = [1,3,6,10]
      print(ranges_list(ranges3)
      print(ranges_list(ranges3, min_condition=0))
      

      将输出:

      # Steps are repeated <= 1 with the condition: min_condition = 1
      # Will output an empty list
      []
      # With min_condition = 0
      # Will output the ranges using: zip(data, data[1:])
      [range(1, 4, 2), range(3, 7, 3), range(6, 11, 4)]
      

      随意使用此解决方案并采用或修改它以满足您的需求。

      【讨论】:

      • 第二个序列不应该产生range(10, 21, 10)吗?
      • 是的,当我设置条件 min_confirmation = 0 时,它会输出:[range(1, 5), range(4, 7, 2), range(6, 8), range(7, 11, 3), range(10, 21, 10), range(20, 29, 4), range(28, 51, 22), range(50, 52), range(51, 60, 8), range(59, 61)] 所以包含range(10, 21, 10)。这列在第三个序列中的限制下,我假设这将产生不希望的输出。我仍在等待 OP 评论以保留这样的代码或修改它。
      【解决方案4】:
      def ranges(data):
          result = []
          if not data:
              return result
          idata = iter(data)
          first = prev = next(idata)
          for following in idata:
              if following - prev == 1:
                  prev = following
              else:
                  result.append((first, prev + 1))
                  first = prev = following
          # There was either exactly 1 element and the loop never ran,
          # or the loop just normally ended and we need to account
          # for the last remaining range.
          result.append((first, prev+1))
          return result
      

      测试:

      >>> data = range(1, 5) + range(6, 9) + range(20, 24)
      >>> print ranges(data)
      [(1, 5), (6, 9), (20, 24)]
      

      【讨论】:

      • @JaredGoguen:作为练习,添加一个设置步骤的参数。步骤的自动检测需要初步全面扫描。
      • 我不相信这是真的。查看其他答案,包括我的。
      • @JaredGoguen:您的解决方案假定输入数据以长于 1 个元素的序列开头,因此第一个差异是步骤。考虑输入[10, 15, 5, 6, 7, 30, 31, 32]。可以确定步骤,例如作为元素之间的最小差异(需要完整扫描),但是当范围接触时会导致问题:[1, 2, 3, 3, 4, 5],甚至相交:[10, 11, 12, 11, 12, 13]。与任何模式识别问题一样,步进自动检测并非易事,需要明确说明一些假设。
      猜你喜欢
      • 1970-01-01
      • 2012-04-08
      • 2018-01-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-07-13
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多