【问题标题】:Python split list into n chunksPython 将列表拆分为 n 个块
【发布时间】:2014-06-30 05:00:25
【问题描述】:

我知道这个问题已经讨论过很多次了,但我的要求不同。

我有一个类似的列表:range(1, 26)。我想把这个列表分成一个固定的数字n。假设 n = 6。

>>> x
[1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25]
>>> l = [ x [i:i + 6] for i in range(0, len(x), 6) ]
>>> l
[[1, 2, 3, 4, 5, 6], [7, 8, 9, 10, 11, 12], [13, 14, 15, 16, 17, 18], [19, 20, 21, 22, 23, 24], [25]]

如您所见,我没有得到 6 个块(包含原始列表元素的六个子列表)。我如何划分一个列表,以便我得到准确的 n 块,这些块可能是均匀的或不均匀的

【问题讨论】:

  • 更一般的,相同的功能:[ np.array(x)[i:i + chunk_size,...] for i in range(0, len(x), chunk_size) ]

标签: python


【解决方案1】:

使用 numpy

>>> import numpy
>>> x = range(25)
>>> l = numpy.array_split(numpy.array(x),6)

>>> import numpy
>>> x = numpy.arange(25)
>>> l = numpy.array_split(x,6);

您也可以使用 numpy.split,但如果长度不能完全整除,则会引发错误。

【讨论】:

  • 我不知道为什么 OP 不接受这个答案。我发现一大堆“答案”漂浮在周围,实际上并没有用,但这个确实有效。
  • 返回一个 numpy 数组的列表,应该明确指出(或显示)。
  • 这仅适用于 numpy 支持的数字类型,不是通用答案。
  • 我也不会接受这个,因为它需要导入.. 这可以通过很多其他方式完成.. 但一如既往,它取决于被拆分的数据.. 地图,[:] 和以此类推。
  • 谢谢你。我在这个问题的更流行版本中偶然发现了非 numpy 答案。
【解决方案2】:

以下解决方案有很多优点:

  • 使用生成器生成结果。
  • 没有进口。
  • 列表是平衡的(如果将长度为 17 的列表拆分为 5 个列表,则永远不会得到 4 个大小为 4 的列表和一个大小为 1 的列表)。
def chunks(l, n):
    """Yield n number of striped chunks from l."""
    for i in range(0, n):
        yield l[i::n]

上面的代码为l = range(16)n = 6 生成以下输出:

[0, 6, 12]
[1, 7, 13]
[2, 8, 14]
[3, 9, 15]
[4, 10]
[5, 11]

如果您需要块是连续的而不是条带的,请使用:

def chunks(l, n):
    """Yield n number of sequential chunks from l."""
    d, r = divmod(len(l), n)
    for i in range(n):
        si = (d+1)*(i if i < r else r) + d*(0 if i < r else i - r)
        yield l[si:si+(d+1 if i < r else d)]

l = range(16)n = 6 产生:

[0, 1, 2]
[3, 4, 5]
[6, 7, 8]
[9, 10, 11]
[12, 13]
[14, 15]

有关发电机优势的更多信息,请参阅this stackoverflow link

【讨论】:

  • 不错的顺序解决方案。我会指出list(chunks(list(range(5)),6)) 产生[[0], [1], [2], [3], [4], []] 这是公平的。
【解决方案3】:

如果顺序无关紧要:

def chunker_list(seq, size):
    return (seq[i::size] for i in range(size))

print(list(chunker_list([1, 2, 3, 4, 5], 2)))
>>> [[1, 3, 5], [2, 4]]

print(list(chunker_list([1, 2, 3, 4, 5], 3)))
>>> [[1, 4], [2, 5], [3]]

print(list(chunker_list([1, 2, 3, 4, 5], 4)))
>>> [[1, 5], [2], [3], [4]]

print(list(chunker_list([1, 2, 3, 4, 5], 5)))
>>> [[1], [2], [3], [4], [5]]

print(list(chunker_list([1, 2, 3, 4, 5], 6)))
>>> [[1], [2], [3], [4], [5], []]

【讨论】:

    【解决方案4】:

    more_itertools.divide 是解决此问题的一种方法:

    import more_itertools as mit
    
    
    iterable = range(1, 26)
    [list(c) for c in mit.divide(6, iterable)]
    

    输出

    [[ 1,  2,  3,  4, 5],                       # remaining item
     [ 6,  7,  8,  9],
     [10, 11, 12, 13],
     [14, 15, 16, 17],
     [18, 19, 20, 21],
     [22, 23, 24, 25]]
    

    如图所示,如果iterable不是整除的,剩下的item会从第一个chunk分配到最后一个chunk。

    详细了解more_itertoolshere

    【讨论】:

      【解决方案5】:

      我的答案是简单地使用python内置的Slice:

      # Assume x is our list which we wish to slice
      x = range(1, 26)
      # Assume we want to slice it to 6 equal chunks
      result = []
      for i in range(0, len(x), 6):
          slice_item = slice(i, i + 6, 1)
          result.append(x[slice_item])
      
      # Result would be equal to 
      

      [[0,1,2,3,4,5], [6,7,8,9,10,11], [12,13,14,15,16,17],[18,19,20,21,22,23], [24, 25]]

      【讨论】:

      • 这没有回答问题。它分成大小为 6 的块,而不是 6 块。
      【解决方案6】:

      试试这个:

      from __future__ import division
      
      import math
      
      def chunked(iterable, n):
          """ Split iterable into ``n`` iterables of similar size
      
          Examples::
              >>> l = [1, 2, 3, 4]
              >>> list(chunked(l, 4))
              [[1], [2], [3], [4]]
      
              >>> l = [1, 2, 3]
              >>> list(chunked(l, 4))
              [[1], [2], [3], []]
      
              >>> l = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
              >>> list(chunked(l, 4))
              [[1, 2, 3], [4, 5, 6], [7, 8, 9], [10]]
      
          """
          chunksize = int(math.ceil(len(iterable) / n))
          return (iterable[i * chunksize:i * chunksize + chunksize]
                  for i in range(n))
      

      它返回一个迭代器而不是一个列表以提高效率(我假设你想循环访问块),但如果你愿意,你可以用列表推导替换它。当项目数不能被块数整除时,最后一个块小于其他块。

      编辑:修复了第二个示例以显示它不处理一种边缘情况

      【讨论】:

        【解决方案7】:

        提示:

        • x 是要拆分的字符串。
        • k 是块的数量

          n = len(x)/k
          
          [x[i:i+n] for i in range(0, len(x), n)]
          

        【讨论】:

          【解决方案8】:

          拿走我的 2 美分..

          from math import ceil
          
          size = 3
          seq = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11]
          
          chunks = [
              seq[i * size:(i * size) + size]
              for i in range(ceil(len(seq) / size))
          ]
          
          # [[1, 2, 3], [4, 5, 6], [7, 8, 9], [10, 11]]
          
          

          【讨论】:

            【解决方案9】:

            一种方法是使最后一个列表不均匀,其余的则均匀。这可以按如下方式完成:

            >>> x
            [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25]
            >>> m = len(x) // 6
            >>> test = [x[i:i+m] for i in range(0, len(x), m)]
            >>> test[-2:] = [test[-2] + test[-1]]
            >>> test
            [[1, 2, 3, 4], [5, 6, 7, 8], [9, 10, 11, 12], [13, 14, 15, 16], [17, 18, 19, 20], [21, 22, 23, 24, 25]]
            

            【讨论】:

            • 可能需要len(x) // 6 以实现 py3 兼容性。
            • x6 的倍数时,这会返回错误的结果,因为在这种情况下,列表的数量是正确的,而无需重新分配最后一个元素。
            • @Bakuriu:可以检查在第一种情况下是否返回了所需数量的列表,如果不应用转换。
            • 如果列表的长度是5,你想分成3块呢?
            【解决方案10】:

            假设你想分成n个块:

            n = 6
            num = float(len(x))/n
            l = [ x [i:i + int(num)] for i in range(0, (n-1)*int(num), int(num))]
            l.append(x[(n-1)*int(num):])
            

            此方法只是将列表的长度除以块的数量,如果长度不是数字的倍数,则在最后一个列表中添加额外的元素。

            【讨论】:

              【解决方案11】:

              如果你想让块的大小尽可能均匀:

              def chunk_ranges(items: int, chunks: int) -> List[Tuple[int, int]]:
                  """
                  Split the items by best effort into equally-sized chunks.
                  
                  If there are fewer items than chunks, each chunk contains an item and 
                  there are fewer returned chunk indices than the argument `chunks`.
              
                  :param items: number of items in the batch.
                  :param chunks: number of chunks
                  :return: list of (chunk begin inclusive, chunk end exclusive)
                  """
                  assert chunks > 0, \
                      "Unexpected non-positive chunk count: {}".format(chunks)
              
                  result = []  # type: List[Tuple[int, int]]
                  if items <= chunks:
                      for i in range(0, items):
                          result.append((i, i + 1))
                      return result
              
                  chunk_size, extras = divmod(items, chunks)
              
                  start = 0
                  for i in range(0, chunks):
                      if i < extras:
                          end = start + chunk_size + 1
                      else:
                          end = start + chunk_size
              
                      result.append((start, end))
                      start = end
              
                  return result
              

              测试用例:

              def test_chunk_ranges(self):
                  self.assertListEqual(chunk_ranges(items=8, chunks=1),
                                       [(0, 8)])
              
                  self.assertListEqual(chunk_ranges(items=8, chunks=2),
                                       [(0, 4), (4, 8)])
              
                  self.assertListEqual(chunk_ranges(items=8, chunks=3),
                                       [(0, 3), (3, 6), (6, 8)])
              
                  self.assertListEqual(chunk_ranges(items=8, chunks=5),
                                       [(0, 2), (2, 4), (4, 6), (6, 7), (7, 8)])
              
                  self.assertListEqual(chunk_ranges(items=8, chunks=6),
                                       [(0, 2), (2, 4), (4, 5), (5, 6), (6, 7), (7, 8)])
              
                  self.assertListEqual(
                      chunk_ranges(items=8, chunks=7),
                      [(0, 2), (2, 3), (3, 4), (4, 5), (5, 6), (6, 7), (7, 8)])
              
                  self.assertListEqual(
                      chunk_ranges(items=8, chunks=9),
                      [(0, 1), (1, 2), (2, 3), (3, 4), (4, 5), (5, 6), (6, 7), (7, 8)])
              

              【讨论】:

              • 由于不使用短变量和测试用例而投票。
              【解决方案12】:

              如果您的列表包含不同类型的元素或存储不同类型值的可迭代对象(例如,某些元素是整数,有些是字符串),如果您使用 numpy 包中的 array_split 函数进行拆分它,你会得到具有相同类型元素的块:

              import numpy as np
              
              data1 = [(1, 2), ('a', 'b'), (3, 4), (5, 6), ('c', 'd'), ('e', 'f')]
              chunks = np.array_split(data1, 3)
              print(chunks)
              # [array([['1', '2'],
              #        ['a', 'b']], dtype='<U11'), array([['3', '4'],
              #        ['5', '6']], dtype='<U11'), array([['c', 'd'],
              #        ['e', 'f']], dtype='<U11')]
              
              data2 = [1, 2, 'a', 'b', 3, 4, 5, 6, 'c', 'd', 'e', 'f']
              chunks = np.array_split(data2, 3)
              print(chunks)
              # [array(['1', '2', 'a', 'b'], dtype='<U11'), array(['3', '4', '5', '6'], dtype='<U11'),
              #  array(['c', 'd', 'e', 'f'], dtype='<U11')]
              

              如果您想在列表拆分后将元素的初始类型分成块,您可以修改numpy包中的array_split函数的source code或使用this implementation

              from itertools import accumulate
              
              def list_split(input_list, num_of_chunks):
                  n_total = len(input_list)
                  n_each_chunk, extras = divmod(n_total, num_of_chunks)
                  chunk_sizes = ([0] + extras * [n_each_chunk + 1] + (num_of_chunks - extras) * [n_each_chunk])
                  div_points = list(accumulate(chunk_sizes))
                  sub_lists = []
                  for i in range(num_of_chunks):
                      start = div_points[i]
                      end = div_points[i + 1]
                      sub_lists.append(input_list[start:end])
                  return (sub_list for sub_list in sub_lists)
              
              result = list(list_split(data1, 3))
              print(result)
              # [[(1, 2), ('a', 'b')], [(3, 4), (5, 6)], [('c', 'd'), ('e', 'f')]]
              
              result = list(list_split(data2, 3))
              print(result)
              # [[1, 2, 'a', 'b'], [3, 4, 5, 6], ['c', 'd', 'e', 'f']]
              

              【讨论】:

                【解决方案13】:

                此解决方案基于 Python 3 文档中的 zip "grouper" 模式。小的补充是,如果 N 不均分列表长度,则所有多余的项都放入第一个块中。

                import itertools
                
                def segment_list(l, N):
                    chunk_size, remainder = divmod(len(l), N)
                    first, rest = l[:chunk_size + remainder], l[chunk_size + remainder:]
                    return itertools.chain([first], zip(*[iter(rest)] * chunk_size))
                

                示例用法:

                >>> my_list = list(range(10))
                >>> segment_list(my_list, 2)
                [[0, 1, 2, 3, 4], (5, 6, 7, 8, 9)]
                >>> segment_list(my_list, 3)
                [[0, 1, 2, 3], (4, 5, 6), (7, 8, 9)]
                >>>
                

                这种解决方案的优点是它保留了原始列表的顺序,并且以函数式风格编写,在调用时只对列表进行一次惰性求值。

                注意,因为它返回一个迭代器,所以结果只能被消费一次。如果你想要一个非惰性列表的便利,你可以将结果包装在list

                >>> x = list(segment_list(my_list, 2))
                >>> x
                [[0, 1, 2, 3, 4], (5, 6, 7, 8, 9)]
                >>> x
                [[0, 1, 2, 3, 4], (5, 6, 7, 8, 9)]
                >>>
                

                【讨论】:

                  【解决方案14】:

                  我会做 (假设你想要 n 个块)

                  import numpy as np
                  
                  # convert x to numpy.ndarray
                  x = np.array(x)
                  l = np.array_split(x, n)
                  

                  它有效,它只有 2 行。

                  示例:

                  # your list
                  x = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25]
                  # amount of chunks you want
                  n = 6
                  x = np.array(x)
                  l = np.array_split(x, n)
                  print(l)
                  
                  >> [array([1, 2, 3, 4, 5]), array([6, 7, 8, 9]), array([10, 11, 12, 13]), array([14, 15, 16, 17]), array([18, 19, 20, 21]), array([22, 23, 24, 25])]
                  

                  如果你想要一个列表列表:

                  l = [list(elem) for elem in l]
                  print(l)
                  
                   >> [[1, 2, 3, 4, 5], [6, 7, 8, 9], [10, 11, 12, 13], [14, 15, 16, 17], [18, 19, 20, 21], [22, 23, 24, 25]]
                  

                  【讨论】:

                  • 这是一个长期存在的问题,有很多答案,请考虑添加解释,说明为什么这是 OP 的最佳答案,因此应按预期标记。
                  • 感谢您的反馈,我刚刚做了(并添加了更正)
                  【解决方案15】:
                  x=[1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25]
                  chunk = len(x)/6
                  
                  l=[]
                  i=0
                  while i<len(x):
                      if len(l)<=4:
                          l.append(x [i:i + chunk])
                      else:
                          l.append(x [i:])
                          break
                      i+=chunk   
                  
                  print l
                  
                  #output=[[1, 2, 3, 4], [5, 6, 7, 8], [9, 10, 11, 12], [13, 14, 15, 16], [17, 18, 19, 20], [21, 22, 23, 24, 25]]
                  

                  【讨论】:

                    【解决方案16】:
                    arr1=[-20, 20, -10, 0, 4, 8, 10, 6, 15, 9, 18, 35, 40, -30, -90, 99]
                    n=4
                    final = [arr1[i * n:(i + 1) * n] for i in range((len(arr1) + n - 1) // n )]
                    print(final)
                    

                    输出:

                    [[-20, 20, -10, 0], [4, 8, 10, 6], [15, 9, 18, 35], [40, -30, -90, 99]]

                    【讨论】:

                    • 此答案似乎并未尝试回答原始问题,因为结果复制了与提出问题以解决相同的问题。
                    【解决方案17】:

                    此函数将返回一个列表(块)中具有设置的最大值数量的列表列表。

                    def chuncker(list_to_split, chunk_size):
                        list_of_chunks =[]
                        start_chunk = 0
                        end_chunk = start_chunk+chunk_size
                        while end_chunk <= len(list_to_split)+chunk_size:
                            chunk_ls = list_to_split[start_chunk: end_chunk]
                            list_of_chunks.append(chunk_ls)
                            start_chunk = start_chunk +chunk_size
                            end_chunk = end_chunk+chunk_size    
                        return list_of_chunks
                    

                    例子:

                    ls = list(range(20))
                    
                    chuncker(list_to_split = ls, chunk_size = 6)
                    

                    输出:

                    [[0, 1, 2, 3, 4, 5], [6, 7, 8, 9, 10, 11], [12, 13, 14, 15, 16, 17], [18, 19 ]]

                    【讨论】:

                      【解决方案18】:

                      我想出了以下解决方案:

                      l = [x[i::n] for i in range(n)]
                      

                      例如:

                      n = 6
                      x = list(range(26))
                      
                      l = [x[i::n] for i in range(n)]
                      print(l)
                      

                      输出:

                      [[0, 6, 12, 18, 24], [1, 7, 13, 19, 25], [2, 8, 14, 20], [3, 9, 15, 21], [4, 10, 16, 22], [5, 11, 17, 23]]
                      

                      如您所见,输出由 n 块组成,它们的元素数量大致相同。


                      它是如何工作的?

                      诀窍是使用列表切片步长(两个分号后的数字)并增加步进切片的偏移量。首先,它需要从第一个开始的每个 n 元素,然后是从第二个开始的每个 n 元素,依此类推。这样就完成了任务。

                      【讨论】:

                        【解决方案19】:

                        对于在没有导入的情况下在 python 3(.6) 中寻找答案的人。
                        x 是要拆分的列表。
                        n 是块的长度。
                        L 是新列表。

                        n = 6
                        L = [x[i:i + int(n)] for i in range(0, (n - 1) * int(n), int(n))]
                        
                        #[[1, 2, 3, 4, 5, 6], [7, 8, 9, 10, 11, 12], [13, 14, 15, 16, 17, 18], [19, 20, 21, 22, 23, 24], [25]]
                        

                        【讨论】:

                        • 我相信问题是分成n 块。这个答案分成大小为n的块。
                        猜你喜欢
                        • 1970-01-01
                        • 1970-01-01
                        • 1970-01-01
                        • 2021-09-19
                        • 1970-01-01
                        • 2021-10-23
                        • 2020-08-07
                        • 1970-01-01
                        • 1970-01-01
                        相关资源
                        最近更新 更多