【问题标题】:Python idiom: List comprehension with limit of itemsPython 成语:项目限制的列表理解
【发布时间】:2010-08-20 20:32:09
【问题描述】:

我基本上是在尝试这样做(伪代码,无效的 python):

limit = 10
results = [xml_to_dict(artist) for artist in xml.findall('artist') while limit--]

那么我怎样才能以简洁有效的方式编写代码呢? XML 文件可以包含 0 到 50 个艺术家之间的任何内容,我无法控制一次获得多少个艺术家,而且 AFAIK,没有 XPATH 表达式可以说“让我最多 10 个节点”。

谢谢!

【问题讨论】:

    标签: python list-comprehension idioms


    【解决方案1】:

    你在使用lxml吗?您可以使用 XPath 来限制查询级别中的项目,例如

    >>> from lxml import etree
    >>> from io import StringIO
    >>> xml = etree.parse(StringIO('<foo><bar>1</bar><bar>2</bar><bar>4</bar><bar>8</bar></foo>'))
    >>> [bar.text for bar in xml.xpath('bar[position()<=3]')]
    ['1', '2', '4']
    

    你也可以use itertools.islice to limit any iterable,例如

    >>> from itertools import islice
    >>> [bar.text for bar in islice(xml.iterfind('bar'), 3)]
    ['1', '2', '4']
    >>> [bar.text for bar in islice(xml.iterfind('bar'), 5)]
    ['1', '2', '4', '8']
    

    【讨论】:

    • 您认为 XPath 解决方案是否比切片替代方案更快?我认为元素是懒惰的,但是完全不获取元素可能会更快?不确定
    • @Infinity:我认为islice 更快,因为 XPath 更复杂。虽然我还没有验证。您需要自己进行基准测试。
    【解决方案2】:

    假设xml 是一个ElementTree 对象,findall() 方法返回一个列表,所以只需对该列表进行切片:

    limit = 10
    limited_artists = xml.findall('artist')[:limit]
    results = [xml_to_dict(artist) for artist in limited_artists]
    

    【讨论】:

      【解决方案3】:

      对于因试图限制从无限生成器返回的项目而发现此问题的其他所有人:

      from itertools import takewhile
      ltd = takewhile(lambda x: x[0] < MY_LIMIT, enumerate( MY_INFINITE_GENERATOR ))
      # ^ This is still an iterator. 
      # If you want to materialize the items, e.g. in a list, do:
      ltd_m = list( ltd )
      # If you don't want the enumeration indices, you can strip them as usual:
      ltd_no_enum = [ v for i,v in ltd_m ]
      

      编辑:实际上,islice 是一个更好的选择。

      【讨论】:

        【解决方案4】:
        limit = 10
        limited_artists = [artist in xml.findall('artist')][:limit]
        results = [xml_to_dict(artist) for limited_artists]
        

        【讨论】:

          【解决方案5】:

          这避免了切片问题:它不会更改操作顺序,也不会构造新列表,如果您要过滤列表推导式,这对于大型列表可能很重要。

          def first(it, count):
              it = iter(it)
              for i in xrange(0, count):
                  yield next(it)
              raise StopIteration
          
          print [i for i in first(range(1000), 5)]
          

          它也适用于生成器表达式,由于内存使用,切片会失败:

          exp = (i for i in first(xrange(1000000000), 10000000))
          for i in exp:
              print i
          

          【讨论】:

          • 你真的需要提高StopIteration。只需结束函数即可。
          猜你喜欢
          • 1970-01-01
          • 2012-09-22
          • 1970-01-01
          • 2023-03-13
          • 2023-03-26
          • 1970-01-01
          • 2023-02-09
          • 2020-03-22
          • 2023-01-02
          相关资源
          最近更新 更多