【问题标题】:Concise way to split a string into a list of fixed number of tokens将字符串拆分为固定数量标记列表的简洁方法
【发布时间】:2014-03-10 18:50:04
【问题描述】:

我正在编写一段代码,需要将连字符分隔的字符串拆分为最多三个标记。如果拆分后的token少于3个,则需要追加足够数量的空字符串才能组成3个token。

例如,'foo-bar-baz' 应拆分为 ['foo', 'bar', 'baz'],但 foo-bar 应拆分为 ['foo', 'bar', '']。

这是我写的代码。

def three_tokens(s):
    tokens = s.split('-', 2)
    if len(tokens) == 1:
        tokens.append('')
        tokens.append('')
    elif len(tokens) == 2:
        tokens.append('')
    return tokens

print(three_tokens(''))
print(three_tokens('foo'))
print(three_tokens('foo-bar'))
print(three_tokens('foo-bar-baz'))
print(three_tokens('foo-bar-baz-qux'))

这是输出:

['', '', '']
['foo', '', '']
['foo', 'bar', '']
['foo', 'bar', 'baz']
['foo', 'bar', 'baz-qux']

我的问题是我编写的three_tokens 函数对于这个小任务来说似乎太冗长了。有没有一种 Python 的方式来写这个,或者有一些 Python 函数或类专门用于完成这种使代码更简洁的任务?

【问题讨论】:

    标签: python python-3.x


    【解决方案1】:

    您可以使用简单的while 循环:

    def three_tokens(s):
        tokens = s.split('-', 2)
        while len(tokens) < 3:
            tokens.append('')
        return tokens
    

    或使用计算出的空字符串数来扩展列表:

    def three_tokens(s):
        tokens = s.split('-', 2)
        tokens.extend([''] * (3 - len(tokens)))
        return tokens
    

    或者使用连接,这样你就可以把它放在返回语句中:

    def three_tokens(s):
        tokens = s.split('-', 2)
        return tokens + [''] * (3 - len(tokens))
    

    【讨论】:

      【解决方案2】:

      这可能有点矫枉过正,但您可以使用itertools 中的一些方法。

      list(itertools.islice(itertools.chain(s.split('-', 2), itertools.repeat('')), 3)
      

      【讨论】:

        【解决方案3】:

        使用str.partition:

        def three_tokens(s):
            t1, unused, t2 = s.partition('-')
            t2, unused, t3 = t2.partition('-')
            return [t1, t2, t3]
        

        【讨论】:

          【解决方案4】:

          这可以工作。

          tokens = s.split('-', 2)
          tokens += [''] * max(0, 3 - len(tokens))
          

          【讨论】:

          • 不需要max(); s.split() 的限制确保在任何情况下tokens 中的元素都不会超过 3 个。
          • 此外,将一个列表乘以一个负数也会导致一个空列表; max() 在这里完全是矫枉过正。
          【解决方案5】:
          >>> n = 3
          >>> a = '123-abc'
          >>> b = a.split('-', n)
          >>> if len(b) < n-1:
          ...     b = b + ['']*(n-len(b))
          ...
          >>> b
          ['123', 'abc', '']
          >>>
          

          【讨论】:

          • 为什么要使用if 语句和b[:] 副本? [''] * 0 是一个空列表,list-copy-by-slice 完全是多余的,浪费循环。
          • 是的。 b[:] 不是必需的。 b+['']*max(1,2-len(b)) 应该可以解决问题。
          • 看我的comment on another answer to this question,你不需要max()。
          【解决方案6】:
          def three_tokens(s):
              tokens = s.split('-', 2)
              return [tokens.pop(0) if len(tokens) else '' for _ in range(0, 3)]
          

          ...产量

          >>> three_tokens('foo')
          ['foo', '', '']
          
          >>> three_tokens('foo-bar')
          ['foo', 'bar', '']
          
          >>> three_tokens('foo-bar-baz')
          ['foo', 'bar', 'baz']
          
          >>> three_tokens('foo-bar-baz-buzz')
          ['foo', 'bar', 'baz-buzz']
          

          【讨论】:

            【解决方案7】:

            这个怎么样?

            def three_tokens(s):
                output = ['', '', '']
                tokens = s.split('-', 2)
                output[0:len(tokens)] = tokens
                return output
            

            还有一个单线:

            three_tokens = lambda s: (s.split('-', 2) + ['', ''])[:3]
            

            顺便说一句,我在您的解决方案中找不到任何非 Python 的东西。有点冗长,但意图很明确。

            还有一个:

            def three_tokens(s):
               it = iter(s.split('-', 2))
               return [ next(it, '') for _ in range(3) ]
            

            【讨论】:

              猜你喜欢
              • 1970-01-01
              • 1970-01-01
              • 2011-04-24
              • 1970-01-01
              • 2020-11-13
              • 2013-09-15
              • 1970-01-01
              • 1970-01-01
              相关资源
              最近更新 更多