【问题标题】:How can I simplify and format this function?如何简化和格式化此功能?
【发布时间】:2014-08-22 07:23:17
【问题描述】:

所以我有这个凌乱的代码,我想从 frankenstein.txt 中获取每个单词,按字母顺序排序,删除一个和两个字母单词,然后将它们写入一个新文件。

def Dictionary():

    d = []
    count = 0

    bad_char = '~!@#$%^&*()_+{}|:"<>?\`1234567890-=[]\;\',./ '
    replace = ' '*len(bad_char)
    table = str.maketrans(bad_char, replace)

    infile = open('frankenstein.txt', 'r')
    for line in infile:
        line = line.translate(table)
        for word in line.split():
            if len(word) > 2:
                d.append(word)
                count += 1
    infile.close()
    file = open('dictionary.txt', 'w')
    file.write(str(set(d)))
    file.close()

Dictionary() 

如何简化它并使其更具可读性,以及如何使单词垂直写入新文件(它写入水平列表):

abbey
abhorred
about
etc....

【问题讨论】:

    标签: python-3.x split simplify


    【解决方案1】:

    以下一些改进:

    from string import digits, punctuation
    
    def create_dictionary():
    
        words = set()
    
        bad_char = digits + punctuation + '...' # may need more characters
        replace = ' ' * len(bad_char)
        table = str.maketrans(bad_char, replace)
    
        with open('frankenstein.txt') as infile:
            for line in infile:
                line = line.strip().translate(table)
                for word in line.split():
                    if len(word) > 2:
                        words.add(word)
    
        with open('dictionary.txt', 'w') as outfile:
            outfile.writelines(sorted(words)) # note 'lines'
    

    几点说明:

    • 关注the style guide
    • string 包含可用于提供“坏字符”的常量;
    • 你从未使用过count(反正只是len(d));
    • 使用with 上下文管理器进行文件处理;和
    • 从一开始就使用set 可以防止重复,但它们不是有序的(因此sorted)。

    【讨论】:

      【解决方案2】:

      使用 re 模块。

      import re
      
      words = set()
      
      with open('frankenstein.txt') as infile:
          for line in infile:
              words.extend([x for x in re.split(r'[^A-Za-z]*', line) if len(x) > 2])
      
      with open('dictionary.txt', 'w') as outfile:
          outfile.writelines(sorted(words))
      

      从 re.split 中的 r'[^A-Za-z]*' 中,将 'A-Za-z' 替换为您想要的字符包含在 dictionary.txt 中。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2022-10-13
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2021-08-14
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多