【问题标题】:How to extract the substring from the end on the string?如何从字符串的末尾提取子字符串?
【发布时间】:2020-12-07 20:45:35
【问题描述】:

我正在尝试从我的文件名末尾提取时间戳。

(示例文件名:ABC_xyz_march_2020.xlsx 或 xyz_mno__20_07_2019.xlsx、xyz_spa_20-07-2019.xlsx、xyz-mar_2019.csv、ABC-dec-5.csv 等) (输出:march_2020、20_07_2019、20-07-2019、mar_2019、dec-5 等)

我正在使用字符串的split()函数,但我没有提取单词中的月份。

谁能提出不同的方法?

【问题讨论】:

    标签: python string split


    【解决方案1】:

    我不经常使用正则表达式,但有时这是适合这项工作的工具:

    import re
    
    PATTERNS = [
        re.compile(pattern)
        for pattern in (
            r"([A-Za-z]+_\d{4})\..*",  # {month_name}_{year}
            r"(\d{2}_\d{2}_\d{4})\..*",  # {day}_{month}_{year}
            r"(\d{2}-\d{2}-\d{4})\..*",  # {day}-{month}-{year}
            r"([A-Za-z]+-\d{1,2})\..*",  # {month_name}-day
        )
    ]
    
    inputs = [
        "ABC_xyz_march_2020.xlsx",
        "xyz_mno__20_07_2019.xlsx",
        "xyz_spa_20-07-2019.xlsx",
        "xyz-mar_2019.csv",
        "ABC-dec-5.csv",
    ]
    
    outputs = [
        "march_2020",
        "20_07_2019",
        "20-07-2019",
        "mar_2019",
        "dec-5",
    ]
    
    for filename, expected_output in zip(inputs, outputs):
        for pattern in PATTERNS:
            match = pattern.search(filename)
            if not match:
                continue
            matched_date = match.group(1)
            if matched_date != expected_output:
                raise ValueError(
                    f"In date {filename=}, got {matched_date=} instead of {expected_output=}"
                )
            print(f"Looked at {filename=} and found {matched_date=}")
            break
    

    这会构建一个正则表达式对象列表。然后对于每个输入文件名,它会尝试将文件名与每个正则表达式匹配,直到匹配。错误处理(或在 no 模式匹配时决定做什么)留给读者。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2010-11-05
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-04-20
      相关资源
      最近更新 更多