【问题标题】:Python imaplib - get_filename() not working when attachment has UTF-8 charactersPython imaplib - 当附件包含 UTF-8 字符时,get_filename() 不起作用
【发布时间】:2020-09-10 15:41:54
【问题描述】:

我有这个功能,可以使用 imaplib 从给定电子邮件中下载所有附件

# Download all attachment files for a given email
def downloaAttachmentsInEmail(m, emailid, outputdir, markRead):
    resp, data = m.uid("FETCH", emailid, "(BODY.PEEK[])")
    email_body = data[0][1]
    mail = email.message_from_bytes(email_body)
    if mail.get_content_maintype() != 'multipart':
        return
    for part in mail.walk():
        if part.get_content_maintype() != 'multipart' and part.get('Content-Disposition') is not None:
            open(outputdir + '/' + part.get_filename(), 'wb').write(part.get_payload(decode=True)
    if(markRead):
        m.uid("STORE", emailid, "+FLAGS", "(\Seen)")

问题是当我尝试下载文件名中包含 UTF-8 字符的文件时它不起作用。我收到此错误,我猜这是因为 part.get_filename() 没有正确读取名称:

    OSError: [Errno 22] Invalid argument: './temp//=?UTF-8?B?QkQgUmVsYXTDs3JpbyAywqogRmFzZS5kb2M=?=\r\n\t=?UTF-8?B?eA==?='

我能做些什么来解决这个问题?

【问题讨论】:

    标签: python python-3.x imaplib


    【解决方案1】:

    这是一个老问题,但我正面临这个问题并且很难找到解决方案......也许这可以帮助其他人!

    编辑:这仅涵盖将文件名“解码”为正确名称的部分!

    import re
    import base64
    import quopri
    
    def encoded_words_to_text(encoded_words):
        try:
            encoded_word_regex = r'=\?{1}(.+)\?{1}([B|Q])\?{1}(.+)\?{1}='
            charset, encoding, encoded_text = re.match(encoded_word_regex, encoded_words).groups()
            if encoding is 'B':
                byte_string = base64.b64decode(encoded_text)
            elif encoding is 'Q':
                byte_string = quopri.decodestring(encoded_text)
            return byte_string.decode(charset)
        except:
            return encoded_words
    

    结果:

    test_string = '=?utf-8?B?SUJUIFB1cmNoYXNlIE9yZGVyLnBkZg==?='
    encoded_words_to_text(test_string)
    'IBT Purchase Order.pdf'
    

    【讨论】:

    • 可以添加所有导入的模块吗?
    【解决方案2】:

    我找到了解决办法

    # Download all attachment files for a given email
    def downloaAttachmentsInEmail(m, emailid, outputdir, markRead):
        resp, data = m.uid("FETCH", emailid, "(BODY.PEEK[])")
        email_body = data[0][1]
        mail = email.message_from_bytes(email_body)
        if mail.get_content_maintype() != 'multipart':
            return
        for part in mail.walk():
            if part.get_content_maintype() != 'multipart' and part.get('Content-Disposition') is not None:
                filename, encoding = decode_header(part.get_filename())[0]
                if(encoding is None):
                    open(outputdir + '/' + filename, 'wb').write(part.get_payload(decode=True))
                else:
                    open(outputdir + '/' + filename.decode(encoding), 'wb').write(part.get_payload(decode=True))
        if(markRead):
            m.uid("STORE", emailid, "+FLAGS", "(\Seen)")**
    

    【讨论】:

    • 注意 decode_header 函数在 email.header 导入中找到。
    猜你喜欢
    • 2019-05-26
    • 2012-02-06
    • 1970-01-01
    • 2013-01-29
    • 2021-01-08
    • 1970-01-01
    • 1970-01-01
    • 2021-10-15
    • 2016-04-24
    相关资源
    最近更新 更多