【问题标题】:Correctly implementing Python Multiprocessing正确实现 Python 多处理
【发布时间】:2013-02-25 19:36:40
【问题描述】:

我正在阅读多处理教程:http://pymotw.com/2/multiprocessing/basics.html

我写了下面的脚本作为练习。该脚本似乎正在运行,我确实看到在 taskmgr 中运行了 5 个新的 python 进程。但是,我的 print 语句会输出多次搜索的同一个文件夹。

我怀疑我不是在不同进程之间拆分工作,而是将整个工作负载分配给每个进程。我很确定我做错了什么并且效率低下。有人可以指出我的错误吗?

到目前为止我所拥有的:

def email_finder(msg_id):
    for folder in os.listdir(sample_path):
        print "Searching through folder: ", folder
        folder_path = sample_path + '\\' + folder
        for file in os.listdir(os.listdir(folder_path)):
            if file.endswith('.eml'):
                file_path = folder_path + '\\' + file
                email_obj = email.message_from_file(open(file_path))
                if msg_id in email_obj.as_string().lower()
                    shutil.copy(file_path, tmp_path + '\\' + file)
                    return 'Found: ', file_path
    else:
        return 'Not Found!'

def worker():
    msg_ids = cur.execute("select msg_id from my_table").fetchall()
    for i in msg_ids:
        msg_id = i[0].encode('ascii')
        if msg_id != '':
            email_finder(msg_id)
    return

if __name__ == '__main__':
    jobs = []
    for i in range(5):
        p = multiprocessing.Process(target=worker)
        jobs.append(p)
        p.start()

【问题讨论】:

    标签: python python-2.7 multiprocessing


    【解决方案1】:

    您的每个子流程都有自己的光标,因此会遍历整个 ID 集。

    您需要从数据库中读取一次msg_ids,然后将其生成到子进程中,而不是让每个子进程自行查询。

    【讨论】:

    • 啊,有道理 :) 谢谢!
    • 完成 :) UI 迫使我在接受之前等待一分钟。
    猜你喜欢
    • 2018-06-30
    • 2021-07-31
    • 2017-07-16
    • 1970-01-01
    • 2016-06-05
    • 1970-01-01
    • 2020-08-15
    • 2021-12-28
    • 1970-01-01
    相关资源
    最近更新 更多