【问题标题】:How to download multiple files simultaneously and trigger specific actions for each one done?如何同时下载多个文件并为每个完成触发特定操作?
【发布时间】:2018-09-07 11:43:35
【问题描述】:

我在尝试实现的功能方面需要帮助,不幸的是我对多线程不是很满意。

我的脚本从 Internet 下载 4 个不同的文件,并为每个文件调用一个专用函数,然后全部保存。 问题是我是一步一步做的,因此我必须等待每次下载完成才能继续下一个。

我知道我应该做些什么来解决这个问题,但我没有成功编写代码。

实际行为:

url_list = [Url1, Url2, Url3, Url4]
files_list = []

files_list.append(downloadFile(Url1))
handleFile(files_list[-1], type=0)
...
files_list.append(downloadFile(Url4))
handleFile(files_list[-1], type=3)
saveAll(files_list)

需要的行为:

url_list = [Url1, Url2, Url3, Url4]
files_list = []

for url in url_list:
    callThread(files_list.append(downloadFile(url)),             # function
               handleFile(files_list[url.index], type=url.index) # trigger
    #use a thread for downloading
    #once file is downloaded, it triggers his associated function
#wait for all files to be treated
saveAll(files_list)

感谢您的帮助!

【问题讨论】:

    标签: python multithreading python-2.7


    【解决方案1】:

    典型的做法是将IO繁重的部分,如通过互联网获取数据和数据处理放在同一个函数中:

    import random
    import threading
    import time
    from concurrent.futures import ThreadPoolExecutor
    
    import requests
    
    
    def fetch_and_process_file(url):
        thread_name = threading.currentThread().name
    
        print(thread_name, "fetch", url)
        data = requests.get(url).text
    
        # "process" result
        time.sleep(random.random() / 4)  # simulate work
        print(thread_name, "process data from", url)
    
        result = len(data) ** 2
        return result
    
    
    threads = 2
    urls = ["https://google.com", "https://python.org", "https://pypi.org"]
    
    executor = ThreadPoolExecutor(max_workers=threads)
    with executor:
        results = executor.map(fetch_and_process_file, urls)
    
    print()
    print("results:", list(results))
    

    输出:

    ThreadPoolExecutor-0_0 fetch https://google.com
    ThreadPoolExecutor-0_1 fetch https://python.org
    ThreadPoolExecutor-0_0 process data from https://google.com
    ThreadPoolExecutor-0_0 fetch https://pypi.org
    ThreadPoolExecutor-0_0 process data from https://pypi.org
    ThreadPoolExecutor-0_1 process data from https://python.org
    

    【讨论】:

    • 谢谢,这似乎是我需要的,一旦我能够正确实施它,我会将您的帖子标记为答案。
    • 我没有找到任何与我的 python 版本(2.7)兼容的 concurrent.features 版本,并且我不能使用 3.x 来实现这个功能。您是否有其他解决方案或可能已被反向移植?
    • 我成功实现了它,但我必须警告异步行为。例如,我最初习惯将这些文件直接放在一个 zipfile 中,但在 python 2.7 中它不是线程安全的。然后我把它们放在一个列表中,并一个一个地压缩。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-12-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多