【问题标题】:Python Process blocked by urllib2Python 进程被 urllib2 阻止
【发布时间】:2010-01-26 02:31:39
【问题描述】:

我设置了一个进程来读取传入 url 下载队列,但是当 urllib2 打开连接时系统挂起。

import urllib2, multiprocessing
from threading import Thread
from Queue import Queue
from multiprocessing import Queue as ProcessQueue, Process

def download(url):
    """Download a page from an url.
    url [str]: url to get.
    return [unicode]: page downloaded.
    """
    if settings.DEBUG:
        print u'Downloading %s' % url
    request = urllib2.Request(url)
    response = urllib2.urlopen(request)
    encoding = response.headers['content-type'].split('charset=')[-1]
    content = unicode(response.read(), encoding)
    return content

def downloader(url_queue, page_queue):
    def _downloader(url_queue, page_queue):
        while True:
            try:
                url = url_queue.get()
                page_queue.put_nowait({'url': url, 'page': download(url)})
            except Exception, err:
                print u'Error downloading %s' % url
                raise err
            finally:
                url_queue.task_done()

    ## Init internal workers
    internal_url_queue = Queue()
    internal_page_queue = Queue()
    for num in range(multiprocessing.cpu_count()):
        worker = Thread(target=_downloader, args=(internal_url_queue, internal_page_queue))
        worker.setDaemon(True)
        worker.start()

    # Loop waiting closing
    for url in iter(url_queue.get, 'STOP'):
        internal_url_queue.put(url)

    # Wait for closing
    internal_url_queue.join()

# Init the queues
url_queue = ProcessQueue()
page_queue = ProcessQueue()

# Init the process
download_worker = Process(target=downloader, args=(url_queue, page_queue))
download_worker.start()

我可以从另一个模块添加 url,当我需要时,我可以停止进程并等待进程关闭。

import module

module.url_queue.put('http://foobar1')
module.url_queue.put('http://foobar2')
module.url_queue.put('http://foobar3')
module.url_queue.put('STOP')
downloader.download_worker.join()

问题是,当我使用 urlopen ("response = urllib2.urlopen(request)") 时,它仍然全部被阻止。

如果我调用 download() 函数或者当我只使用没有进程的线程时没有问题。

【问题讨论】:

    标签: python multithreading urllib2 multiprocess


    【解决方案1】:

    这里的问题不是 urllib2,而是多处理模块的使用。在 Windows 下使用多处理模块时,您不能使用在导入模块时立即运行的代码 - 而是将内容放在主模块中的 if __name__=='__main__' 块内。参见“安全导入主模块”here。

    对于您的代码,请在下载器模块中进行以下更改:

    #....
    def start():
        global download_worker
        download_worker = Process(target=downloader, args=(url_queue, page_queue))
        download_worker.start()
    

    在主模块中:

    import module
    if __name__=='__main__':
        module.start()
        module.url_queue.put('http://foobar1')
        #....
    

    因为你没有这样做,所以每次启动子进程都会再次运行主代码并启动另一个进程,从而导致挂起。

    【讨论】:

    • 我不使用 Windows,但您建议使用 start() 函数解决了这个问题。谢谢!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-01-04
    • 2014-07-12
    • 1970-01-01
    • 2016-02-24
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多