【问题标题】:HTML Link parsing using BeautifulSoup使用 BeautifulSoup 解析 HTML 链接
【发布时间】:2016-01-28 10:20:45
【问题描述】:

这是我的 Python 代码,我用它来从作为参数发送的 页面链接 中提取 特定 HTML。我正在使用BeautifulSoup。此代码有时可以正常工作,有时会卡住!

import urllib
from bs4 import BeautifulSoup

rawHtml = ''
url = r'http://iasexamportal.com/civilservices/tag/voice-notes?page='
for i in range(1, 49):  
    #iterate url and capture content
    sock = urllib.urlopen(url+ str(i))
    html = sock.read()  
    sock.close()
    rawHtml += html
    print i

我在这里打印循环变量以找出卡住的位置。它告诉我它在任何循环序列中随机卡住。

soup = BeautifulSoup(rawHtml, 'html.parser')
t=''
for link in soup.find_all('a'):
    t += str(link.get('href')) + "</br>"
    #t += str(link) + "</br>"
f = open("Link.txt", 'w+')
f.write(t)
f.close()

可能是什么问题。是 socket 配置的问题还是其他问题。

这是我得到的错误。我检查了这些链接 - python-gaierror-errno-11004,ioerror-errno-socket-error-errno-11004-getaddrinfo-failed 以获得解决方案。但我没有发现它有多大帮助。

 d:\python>python ext.py
Traceback (most recent call last):
  File "ext.py", line 8, in <module>
    sock = urllib.urlopen(url+ str(i))
  File "d:\python\lib\urllib.py", line 87, in urlopen
    return opener.open(url)
  File "d:\python\lib\urllib.py", line 213, in open
    return getattr(self, name)(url)
  File "d:\python\lib\urllib.py", line 350, in open_http
    h.endheaders(data)
  File "d:\python\lib\httplib.py", line 1049, in endheaders
    self._send_output(message_body)
  File "d:\python\lib\httplib.py", line 893, in _send_output
    self.send(msg)
  File "d:\python\lib\httplib.py", line 855, in send
    self.connect()
  File "d:\python\lib\httplib.py", line 832, in connect
    self.timeout, self.source_address)
  File "d:\python\lib\socket.py", line 557, in create_connection
    for res in getaddrinfo(host, port, 0, SOCK_STREAM):
IOError: [Errno socket error] [Errno 11004] getaddrinfo failed

当我在我的个人笔记本电脑上运行它时,它运行得非常好。但是当我在 Office Desktop 上运行它时会出错。另外,我的 Python 版本是 2.7。希望这些信息对您有所帮助。

【问题讨论】:

  • 你说的卡住了是什么意思?是否出现错误并且您有堆栈跟踪?还是只是程序意外挂了很长时间?
  • 似乎对我有用。请解释“卡住”...

标签: python url beautifulsoup filereader filewriter


【解决方案1】:

最后,伙计们......它成功了!当我在其他 PC 上检查时,相同的脚本也有效。所以问题可能是因为我的办公室桌面的防火墙设置或代理设置。阻止了这个网站。

【讨论】:

    猜你喜欢
    • 2013-03-10
    • 1970-01-01
    • 2012-12-13
    • 2010-09-12
    • 2019-06-13
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多