【发布时间】:2018-08-13 16:38:45
【问题描述】:
我正在尝试递归地抓取所有英文文章链接的维基百科 url。我想执行 n 的深度优先遍历,但由于某种原因,我的代码不会在每次传递时都重复出现。知道为什么吗?
def crawler(url, depth):
if depth == 0:
return None
links = bs.find("div",{"id" : "bodyContent"}).findAll("a" , href=re.compile("(/wiki/)+([A-Za-z0-9_:()])+"))
print ("Level ",depth," ",url)
for link in links:
if ':' not in link['href']:
crawler("https://en.wikipedia.org"+link['href'], depth - 1)
这是对爬虫的调用
url = "https://en.wikipedia.org/wiki/Harry_Potter"
html = urlopen(url)
bs = BeautifulSoup(html, "html.parser")
crawler(url,3)
【问题讨论】:
-
我必须使用 beautifulsoup 来解决这个问题...这是我必须完成的任务...我无法使用数据转储
-
函数内部对url的请求在哪里?您必须在每次重复时发送请求。
标签: python beautifulsoup web-crawler