【问题标题】:Python and BeautifulSoup Opening pagesPython 和 BeautifulSoup 打开页面
【发布时间】:2015-12-21 15:16:16
【问题描述】:

我想知道如何使用 BeautifulSoup 在我的列表中打开另一个页面?我关注了this tutorial,但它并没有告诉我们如何打开列表中的另一个页面。另外,如何打开嵌套在类中的“a href”?

这是我的代码:

# coding: utf-8

import requests
from bs4 import BeautifulSoup

r = requests.get("")
soup = BeautifulSoup(r.content)
soup.find_all("a")

for link in soup.find_all("a"):
    print link.get("href")

    for link in soup.find_all("a"):
        print link.text

    for link in soup.find_all("a"):
        print link.text, link.get("href")

    g_data = soup.find_all("div", {"class":"listing__left-column"})

    for item in g_data:
        print item.contents

    for item in g_data:
        print item.contents[0].text
        print link.get('href')

    for item in g_data:
        print item.contents[0]

我正在尝试从每个企业的标题中收集 href,然后打开它们并抓取该数据。

【问题讨论】:

  • 首先,我不明白你在问什么。那么,也许你想看看the document。
  • 您需要让我们知道您希望抓取哪个页面。需要r = requests.get("http://www.yellowpages.com/") 之类的东西。
  • 我应该多解释一下,我想做的是在一个 div ect 中打开一个 href。 puu.sh/kmgxZ/15fc324654.png我想调用每个有链接的href并打开它们的页面然后开始报废

标签: python python-2.7 web-scraping beautifulsoup


【解决方案1】:

我仍然不确定您从哪里获取 HTML,但如果您尝试提取所有 href 标记,那么根据您发布的图像,以下方法应该可以工作:

import requests
from bs4 import BeautifulSoup

r = requests.get("<add your URL here>")
soup = BeautifulSoup(r.content)

for a_tag in soup.find_all('a', class_='listing-name', href=True):
    print 'href: ', a_tag['href']

通过将href=True 添加到find_all(),它确保只返回包含href 属性的a 元素,因此无需将其作为属性进行测试。

提醒您一下,您可能会发现一些网站会在一两次尝试后将您锁定,因为它们能够检测到您正在尝试通过脚本访问网站,而不是作为人。如果您觉得您没有得到正确的响应,我建议您打印您返回的 HTML,以确保它仍然如您所愿。

如果您想获取每个链接的 HTML,可以使用以下内容:

import requests
from bs4 import BeautifulSoup

# Configure this to be your first request URL
r = requests.get("http://www.mywebsite.com/search/")
soup = BeautifulSoup(r.content)

for a_tag in soup.find_all('a', class_='listing-name', href=True):
    print 'href: ', a_tag['href']

# Configure this to the root of the above website, e.g. 'http://www.mywebsite.com'
base_url = "http://www.mywebsite.com"

for a_tag in soup.find_all('a', class_='listing-name', href=True):
    print '-' * 60      # Add a line of dashes
    print 'href: ', a_tag['href']
    request_href = requests.get(base_url + a_tag['href'])
    print request_href.content

使用 Python 2.x 测试,对于 Python 3.x,请在打印语句中添加括号。

【讨论】:

  • 好吧,经过长时间阅读,我发现我想使用可以做到这一点的东西打开 href 或类。有人告诉我 Requests 可以做到这一点。因此,如果我收到在该页面上打开 href 的请求,然后用 BS 废弃该页面,它将起作用
  • 谢谢,他们在抓取网站和侧页中的页面方面并不多。只是很多刮一页的教程。对于 Python 和 Scraping,您会推荐什么书或教程系列。
  • 最重要的是要了解 HTML 的结构。然后你就会知道要寻找什么。
  • 您好 Martin,我现在已经获得了 HTML 并提取了数据,但现在正在寻找使用 beautifulsoup 的方法,我们可以有多个具有 BS 属性的类吗?例如,我将 request_href.content 变成了一个变量,现在想从中提取内容。我可以看到我无法添加诸如 newpage.findAll 之类的东西
  • 我建议您通读整个Beautifulsoup tutorial,它解释了一切。您可能想点击我的回答打勾,然后您可以提出一个新问题。
【解决方案2】:
  1. 我遇到了同样的问题,我想分享我的发现,因为我确实尝试了答案,但由于某些原因它不起作用,但经过一些研究,我发现了一些有趣的东西。

  2. 您可能需要找到“href”链接本身的属性: 在您的情况下,您将需要包含 href 链接的确切 class 链接,我在想="class":"listing__left-column" 并将其等同于变量say "全部”例如:

from bs4 import BeautifulSoup
all = soup.find_all("div", {"class":"listing__left-column"})
for item in all:
  for link in item.find_all("a"):
    if 'href' in link.attrs:
        a = link.attrs['href']
        print(a)
        print("")

我这样做了,我能够进入另一个嵌入主页的链接

【讨论】:

  • 你好 - 你把 URL 放在哪里了!?
  • r = requests.get("")
猜你喜欢
  • 1970-01-01
  • 2013-01-29
  • 1970-01-01
  • 1970-01-01
  • 2014-12-17
  • 2015-06-07
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多