【问题标题】:How to scrape all the home page text content of a website?如何抓取一个网站的所有首页文字内容?
【发布时间】:2020-06-13 17:17:30
【问题描述】:

所以我是网页抓取的新手,我想只抓取主页的所有文本内容。

这是我的代码,但它现在可以正常工作了。

from bs4 import BeautifulSoup
import requests


website_url = "http://www.traiteurcheminfaisant.com/"
ra = requests.get(website_url)
soup = BeautifulSoup(ra.text, "html.parser")

full_text = soup.find_all()

print(full_text)

当我打印“full_text”时,它给了我很多 html 内容,但不是全部,当我 ctrl + f " traiteurcheminfaisant@hotmail.com" 主页上的电子邮件地址时(页脚) 在全文中找不到。

感谢您的帮助!

【问题讨论】:

  • 如果您打印(ra.text) 或(soup.text),您将获得包括电子邮件地址在内的完整html。我不确定为什么 BS4 没有返回电子邮件地址,但我猜这与 BS4 find_function 的工作方式有关。

标签: python web-scraping data-mining


【解决方案1】:

快速浏览一下您试图从中抓取的网站让我怀疑在通过请求模块发送简单的获取请求时并非所有内容都已加载。换句话说,网站上的某些组件(例如您提到的页脚)似乎是通过 Javascript 异步加载的。

如果是这种情况,您可能需要使用某种自动化工具导航到页面,等待它加载,然后解析完全加载的源代码。为此,最常用的工具是 Selenium。第一次设置可能有点棘手,因为您还需要为您想使用的任何浏览器安装单独的 web 驱动程序。也就是说,我上次设置它非常容易。这是一个粗略的示例,说明这对您来说可能是什么样子(一旦您正确设置了 Selenium):

from bs4 import BeautifulSoup
from selenium import webdriver

import time

driver = webdriver.Firefox(executable_path='/your/path/to/geckodriver')
driver.get('http://www.traiteurcheminfaisant.com')
time.sleep(2)

source = driver.page_source
soup = BeautifulSoup(source, 'html.parser')

full_text = soup.find_all()

print(full_text)

【讨论】:

    【解决方案2】:

    我之前没有使用过 BeatifulSoup,但请尝试使用 urlopen。这会将网页存储为字符串,您可以使用它来查找电子邮件。

    from urllib.request import urlopen
    
    try:
        response = urlopen("http://www.traiteurcheminfaisant.com")
        html = response.read().decode(encoding = "UTF8", errors='ignore')
        print(html.find("traiteurcheminfaisant@hotmail.com"))
    except:
        print("Cannot open webpage")
    
    
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-05-08
      • 1970-01-01
      • 1970-01-01
      • 2011-03-12
      • 2018-05-31
      • 2010-10-09
      相关资源
      最近更新 更多