【问题标题】:How to scrape websites based on the site's title using python?如何使用python根据网站标题抓取网站?
【发布时间】:2019-09-22 21:06:48
【问题描述】:

我正在为包含特定标题的网站抓取网站。 我将如何制作它,例如,检查“example.com/xxxxxxxxxx”,其中“x”是一个随机数,如果它有标题 404?

【问题讨论】:

  • 你的代码在哪里?

标签: python http screen-scraping


【解决方案1】:

这会找到页面的标题:

import requests
from lxml.html import fromstring

def Get_PageTitle(url):
    req = requests.get(url)
    tree = fromstring(req.content)
    title = tree.findtext('.//title')
    return title


url = "http://www.google.com"
title = Get_PageTitle(url)

if "404" in title:
    #title has 404
    print("Title has 404 in it")

else:
    #no 404 in title
    pass

编辑:

上面的代码检查标题是否有 404 in。如果您想知道标题是否为 404,请使用以下代码:

import requests
from lxml.html import fromstring

def Get_PageTitle(url):
    req = requests.get(url)
    tree = fromstring(req.content)
    title = tree.findtext('.//title')
    return title


url = "http://www.google.com"
title = Get_PageTitle(url)

if "404" is title:
    #title is 404
    print("Title is 404 in it")
    print(title)

else:
    #title is not 404
    pass

How to get page title in requests

【讨论】:

  • 我不是在找网站代码,我是在找页面标题
  • 好的...这将检查标题以查看其中是否包含 404...这不是您想要的吗?
猜你喜欢
  • 2018-08-02
  • 1970-01-01
  • 2020-09-28
  • 2016-05-27
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多