【问题标题】:How to print Google Search results properly with bs4?如何使用 bs4 正确打印 Google 搜索结果?
【发布时间】:2021-06-03 05:11:58
【问题描述】:

我有一个工作代码,它首先打印搜索标题,然后打印 url,但它会在网站标题之间打印很多 url。但是如何以这样的格式打印它们并避免每次打印相同的 url 10 次:

1) Title url
2) Title url
and so on... 

我的代码:

search = input("Search:")

page = requests.get(f"https://www.google.com/search?q=" + search)

soup = BeautifulSoup(page.content, "html5lib")

links = soup.findAll("a")

heading_object = soup.find_all('h3')

for info in heading_object:
    x = info.getText()
    print(x)
    for link in links:
        link_href = link.get('href')
        if "url?q=" in link_href:
            y = (link.get('href').split("?q=")[1].split("&sa=U")[0])
            print(y)

【问题讨论】:

  • 我不希望这仅适用于 beautifilsoup 和简单的 get 请求,因为 google 使用 JS 呈现了很多页面。你可能想看看Search API。
  • 代码为:heading_object 中的信息:x = info.getText() 链接中的链接:link_href = link.get('href') if "url?q=" in link_href: y = (link.get('href').split("?q=")[1].split("&sa=U")[0]) print(x, y) 给我需要的结果,但打印出来的除外来自 Google 的每个结果大约有 10 个副本。
  • 始终将代码、数据和完整的错误消息作为文本(不是截图,不是链接)放在有问题的地方(不是评论)。
  • 您应该以不同的方式搜索它 - 首先找到同时保留 title 和 url 的对象,然后在此对象中搜索单个 title 和 url 以将其作为对。最终你应该使用zip(heading_object, links) 来创建对,但如果页面上的某些项目(标题或链接)为空,它可能会给出错误的结果,因为它会将其他项目移动到这个地方。
  • 我编辑了代码。问题是它会打印出在 Google 搜索页面上找到的所有 url。

标签: python beautifulsoup request


【解决方案1】:

如果您获得单独的标题和链接,则可以使用zip() 将它们成对分组

for info, link in zip(heading_object, links):
    info = info.getText()

    link = link.get('href')
    if "?q=" in link:
        link = link.split("?q=")[1].split("&sa=U")[0]

    print(info, link)

但是当页面上不存在某些标题或链接时,这可能会出现问题,因为这样会创建错误的对。它将标题与下一个元素的链接配对。您应该搜索同时保留标题和链接的元素,并在每个元素中搜索单个标题和单个链接以创建对。如果没有标题或链接,那么您可以输入一些默认值,它不会创建错误的对。

【讨论】:

  • 谢谢,它现在在一行上打印标题和链接。但唯一的问题是它打印出的不是完整的网址,例如“/?sa=X&ved=0ahUKEwirhbTIh5fvAhVmyDgGHYGlCscQOwgC”或“/?output=search&ie=UTF-8&sa=X&ved=0ahUKEwirhbTIh5fvAhVmyDgGHYGlCscQPAgE”。我不确定如何解决它。
  • 我不明白 - 如果您需要完整,请删除 if。你会看到你真正从服务器得到了什么。它可能会为不同的设备发送不同的 HTML(特别是如果他们使用错误的标头 user-agent 并且他们不使用 JavaScript)
【解决方案2】:

你正在寻找这个:

for result in soup.select('.yuRUbf'):
  title = result.select_one('.DKV0Md').text
  url = result.a['href']
  print(f'{title}, {url}\n') # prints TITLE, URL followed by a new line.

如果您使用的是f-string,那么appropriate way 就是这样使用它:

page = requests.get(f"https://www.google.com/search?q=" + search) # not proper f-string
page = requests.get(f"https://www.google.com/search?q={search}")  # proper f-string

代码:

import requests, lxml
from bs4 import BeautifulSoup

headers = {
  'User-agent':
  "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.102 Safari/537.36 Edge/18.19582"
}

params = {
  "q": "python memes",
  "hl": "en"
}

soup = BeautifulSoup(requests.get('https://www.google.com/search', headers=headers, params=params).text, 'lxml')

for result in soup.select('.yuRUbf'):
  title = result.select_one('.DKV0Md').text
  url = result.a['href']
  print(f'{title}, {url}\n')

--------
'''
35 Funny And Best Python Programming Memes - CodeItBro, https://www.codeitbro.com/funny-python-programming-memes/

ML Memes (@python.memes_) • Instagram photos and videos, https://www.instagram.com/python.memes_/?hl=en

28 Python Memes ideas - Pinterest, https://in.pinterest.com/codeitbro/python-memes/
'''

或者,您可以使用来自 SerpApi 的 Google Organic Results API 来做同样的事情。这是一个带有免费计划的付费 API。

其中一个区别是您只需要遍历 JSON,而不是弄清楚如何抓取内容。

要集成的代码:

from serpapi import GoogleSearch
import os

params = {
  "api_key": os.getenv("API_KEY"),
  "engine": "google",
  "q": "python memes",
  "hl": "en"
}

search = GoogleSearch(params)
results = search.get_dict()

for result in results['organic_results']:
  title = result['title']
  url = result['link']
  print(f'{title}, {url}\n')

-------
'''
35 Funny And Best Python Programming Memes - CodeItBro, https://www.codeitbro.com/funny-python-programming-memes/

ML Memes (@python.memes_) • Instagram photos and videos, https://www.instagram.com/python.memes_/?hl=en

28 Python Memes ideas - Pinterest, https://in.pinterest.com/codeitbro/python-memes/
'''

免责声明,我为 SerpApi 工作。

【讨论】:

    猜你喜欢
    • 2021-10-28
    • 2021-08-23
    • 1970-01-01
    • 1970-01-01
    • 2014-08-03
    • 2019-08-04
    • 2017-09-29
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多