【问题标题】:How to simulate a button click in a request?如何模拟请求中的按钮单击?
【发布时间】:2020-06-06 17:41:04
【问题描述】:

请不要关闭这个问题 - 这不是重复的。我需要使用 Python 请求而不是 Selenium 来单击按钮,如 here

我正在尝试抓取Reverso Context translation examples page。我有一个问题:我只能获得 20 个示例,然后我需要在页面上存在很多次时单击“显示更多示例”按钮才能获得完整的结果列表。可以简单地使用网络浏览器完成,但是如何使用 Python Requests 库来完成呢?

我查看了按钮的 HTML 代码,但找不到 onclick 属性来查看附加到它的 JS 脚本,我不明白我需要发送什么请求:

<button id="load-more-examples" class="button load-more " data-default-size="14px">Display more examples</button>

这是我的 Python 代码:

from bs4 import BeautifulSoup
import requests
import re


with requests.Session() as session:  # Create a Session
    # Log in
    login_url = 'https://account.reverso.net/login/context.reverso.net/it?utm_source=contextweb&utm_medium=usertopmenu&utm_campaign=login'
    session.post(login_url, "Email=reverso.scraping@yahoo.com&Password=sample",
           headers={"User-Agent": "Mozilla/5.0", "content-type": "application/x-www-form-urlencoded"})

    # Get the HTML
    html_text = session.get("https://context.reverso.net/translation/russian-english/cat", headers={"User-Agent": "Mozilla/5.0"}).content

    # And scrape it
    for word_pair in BeautifulSoup(html_text).find_all("div", id=re.compile("^OPENSUBTITLES")):
        print(word_pair.find("div", class_="src ltr").text.strip(), "=", word_pair.find("div", class_="trg ltr").text.strip())

注意: 您需要登录,否则只会显示前10个示例,不会显示按钮。您可以使用此真实身份验证数据:
电子邮件: reverso.scraping@yahoo.com
密码:示例

【问题讨论】:

  • 这能回答你的问题吗? invoking onclick event with beautifulsoup python
  • 你不能用请求来做到这一点stackoverflow.com/a/37167063/8619959
  • 它可以简单地使用网络浏览器完成,但是我如何使用 Python Requests 库来做到这一点? 你不能,Requests 不会执行 JavaScript 或类似的东西.
  • 这能回答你的问题吗? "Clicking" button with requests
  • 在某些情况下(例如这个例子),可以更详细地探索浏览器的行为:获取它发送的请求并尝试使用 Python 请求做同样的事情 当然,但我不会称之为模拟按钮按下。

标签: python web-scraping beautifulsoup python-requests urllib


【解决方案1】:

这是一个解决方案,它使用requests 获取所有例句,并使用BeautifulSoup 从中删除所有HTML标签:

from bs4 import BeautifulSoup
import requests
import json


headers = {
    "Connection": "keep-alive",
    "Accept": "application/json, text/javascript, */*; q=0.01",
    "X-Requested-With": "XMLHttpRequest",
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36",
    "Content-Type": "application/json; charset=UTF-8",
    "Content-Length": "96",
    "Origin": "https://context.reverso.net",
    "Sec-Fetch-Site": "same-origin",
    "Sec-Fetch-Mode": "cors",
    "Referer": "https://context.reverso.net/^%^D0^%^BF^%^D0^%^B5^%^D1^%^80^%^D0^%^B5^%^D0^%^B2^%^D0^%^BE^%^D0^%^B4/^%^D0^%^B0^%^D0^%^BD^%^D0^%^B3^%^D0^%^BB^%^D0^%^B8^%^D0^%^B9^%^D1^%^81^%^D0^%^BA^%^D0^%^B8^%^D0^%^B9-^%^D1^%^80^%^D1^%^83^%^D1^%^81^%^D1^%^81^%^D0^%^BA^%^D0^%^B8^%^D0^%^B9/cat",
    "Accept-Encoding": "gzip, deflate, br",
    "Accept-Language": "ru-RU,ru;q=0.9,en-US;q=0.8,en;q=0.7",
}

data = {
    "source_text": "cat",
    "target_text": "",
    "source_lang": "en",
    "target_lang": "ru",
    "npage": 1,
    "mode": 0
}

npages = requests.post("https://context.reverso.net/bst-query-service", headers=headers, data=json.dumps(data)).json()["npages"]
for npage in range(1, npages + 1):
    data["npage"] = npage
    page = requests.post("https://context.reverso.net/bst-query-service", headers=headers, data=json.dumps(data)).json()["list"]
    for word in page:
        print(BeautifulSoup(word["s_text"]).text, "=", BeautifulSoup(word["t_text"]).text)

一开始,我收到了 Google Chrome DevTools 的请求:

  1. 按F12键进入并选择网络标签
  2. 点击了“显示更多示例”按钮
  3. 找到最后一个请求(“bst-query-service”)
  4. 右键单击它并选择 复制 > 复制为 cURL (cmd)

然后,我打开this online-tool,将复制的cURL插入到左侧的文本框中,并复制右侧的输出(使用Ctrl-C热键,否则可能不起作用) .

之后我将其插入 IDE 并:

  1. 删除了 cookies 字典 - 这里没有必要
  2. 重要提示:将data字符串重写为Python字典,并用json.dumps(data)包裹,否则返回一个空词列表的请求。
  3. 添加了一个脚本,该脚本:获取多次获取单词(“页面”)并创建了一个for 循环,该循环获取该次数的单词并在不使用 HTML 标记的情况下打印它们(使用 BeautifulSoup)

UPD:
对于那些访问该问题以了解如何使用Reverso Context(不仅仅是模拟其他网站上的按钮单击请求)的人,发布了Reverso API 的Python 包装器:Reverso-API。它可以做与上面相同的事情,但要简单得多:

from reverso_api.context import ReversoContextAPI


api = ReversoContextAPI("cat", "", "en", "ru")
for source, target in api.get_examples_pair_by_pair():
    print(highlight_example(source.text), "==", highlight_example(target.text))

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2019-12-24
    • 1970-01-01
    • 2017-09-30
    • 2016-09-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多