【问题标题】:How to scrape HTML from a javascript webpage using Python如何使用 Python 从 javascript 网页中抓取 HTML
【发布时间】:2020-06-29 05:45:50
【问题描述】:

试图解析 html 以便从嵌套在标签内的标签中获取数据,但是当我美化时,我得到了 javascript。如何从此 javascript 中获取信息?我如何把它变成html?有没有更好的方法来获取这些信息?这是我的第一个问题,如果我犯了任何错误,我深表歉意。谢谢。

这是我的代码:

from bs4 import BeautifulSoup as bs
import requests

html = requests.get(url)
soup = bs(html.content, 'html.parser')
print(soup.prettify())

响应是: 看起来像字节/字符串的预美化代码,后跟

<html>
<head>
</head>
<script language="javascript">
var strUrl = window.location.href;


if (strUrl.indexOf("modisoftinc.com") > 0)
    window.location.replace("https://www.modisoftinc.com/home.html");
if (strUrl.indexOf("www.modisoftinc.com") > 0)
    window.location.replace("https://www.modisoftinc.com/home.html");
if (strUrl.indexOf("http://modisoftinc.com") > 0)
    window.location.replace("https://www.modisoftinc.com/home.html");
if (strUrl.indexOf("www.modisoftinc.com") > 0)
    window.location.replace("https://www.modisoftinc.com/home.html");


if (strUrl.indexOf("echecks.modisoftinc.com") > 0)
    window.location.replace("https://echecks.modisoftinc.com/Account/Logon");


if (strUrl.indexOf("pos.modisoftinc.com") > 0)
    window.location.replace("https://pos.modisoftinc.com/Account/Logon");


if (strUrl.indexOf("clock.modisoftinc.com") > 0)
    window.location.replace("https://clock.modisoftinc.com/Account/Logon");


if (strUrl.indexOf("admin11.modisoftinc.com") > 0)
    window.location.replace("https://admin11.modisoftinc.com/Account/Logon");




if (strUrl.indexOf("modisoft.com") > 0)
    window.location.replace("https://www.modisoft.com/home.html");
if (strUrl.indexOf("www.modisoft.com") > 0)
    window.location.replace("https://www.modisoft.com/home.html");
if (strUrl.indexOf("http://modisoft.com") > 0)
    window.location.replace("https://www.modisoft.com/home.html");
if (strUrl.indexOf("www.modisoft.com") > 0)
    window.location.replace("https://www.modisoft.com/home.html");


if (strUrl.indexOf("echecks.modisoft.com") > 0)
    window.location.replace("https://echecks.modisoft.com/Account/Logon");

if (strUrl.indexOf("app.modisoft.com") > 0)
    window.location.replace("https://app.modisoft.com/Account/Logon");

if (strUrl.indexOf("app1.modisoft.com") > 0)
    window.location.replace("https://app1.modisoft.com/Account/Logon");

if (strUrl.indexOf("app2.modisoft.com") > 0)
    window.location.replace("https://app2.modisoft.com/Account/Logon");

if (strUrl.indexOf("pos.modisoft.com") > 0)
    window.location.replace("https://pos.modisoft.com/Account/Logon");

if (strUrl.indexOf("clock.modisoft.com") > 0)
    window.location.replace("https://clock.modisoft.com/Account/Logon");

    if (strUrl.indexOf("admin11.modisoft.com") > 0)
    window.location.replace("https://admin11.modisoft.com/Account/Logon");



if (strUrl.indexOf("modisoftrewards.com") > 0)
    window.location.replace("https://www.modisoftrewards.com/index.html");
if (strUrl.indexOf("www.modisoftrewards.com") > 0)
    window.location.replace("https://www.modisoftrewards.com/index.html");
if (strUrl.indexOf("http://modisoftrewards.com") > 0)
    window.location.replace("https://www.modisoftrewards.com/index.html");
if (strUrl.indexOf("www.modisoftrewards.com") > 0)
    window.location.replace("https://www.modisoftrewards.com/index.html");






   if (strUrl.indexOf("localhost") > 0)
       window.location.replace("Account/Logon");
</script>
<body>
</body>
</html>

【问题讨论】:

  • 不能转成html。根据您想要从页面中获得的内容,您需要自动化浏览器以让 javascript 在页面上运行,然后获取您想要的内容,或者使用网络选项卡/网络监控工具查看您想要的内容是否可从另一个 uri 获得通过 xhr 并向该端点发出请求

标签: javascript python html linux screen-scraping


【解决方案1】:

如何从这个 javascript 中获取信息?怎么转成html?

是的,您需要一个浏览器自动化(selenium、无头 Chrome)来执行现场 JS。然后,JS 用缺失的数据填充 HTML。 例如:

  1. https://webscraping.pro/javascript-rendering-library-for-scraping-javascript-sites/

  2. https://webscraping.pro/java-library-to-scrape-linkedin-its-data-affiliates/

破解

在某些情况下,您可能会使用a bare coding(python、php)来模仿 JS 请求(通常是 XHR/Ajax)并获取缺失的信息。例如。 Scrape a JS Lazy load page by Python requests

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2019-01-20
    • 1970-01-01
    • 2011-12-24
    相关资源
    最近更新 更多