【问题标题】:BeautifulSoup returns Null resultsBeautifulSoup 返回 Null 结果
【发布时间】:2020-04-23 00:47:40
【问题描述】:

我对使用 beauifulsoup 很陌生,我正在尝试使用下面的代码从网站上抓取文本。 但是,find_all 什么也不返回。

import bs4 as bs
import urllib.request
source = urllib.request.urlopen('https://beta.regulations.gov/document/USCIS-2019-0010-9175').read()
soup = BeautifulSoup(page.content,'html.parser')
text = soup.find_all(class_="px-2")
print(text)

html for website

【问题讨论】:

标签: python beautifulsoup


【解决方案1】:

如 cmets 中所述,数据是通过 Javascript 动态加载的。但是当你打开 Firefox/Chrome 的网络标签时,你可以看到数据来自哪里:

import requests

url = 'https://beta.regulations.gov/document/USCIS-2019-0010-9175'
ajax_url = 'https://beta.regulations.gov/api/documentdetails/{}'

document_id = url.split('/')[-1]
data = requests.get(ajax_url.format(document_id)).json()

# from pprint import pprint # <-- uncoment to see all data
# pprint(data)

print(data['data']['attributes']['content'])

打印:

Rescind the increase in fees. This is draconian. For all intents and purposes, denying access to this information will prevent many Americans from knowing where they came from. This is an outrage. This is not the mark of a democracy. I strongly disagree with this fee increase

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-11-21
    • 2015-05-18
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多