【发布时间】:2021-01-20 02:33:13
【问题描述】:
我正在尝试抓取以下网站:voxnews.info
import requests
from bs4 import BeautifulSoup
from concurrent.futures import ThreadPoolExecutor
import pandas as pd
web='https://voxnews.info'
def main(req, num, web):
r = req.get(web+"/page/{}/".format(num))
soup = BeautifulSoup(r.content, 'html.parser')
goal = [(x.time.text, x.h1.a.get_text(strip=True), x.select_one("span.cat-links").get_text(strip=True), x.p.get_text(strip=True))
for x in soup.select("div.site-content")]
return goal
with ThreadPoolExecutor(max_workers=30) as executor:
with requests.Session() as req:
fs = [executor.submit(main, req, num) for num in range(1, 2)] # need to scrape all the webpages in the website
allin = []
for f in fs:
allin.extend(f.result())
df = pd.DataFrame.from_records(
allin, columns=["Date", "Title", "Category", "Content"])
print(df)
但是代码有两个问题:
- 第一个是我没有抓取所有页面(我目前将 1 和 2 放在范围内,但我需要所有页面);
- 它没有正确保存日期。
如果可以看看代码并告诉我如何改进它以解决这两个问题,那就太棒了。
【问题讨论】:
-
我假设您想从页面上的所有文章中抓取数据?
-
你好守门员 1998,是的,没错。我需要在所有页面中抓取每篇文章的日期标题类别和内容。
-
好的,我会处理一些事情。我首先要说的是,这些文章没有我可以看到的与之关联的 p 标签。您希望日期采用什么格式?
-
感谢守门员 1998。我已尽力写下一些可行的代码,但不幸的是,它并没有达到我的预期。感谢您的宝贵时间和帮助
标签: python web-scraping beautifulsoup