【发布时间】:2019-09-21 17:09:42
【问题描述】:
我正在尝试在网站上搜索游戏名称以及其他项目,但为了简洁起见,仅提供游戏名称。
我曾尝试同时使用 selenium 和 beautiful soup 来获取标题,但无论我做什么,我似乎都无法获得所有 9 月的版本。事实上,我也获得了一些 8 月的游戏名称。我认为这与网站没有尽头的事实有关。我将如何仅获得 9 月的冠军头衔?以下是我使用的代码,我尝试过使用 Scrolling,但我认为我不明白如何正确使用它。
编辑:我的目标是能够通过更改几行代码最终得到每个月。
from selenium import webdriver
from bs4 import BeautifulSoup
titles = []
chromedriver = 'C:/Users/Chase The Great/Desktop/Podcast/chromedriver.exe'
driver = webdriver.Chrome(chromedriver)
driver.get('https://www.releases.com/l/Games/2019/9/')
res = driver.execute_script("return document.documentElement.outerHTML")
driver.quit()
soup = BeautifulSoup(res, 'lxml')
for title in soup.find_all(class_= 'calendar-item-title'):
titles.append(title.text)
预计我将获得 133 个标题,而我将获得一些 8 月的标题以及仅部分标题:
['SubaraCity', 'AER - Memories of Old', 'Vambrace: Cold Soul', 'Agent A: A Puzzle in Disguise', 'Bubsy: Paws on Fire!', 'Grand Brix Shooter', 'Legend of the Skyfish', 'Vambrace: Cold Soul', 'Obakeidoro!', 'Pokemon Masters', 'Decay of Logos', 'The Lord of the Rings: Adventure ...', 'Heave Ho', 'Newt One', 'Blair Witch', 'Bulletstorm: Duke of Switch Edition', 'The Ninja Saviors: Return of the ...', 'Re:Legend', 'Risk of Rain 2', 'Decay of Logos', 'Unlucky Seven', 'The Dark Pictures Anthology: Man ...', 'Legend of the Skyfish', 'Astral Chain', 'Torchlight II', 'Final Fantasy VIII Remastered', 'Catherine: Full Body', 'Root Letter: Last Answer', 'Children of Morta', 'Himno', 'Spyro Reignited Trilogy', 'RemiLore: Lost Girl in the Lands ...', 'Divinity: Original Sin 2 - Defini...', 'Monochrome Order', 'Throne Quest Deluxe', 'Super Kirby Clash', 'Himno', 'Post War Dreams', 'The Long Journey Home', 'Spice and Wolf VR', 'WRC 8', 'Fantasy General II', 'River City Girls', 'Headliner: NoviNews', 'Green Hell', 'Hyperforma', 'Atomicrops', 'Remothered: Tormented Fathers']
【问题讨论】:
-
“网站没有结尾”是指“无限滚动”,当您滚动到屏幕底部时会加载新的内容页面,对吧?
-
您想获得多远的时间?本月和上月?这就是这类网站的挑战。你什么时候停下来?如果我们知道,我认为我们可以为您提供更好的帮助。
-
我只想要 9 月的内容。我的目标是每个月,我只想更改一行代码,以便下个月是我希望的。
-
这是很好的信息,但请将其添加到问题文本而不是评论中。其他人不一定会读cmets,所以他们会错过这一点。
标签: python selenium web-scraping beautifulsoup infinite-scroll