【问题标题】:Scraping data from a website with Infinite Scroll?使用 Infinite Scroll 从网站抓取数据?
【发布时间】:2019-09-21 17:09:42
【问题描述】:

我正在尝试在网站上搜索游戏名称以及其他项目,但为了简洁起见,仅提供游戏名称。

我曾尝试同时使用 selenium 和 beautiful soup 来获取标题,但无论我做什么,我似乎都无法获得所有 9 月的版本。事实上,我也获得了一些 8 月的游戏名称。我认为这与网站没有尽头的事实有关。我将如何仅获得 9 月的冠军头衔?以下是我使用的代码,我尝试过使用 Scrolling,但我认为我不明白如何正确使用它。

编辑:我的目标是能够通过更改几行代码最终得到每个月。

from selenium import webdriver
from bs4 import BeautifulSoup

titles = []

chromedriver = 'C:/Users/Chase The Great/Desktop/Podcast/chromedriver.exe'
driver = webdriver.Chrome(chromedriver)
driver.get('https://www.releases.com/l/Games/2019/9/')
res = driver.execute_script("return document.documentElement.outerHTML")
driver.quit()
soup = BeautifulSoup(res, 'lxml')

for title in soup.find_all(class_= 'calendar-item-title'):
    titles.append(title.text)

预计我将获得 133 个标题,而我将获得一些 8 月的标题以及仅部分标题:

['SubaraCity', 'AER - Memories of Old', 'Vambrace: Cold Soul', 'Agent A: A Puzzle in Disguise', 'Bubsy: Paws on Fire!', 'Grand Brix Shooter', 'Legend of the Skyfish', 'Vambrace: Cold Soul', 'Obakeidoro!', 'Pokemon Masters', 'Decay of Logos', 'The Lord of the Rings: Adventure ...', 'Heave Ho', 'Newt One', 'Blair Witch', 'Bulletstorm: Duke of Switch Edition', 'The Ninja Saviors: Return of the ...', 'Re:Legend', 'Risk of Rain 2', 'Decay of Logos', 'Unlucky Seven', 'The Dark Pictures Anthology: Man ...', 'Legend of the Skyfish', 'Astral Chain', 'Torchlight II', 'Final Fantasy VIII Remastered', 'Catherine: Full Body', 'Root Letter: Last Answer', 'Children of Morta', 'Himno', 'Spyro Reignited Trilogy', 'RemiLore: Lost Girl in the Lands ...', 'Divinity: Original Sin 2 - Defini...', 'Monochrome Order', 'Throne Quest Deluxe', 'Super Kirby Clash', 'Himno', 'Post War Dreams', 'The Long Journey Home', 'Spice and Wolf VR', 'WRC 8', 'Fantasy General II', 'River City Girls', 'Headliner: NoviNews', 'Green Hell', 'Hyperforma', 'Atomicrops', 'Remothered: Tormented Fathers']

【问题讨论】:

  • “网站没有结尾”是指“无限滚动”,当您滚动到屏幕底部时会加载新的内容页面,对吧?
  • 您想获得多远的时间?本月和上月?这就是这类网站的挑战。你什么时候停下来?如果我们知道,我认为我们可以为您提供更好的帮助。
  • 我只想要 9 月的内容。我的目标是每个月,我只想更改一行代码,以便下个月是我希望的。
  • 这是很好的信息,但请将其添加到问题文本而不是评论中。其他人不一定会读cmets,所以他们会错过这一点。

标签: python selenium web-scraping beautifulsoup infinite-scroll


【解决方案1】:

在我看来,为了只获得 9 月,首先您只想获取 9 月的部分:

section = soup.find('section', {'class': 'Y2019-M9 calendar-sections'})

然后,一旦您获取 9 月的部分,就可以获取 <a> 标签中的所有标题,如下所示:

for title in section.find_all('a', {'class': ' calendar-item-title subpage-trigg'}):
    titles.append(title.text)

请注意,以前的都没有经过测试。

更新: 问题是每次您要加载页面时,它只会为您提供仅包含 24 个项目的第一部分,为了访问它们,您必须向下滚动(无限滚动)。 如果您打开浏览器开发工具,选择Network,然后选择XHR,您会注意到每次滚动并加载下一个“页面”时,都会有一个带有url 的请求,类似于:

https://www.releases.com/calendar/nextAfter?blockIndex=139&itemIndex=23&category=Games&regionId=us

我的猜测是,blockIndex 用于该月,itemIndex 用于加载的每个页面,如果您只寻找 9 月的月份,blockIndex 在该请求中将始终为139是为下一页获取下一个itemIndex,以便您可以构建下一个请求。 下一个itemIndex 将始终是上一个请求的最后一个itemIndex。

我确实制作了一个脚本,它只使用BeautifulSoup 来执行您想要的操作。自行决定使用它,有一些常量可以动态提取,但我认为这可以让您抢先一步:

import json

import requests
from bs4 import BeautifulSoup

DATE_CODE = 'Y2019-M9'
LAST_ITEM_FIRST_PAGE = f'calendar-item col-xs-6 to-append first-item calendar-last-item {DATE_CODE}-None'
LAST_ITEM_PAGES = f'calendar-item col-xs-6 to-append calendar-last-item {DATE_CODE}-None'
INITIAL_LINK = 'https://www.releases.com/l/Games/2019/9/'
BLOCK = 139
titles = []


def get_next_page_link(div: BeautifulSoup):
    index = div['item-index']
    return f'https://www.releases.com/calendar/nextAfter?blockIndex={BLOCK}&itemIndex={index}&category=Games&regionId=us'


def get_content_from_requests(page_link):
    headers = requests.utils.default_headers()
    headers['User-Agent'] = 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/71.0.3578.98 Safari/537.36'
    req = requests.get(page_link, headers=headers)
    return BeautifulSoup(req.content, 'html.parser')


def scroll_pages(link: str):
    print(link)
    page = get_content_from_requests(link)
    for div in page.findAll('div', {'date-code': DATE_CODE}):
        item = div.find('a', {'class': 'calendar-item-title subpage-trigg'})
        if item:
            # print(f'TITLE: {item.getText()}')
            titles.append(item.getText())
    last_index_div = page.find('div', {'class': LAST_ITEM_FIRST_PAGE})
    if not last_index_div:
        last_index_div = page.find('div', {'class': LAST_ITEM_PAGES})
    if last_index_div:
        scroll_pages(get_next_page_link(last_index_div))
    else:
        print(f'Found: {len(titles)} Titles')
        print('No more pages to scroll finishing...')


scroll_pages(INITIAL_LINK)
with open(f'titles.json', 'w') as outfile:
    json.dump(titles, outfile)

如果您的目标是使用Selenium,我认为同样的原则可能适用,除非它在加载页面时具有滚动功能。 相应地替换 INITIAL_LINK、DATE_CODE 和 BLOCK 也会得到其他月份。

【讨论】:

  • 它改进了我的搜索以删除八月,但只给了我九月的一部分。仅供参考,我稍微更改了您的编码以适应我的编写方式,但应该按照您的建议做同样的事情。
  • 你现在得到了什么?我从你的问题中注意到没有calendar-item-title 类,而是`calendar-item-title subpage-trigg`,开头有一个空格
  • 这很奇怪,如果我使用 calender-item-title 或您建议的那个,我会得到相同的答案。我没有得到所有的标题。我认为这是因为该网站不会生成整个网站。这就是为什么我使用硒,而不仅仅是美丽的汤。
  • 如果您查询https://www.releases.com/l/Games/2019/,整个页面,然后修剪掉 9 月的部分可能会做到吗?
  • 是的,或多或少是我所说的,但要好得多。谢谢!
猜你喜欢
  • 1970-01-01
  • 2015-01-19
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2014-07-06
相关资源
最近更新 更多