【发布时间】:2020-11-25 15:34:52
【问题描述】:
我正在尝试从代码中的链接中抓取视频标题。
基本上是想滚动+刮。
我的代码运行了,但它刮掉了页面的一半,而不是刮掉剩下的一半,而是重复前半部分。
import time
from selenium import webdriver
from bs4 import BeautifulSoup
import requests
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from webdriver_manager.chrome import ChromeDriverManager
options = Options()
driver = webdriver.Chrome(ChromeDriverManager().install())
url='https://www.youtube.com/user/OakDice/videos'
driver.get(url)
content=driver.page_source.encode('utf-8').strip()
soup=BeautifulSoup(content,'lxml')
SCROLL_PAUSE_TIME = 2
# Get scroll height
last_height = driver.execute_script("return document.documentElement.scrollHeight")
while True:
# Scroll down to bottom
time.sleep(2)
driver.execute_script("window.scrollTo(0, arguments[0]);", last_height)
# Wait to load page
time.sleep(SCROLL_PAUSE_TIME)
titles = soup.findAll('a', id='video-title')
for title in titles:
print(title.text)
# Calculate new scroll height and compare with last scroll height
new_height = driver.execute_script("return document.documentElement.scrollHeight")
if new_height == last_height:
break
last_height = new_height
【问题讨论】:
-
正如我在您的另一个问题中所说的 (stackoverflow.com/questions/65008223/…)。我认为您不想使用 driver.page_source