【问题标题】:Scrape whole table with selenium and BeautifulSoup?用硒和 BeautifulSoup 刮掉整张桌子?
【发布时间】:2022-01-13 18:41:48
【问题描述】:

我想在这个网站的中间刮掉整个桌子: https://www.brilliantearth.com/lab-diamonds-search/

我使用以下代码进行了尝试 - 但我只获得了表的前 200 行:

import time
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
import os, sys
import xlwings as xw
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager
from fake_useragent import UserAgent

if __name__ == '__main__':
  WAIT = 3  
  ua = UserAgent()
  userAgent = ua.random
  options = Options()
  # options.add_argument('--headless')
  options.add_experimental_option ('excludeSwitches', ['enable-logging'])
  options.add_argument("start-maximized")
  options.add_argument('window-size=1920x1080')                               
  options.add_argument('--no-sandbox')
  options.add_argument('--disable-gpu')  
  options.add_argument(f'user-agent={userAgent}')   
  srv=Service(ChromeDriverManager().install())
  driver = webdriver.Chrome (service=srv, options=options)    
  waitWebDriver = WebDriverWait (driver, 10)         
  
  lElems = []
  link = f"https://www.brilliantearth.com/lab-diamonds-search/" 
  # driver.minimize_window()        # optional
  driver.get (link)       
  time.sleep(WAIT) 
  driver.find_element(By.XPATH,"(//button[@title='Accept All'])[1]").click() 
  time.sleep(WAIT) 
  soup = BeautifulSoup (driver.page_source, 'html.parser')     
  time.sleep(WAIT)   
  tmpSearch = soup.find("div", {"id": "diamond_search_wrapper"})
  tmpDIVs = tmpSearch.select("div.inner.item")
  for idx,elem in enumerate(tmpDIVs):
    tmpTD = elem.find_all("td")  
    row = []    
    for e2 in tmpTD:
      row.append(e2.text)
    print(idx, row)

我想用 selenium 向下滚动到该表的最底部。 但是当我向下滚动时,只有整个页面向下滚动,而不是里面的表格。

如何在表格中向下滚动到底部? (然后可能会从表中刮掉所有元素)

【问题讨论】:

    标签: python selenium web-scraping beautifulsoup


    【解决方案1】:

    抓取他们的后端网络调用会更快更容易,在浏览器中探索它们打开开发者工具-网络-获取/XHR并刷新页面或向下滚动您想要的数据,您可以看到正在发生的网络调用.我在下面重新创建了它们并将数据转储到 csv 中:

    import requests
    import pandas as pd
    
    headers =   {
        'accept':'application/json, text/javascript, */*; q=0.01',
        'accept-encoding':'gzip, deflate, br',
        'referer':'https://www.brilliantearth.com/lab-diamonds-search/',
        'sec-fetch-site':'same-origin',
        'user-agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.71 Safari/537.36',
        'x-requested-with':'XMLHttpRequest'
        }
    
    final = []
    for page in range(1,10):
        print(f'Scraping page {page}')
        new_url = f'https://www.brilliantearth.com/lab-diamonds/list/?page={page}&shapes=Round&cuts=Fair%2CGood%2CVery%20Good%2CIdeal%2CSuper%20Ideal&colors=J%2CI%2CH%2CG%2CF%2CE%2CD&clarities=SI2%2CSI1%2CVS2%2CVS1%2CVVS2%2CVVS1%2CIF%2CFL&polishes=Good%2CVery%20Good%2CExcellent&symmetries=Good%2CVery%20Good%2CExcellent&fluorescences=Very%20Strong%2CStrong%2CMedium%2CFaint%2CNone&min_carat=0.30&max_carat=8.18&min_table=45.00&max_table=82.50&min_depth=5.00&max_depth=85.80&min_price=350&max_price=128290&stock_number=&row=0&requestedDataSize=200&order_by=price&order_method=asc&currency=%24&has_v360_video=&dedicated=&min_ratio=1.00&max_ratio=2.75&exclude_quick_ship_suppliers=&MIN_PRICE=350&MAX_PRICE=128290&MIN_CARAT=0.3&MAX_CARAT=8.18&MIN_TABLE=45&MAX_TABLE=82.5&MIN_DEPTH=5&MAX_DEPTH=85.8'
        resp = requests.get(new_url,headers=headers).json()
    
        for diamond in resp['diamonds']:
            diamond.pop('v360_src',None) #remove long video and images links to clean up csv
            diamond.pop('images',None)
            final.append(diamond)
    
    df = pd.DataFrame(final)
    df.to_csv('diamonds.csv',encoding='utf-8',index=False)
    print('Saved to diamonds.csv')
    

    【讨论】:

    • 非常感谢 - 这很完美。也许还有 2 个额外的问题:1)你现在如何使用 ?page 参数? (我检查了网络——在你拥有的 fetch/XHR 中找到了 url——但我的没有 ?page 参数) 2)你怎么知道你必须使用哪些标题? (当我检查 fetch/XHR - 我有更多的请求项目,而不仅仅是你拥有的那个)
    • 我一直向下滚动并查看网络请求以查看带有“page”参数的网络请求,我尝试在没有标题的情况下进行抓取,并收到错误提示“缺少引用标题”等,因此添加标题直到它工作
    • Thx - 很有趣 - 我没有找到一个带有“页面”的链接 - 希望我在正确的地方寻找 => 标题和请求 URL
    • 我的网址总是看起来像:https://www.brilliantearth.com/lab-diamonds/list/?shapes=Round&cuts=Fair,Good,Very Good,Ideal,Super Ideal&colors=J,I,H,G,F,E,D&clarities=SI2,SI1,VS2,VS1,VVS2,VVS1,IF,FL&polishes=Good,Very Good,Excellent&symmetries=Good,Very Good,Excellent&fluorescences=Very Strong,Strong,Medium,Faint,None&min_carat=0.52&max_carat=8.18&min_table=45.00&max_table=82.50&min_depth=5.00&max_depth=85.80&min_price=600&max_price=128290&stock_number=&row=0&page=5&requestedDataSize=200&order_by=price&order_method=asc&currency=$&has_v360_video=&dedicated=&min_ratio=1.00&max..........
    • 我将“page=" 从末尾移到开头,以便您看得更清楚。我可以在你的文章中看到它最后说“page = 5”......
    【解决方案2】:

    您可以使用以下代码滚动内表:

    rows = driver.find_elements_by_css_selector("#diamond_search_wrapper div.inner.item") 
    
    for row in rows :
        driver.execute_script("arguments[0].scrollIntoView();", row )
        #scrape the data etc..
    

    【讨论】:

    • 感谢您的回复 - 但是当我这样做并随后刮擦时 - 我只能像以前一样再次获得 200 个元素。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2019-08-23
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-01-05
    相关资源
    最近更新 更多