【问题标题】:print text inside parent div beautifulsoup在父 div beautifulsoup 中打印文本
【发布时间】:2018-12-15 11:59:41
【问题描述】:

我正在尝试从中获取每个产品的名称和价格 https://www.daraz.pk/catalog/?q=risk 但什么也没显示。

containers = page_soup.find_all("div",{"class":"c2p6A5"})

for container in containers:
  pname = container.findAll("div", {"class": "c29Vt5"})
  name = pname[0].text
  price1 = container.findAll("span", {"class": "c29VZV"})
  price = price1[0].text
  print(name)
  print(price)

【问题讨论】:

  • 您可以获取 json,但您必须知道页数才能获得所有结果。获取页数的唯一方法是首先让页面呈现,例如使用 selenium 然后切换到请求。您也不需要使用正则表达式,因为您可以简单地执行 item = soup.select('script')[2]
  • 是的,谢谢.. 其他人也建议过这个我还在弄清楚:你们如何检查它返回的 json 数据
  • F12 打开开发工具并检查 html。我使用 Ctrl + F (5,900) 在 html 中搜索了第一个价格。这向我展示了该值在脚本标记内的 json 字符串中的出现。您可以从脚本语法中看到这是用于更新页面的。您可以使用语法获取每个页面:daraz.pk/catalog/?page=1&q=risk 并更改页码。但是,如果不使用浏览器 (AFAIK),您无法获得总页数。
  • 所以我会根据时间是否真的是一个问题,使用一种解决方案来呈现页面以获取页码计数,然后切换到请求。您可以从使用选择器 li[class*="ant-pagination-item ant-pagination-item-"] 的 len 中获取页数
  • 非常感谢

标签: python web-scraping beautifulsoup


【解决方案1】:

页面中有JSON数据,你可以使用beautifulsoup在<script>标签中获取,但我认为不需要,因为你可以直接使用json和re获取它

import requests, json, re

html = requests.get('https://.......').text

jsonStr = re.search(r'window.pageData=(.*?)</script>', html).group(1)
jsonObject = json.loads(jsonStr)

for item in jsonObject['mods']['listItems']:
    print(item['name'])
    print(item['price'])

【讨论】:

  • @ewwink 我注意到您对这个主题非常了解。我不太确定什么时候使用 selenium,什么时候仍然可以使用请求(我通常只是在请求没有产生结果时默认使用 selenium)。是否有链接/资源可以帮助我更好地理解您所知道的?有点像一个“清单”,要寻找,要确切知道在什么时候只有硒才是正确的选择?
  • 我不是专家 :D 很简单,如果我在 page source 中找不到硒或无法复制 XHR 或 Ajax 请求,我将使用硒。
  • 谢谢。我不太熟悉 xhr 或 Ajax 拒绝。但只要你这么说,就给了我一些方向。
【解决方案2】:

如果页面是动态的,Selenium 应该处理好这个

from bs4 import BeautifulSoup
import requests
from selenium import webdriver

browser = webdriver.Chrome()
browser.get('https://www.daraz.pk/catalog/?q=risk')

r = browser.page_source
page_soup = bs4.BeautifulSoup(r,'html.parser')

containers = page_soup.find_all("div",{"class":"c2p6A5"})

for container in containers:
  pname = container.findAll("div", {"class": "c29Vt5"})
  name = pname[0].text
  price1 = container.findAll("span", {"class": "c29VZV"})
  price = price1[0].text
  print(name)
  print(price)

browser.close() 

输出:

Risk Strategy Game
Rs. 5,900
Risk Classic Board Game
Rs. 945
RISK - The Game of Global Domination
Rs. 1,295
Risk Board Game
Rs. 1,950
Risk Board Game - Yellow
Rs. 3,184
Risk Board Game - Yellow
Rs. 1,814
Risk Board Game - Yellow
Rs. 2,086
Risk Board Game - The Game of Global Domination
Rs. 975
...

【讨论】:

  • 没有 Selenium 有什么办法吗?我收到以下错误 platform_sensor_reader_win.cc(242)] NOT IMPLEME .. 我猜是因为 chromedriver 问题
  • 嗯。没有把握。我会调查的。
【解决方案3】:

我错了。用于计算页数的信息存在于 json 中,因此您可以获得所有结果。不需要正则表达式,因为您可以提取相关的脚本标签。此外,您可以在循环中创建页面 url。

import requests
from bs4 import BeautifulSoup
import json
import math

def getNameAndPrice(url):
    res = requests.get(url)
    soup = BeautifulSoup(res.content,'lxml')
    data = json.loads(soup.select('script')[2].text.strip('window.pageData='))
    if url == startingPage:
        resultCount = int(data['mainInfo']['totalResults'])
        resultsPerPage = int(data['mainInfo']['pageSize'])
        numPages = math.ceil(resultCount/resultsPerPage)
    result = [[item['name'],item['price']] for item in data['mods']['listItems']]   
    return result

resultCount = 0
resultsPerPage = 0
numPages = 0
link = "https://www.daraz.pk/catalog/?page={}&q=risk"
startingPage = "https://www.daraz.pk/catalog/?page=1&q=risk"
results = []
results.append(getNameAndPrice(startingPage))

for links in [link.format(page) for page in range(2,numPages + 1)]: 
    results.append(getNameAndPrice(links))

【讨论】:

    【解决方案4】:

    向像我这样的新手参考 JSON 答案。 您可以使用 Selenium 导航到搜索结果页面,如下所示:

    PS:非常感谢@ewwink。你拯救了我的一天!

    from selenium import webdriver
    from selenium.webdriver.common.keys import Keys
    import time #time delay when load web
    import requests, json, re
    
    keyword = 'fan'
    
    opt = webdriver.ChromeOptions()
    opt.add_argument('headless')
    driver = webdriver.Chrome(options = opt)
    
    # driver = webdriver.Chrome()
    
    url = 'https://www.lazada.co.th/'
    driver.get(url)
    
    search = driver.find_element_by_name('q')
    search.send_keys(keyword)
    search.send_keys(Keys.RETURN)
    
    time.sleep(3) #wait for web load for 3 secs
    
    page_html = driver.page_source #Selenium way of page_html = webopen.read() for BS
    
    driver.close()
    
    jsonStr = re.search(r'window.pageData=(.*?)</script>', page_html).group(1)
    jsonObject = json.loads(jsonStr)
    
    for item in jsonObject['mods']['listItems']:
        print(item['name'])
        print(item['sellerName'])
    

    【讨论】:

      猜你喜欢
      • 2019-04-24
      • 2017-06-19
      • 1970-01-01
      • 2016-07-02
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多