【问题标题】:I'm trying to extract reviews inside the i-frame using selenium in python but not able to access inner HTML.?我正在尝试在 python 中使用 selenium 提取 i-frame 内的评论,但无法访问内部 HTML。?
【发布时间】:2020-05-08 06:24:01
【问题描述】:

我正在尝试从 iframe 中提取评论,我正在尝试切换到 iframe,我猜我成功切换了,但无法访问更多标签和属性..

我尝试了多种解决方案,即评论,我需要获得特定的评论 div,或者即使我们获得 iframe 的整页源代码,我也可以通过 Beautifulsoup 进行解析。

网址=https://www.aliexpress.com/item/4000295971597.html?spm=a2g0o.productlist.0.0.46754f9bCFP2xJ&s=p&ad_pvid=202005070358123970448424509400000207120_1&algo_pvid=28913e75-ad10-4d7c-a25b-e1c67d3e18be&algo_expid=28913e75-ad10-4d7c-a25b-e1c67d3e18be-0&btsid=0ab6d69f15888490922438266e45ea&ws_ab_test=searchweb0_0,searchweb201602_,searchweb201603_

我的代码:

from telnetlib import EC
from time import sleep
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.wait import WebDriverWait

def get_pro_reviews():
driver = webdriver.Chrome()
driver.get(web_url)
try:
    driver.execute_script("window.scrollTo(0, 200)")
    driver.implicitly_wait(10)
    WebDriverWait(driver, 20).until(EC.element_to_be_clickable(
        (By.XPATH, '/html/body/div[5]/div/div[3]/div[2]/div[2]/div[1]/div/div[1]/ul/li[2]'))).click()
    print('==> Review tab Clicked')
    driver.implicitly_wait(5)
    driver.execute_script("window.scrollTo(200, 1200)")
    sleep(5)
except Exception as error:
    product['Review'] = 'Reviews Not available'
    print(f'Error in Clicking review tab ==>{error}')
try:
    sleep(5)
    # driver.switch_to.frame(driver.find_element_by_tag_name("iframe"))
    # iframes=driver.find_elements_by_tag_name("iframe")
    # chn=driver.switch_to.frame(iframes[0])
    print(driver)
    che = WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it(
        driver.find_element_by_xpath("//*[@id='product-evaluation']")))
    print('==> Swithched to iframe')
    print(che)
    print(driver)
    review_elem = driver.page_source()
    # review_elem = driver.find_element_by_css_selector('div.feedback-list-wrap')
    print(review_elem)
    # review_elem = chn.find_element_by_xpath('/html/body/div/div[5]')
    review_elem_source_code = review_elem.get_attribute("innerHTML")
    review_elem_soup: BeautifulSoup = BeautifulSoup(review_elem_source_code, 'html.parser')
    print(review_elem_soup)

    for review_raw in review_elem_soup.find_elements_by_css_selector('div.feedback-item'):
        print(review_raw)
    driver.switch_to.default_content()
except Exception as error:
    print(f'Error in getting Reviews ==>{error}')

错误:

Error in getting Reviews ==>'str' object is not callable

这些是我需要的评论!

【问题讨论】:

  • 你能放一张你要检索哪些内容的图片吗?
  • @Armando 是的,让我说吧!

标签: python selenium-webdriver iframe web-scraping selenium-chromedriver


【解决方案1】:

我在顶部添加了 web_url 变量,因为没有定义。

我在顶部添加了 product 作为 dict,因为没有定义。我假设是一个字典,因为你做product['Review'] = 'Reviews Not available'。如果您在其他地方定义了它,请将两者都删除。

现在,我们来看看你的错误:

review_elem = driver.page_source() 

错误:

1)page_source 是一个属性而不是类的方法,所以你应该删除 ()

2) 这将返回一个包含所有 html 的字符串,这就是您收到错误的原因:

"Error in getting Reviews ==>'str' object is not callable"

str 因为此时 review_elem 是一个字符串并且不可调用,因为您不能调用属性(()),您只需将其作为类的属性访问即可。

了解了这个概念后,这对你不起作用。如果您继续并按原样运行代码,则在分配此代码时会出现另一个错误:

review_elem_source_code = review_elem.get_attribute("innerHTML")

因为 review_elem(它是一个字符串)没有方法 get_attribute()。 所以,我提出以下建议:

review_elem_source_code = driver.execute_script("return document.body.innerHTML;")

这将获得 iframe 的 innerHtml。

最后,这样做之后你会得到另一个错误,因为for循环中出现了以下错误:

review_elem_soup.find_elements_by_css_selector('div.feedback-item')

find_elements_by_css_selector() 在 BeautifulSoup 中不存在,而在驱动程序中存在。因此,为了使用 BeautifulSoup 获得您想要的所有元素,您应该更改为:

review_elem_soup.select('div.feedback-item')

代码如下:

from time import sleep
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.wait import WebDriverWait

web_url = "https://www.aliexpress.com/item/4000295971597.html?spm=a2g0o.productlist.0.0.46754f9bCFP2xJ&s=p&ad_pvid=202005070358123970448424509400000207120_1&algo_pvid=28913e75-ad10-4d7c-a25b-e1c67d3e18be&algo_expid=28913e75-ad10-4d7c-a25b-e1c67d3e18be-0&btsid=0ab6d69f15888490922438266e45ea&ws_ab_test=searchweb0_0,searchweb201602_,searchweb201603_"
driver = webdriver.Chrome()
driver.get(web_url)
product = {}
try:
    driver.execute_script("window.scrollTo(0, 200)")
    driver.implicitly_wait(10)
    WebDriverWait(driver, 20).until(EC.element_to_be_clickable(
        (By.XPATH, '/html/body/div[5]/div/div[3]/div[2]/div[2]/div[1]/div/div[1]/ul/li[2]'))).click()
    print('==> Review tab Clicked')
    driver.implicitly_wait(5)
    driver.execute_script("window.scrollTo(200, 1200)")
    sleep(5)
except Exception as error:
    product['Review'] = 'Reviews Not available'
    print(f'Error in Clicking review tab ==>{error}')
try:
    sleep(5)
    che = WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it(
        driver.find_element_by_xpath("//*[@id='product-evaluation']")))
    print('==> Switched to iframe')
    review_elem_source_code = driver.execute_script("return document.body.innerHTML;")
    review_elem_soup: BeautifulSoup = BeautifulSoup(review_elem_source_code, 'html.parser')
    for review_raw in review_elem_soup.select('div.feedback-item'):
        print(review_raw)
    driver.switch_to.default_content()
except Exception as error:
    print(f'Error in getting Reviews ==>{error}')

Ps:我删除了你的大部分打印,所以代码更容易解​​释。

希望这能帮助你实现你想要的。

【讨论】:

  • @mobin alhassan 抱歉,我有点忙。之前无法回复,希望对您有所帮助!
  • 谢谢它对我来说很完美......这是故意不定义那个变量,我假设你明白了。
猜你喜欢
  • 2021-03-09
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-06-30
  • 2020-05-22
  • 1970-01-01
相关资源
最近更新 更多