【发布时间】:2021-02-23 15:03:53
【问题描述】:
以下是我尝试使用 python 和 selenium 抓取的一些 html。
<h2 class ="page-title">
Strange Video Titles
<span class="duration">28 min</span>
<span class="video-hd-mark">720p</span>
</h2>
下面是我的代码:
title=driver.find_element_by_class_name('page-title').text
print(title)
但是,当我运行它时,它会打印 h2 标记中的所有内容,包括 span 类中的文本。我尝试在末尾添加 [0] 或 [1] 以指定我只想要第一行文本,但这不起作用。如何只打印跨度类上方的视频标题?
编辑 - 我认为这是解决方案
所以我决定做以下事情:
title=driver.find_element_by_class_name('page-title').text
duration = driver.find_element_by_xpath('/html/body/div/div[4]/h2/span[1]').text
vid_quality =driver.find_element_by_xpath('/html/body/div/div[4]/h2/span[2]').text
if (duration) in title:
title = title.replace(duration, "")
if(vid_quality) in title:
title = title.replace(vid_quality,"")
谢谢。
【问题讨论】:
标签: python selenium xpath css-selectors webdriverwait