【发布时间】:2018-06-23 12:54:06
【问题描述】:
我有以下网页的 HTML 代码:
<div class="align-center">
<a target="_blank" rel="nofollow" class="link-block-2 w-inline-block w-condition-invisible">
<img src="https://global.com/slack-symbol.png" alt="Slack link">
</a>
<a target="_blank" rel="nofollow" href="https://twitter.com/abc" class="link-block-2 w-inline-block">
<img src="https://global.com/twitter.png" width="16" alt="Twitter link">
</a>
<a target="_blank" rel="nofollow" href="https://t.me/abc" class="link-block-2 w-inline-block">
<img src="https://global.com/telegram.png" alt="Telegram link">
</a>
</div>
另外,我的链接名称列表如下:
links_dict = {}
links = ["Slack","Twitter","Telegram"]
我想为每个对应的链接提取href 值。如果没有href(见上面示例代码中的Slack),则表示没有链接。
预期的输出如下:
"Slack" -> "None"
"Twitter" -> "https://twitter.com/abc"
"Telegram" -> "https://t.me/abc"
我无法仅通过a 访问a href,因为有许多其他div 元素与其他a。
我想将BeautifulSoap 或Selenium 与PhantomJS 一起使用。这是我尝试过的:
BeautifulSoap:
res = requests.get("https://myurl.com")
soup = BeautifulSoup(res.content,'html.parser')
tags = soup.find_all(class_="align-center")
for tag in tags:
print tag.text.strip()
硒:
driver = webdriver.PhantomJS()
driver.set_window_size(1120, 550)
driver.get("https://mytest.com")
tags = driver.find_elements_by_class_name("align-center")
for tag in tags:
tag.find_element_by_tag_name("a").click()
url = driver.current_url
print(url)
driver.quit()
【问题讨论】:
-
您尝试过以下解决方案吗?你有什么反馈?
-
@Shahin:我自己找到了解决方案。另外,我不明白为什么我的问题被否决了。请点赞。
标签: python html selenium beautifulsoup phantomjs