【问题标题】:How to extract a href from a given div?如何从给定的div中提取href?
【发布时间】:2018-06-23 12:54:06
【问题描述】:

我有以下网页的 HTML 代码:

<div class="align-center">
<a target="_blank" rel="nofollow" class="link-block-2 w-inline-block w-condition-invisible">
  <img src="https://global.com/slack-symbol.png" alt="Slack link">
</a>
<a target="_blank" rel="nofollow" href="https://twitter.com/abc" class="link-block-2 w-inline-block">
  <img src="https://global.com/twitter.png" width="16" alt="Twitter link">
</a>
<a target="_blank" rel="nofollow" href="https://t.me/abc" class="link-block-2 w-inline-block">
  <img src="https://global.com/telegram.png" alt="Telegram link">
</a>
</div>

另外,我的链接名称列表如下:

links_dict = {}
links = ["Slack","Twitter","Telegram"] 

我想为每个对应的链接提取href 值。如果没有href(见上面示例代码中的Slack),则表示没有链接。

预期的输出如下:

"Slack" -> "None"
"Twitter" -> "https://twitter.com/abc"
"Telegram" -> "https://t.me/abc"

我无法仅通过a 访问a href,因为有许多其他div 元素与其他a。

我想将BeautifulSoap 或Selenium 与PhantomJS 一起使用。这是我尝试过的:

BeautifulSoap:

res = requests.get("https://myurl.com")
soup = BeautifulSoup(res.content,'html.parser')
tags = soup.find_all(class_="align-center")
for tag in tags:
    print tag.text.strip()

硒:

driver = webdriver.PhantomJS()
driver.set_window_size(1120, 550)
driver.get("https://mytest.com")

tags = driver.find_elements_by_class_name("align-center")

for tag in tags:
    tag.find_element_by_tag_name("a").click()
    url = driver.current_url
    print(url)
driver.quit()

【问题讨论】:

  • 您尝试过以下解决方案吗?你有什么反馈?
  • @Shahin:我自己找到了解决方案。另外,我不明白为什么我的问题被否决了。请点赞。

标签: python html selenium beautifulsoup phantomjs


【解决方案1】:

试试下面的脚本。它会为您获取所需的结果。

from bs4 import BeautifulSoup

content="""
<div class="align-center">
<a target="_blank" rel="nofollow" class="link-block-2 w-inline-block w-condition-invisible">
  <img src="https://global.com/slack-symbol.png" alt="Slack link">
</a>
<a target="_blank" rel="nofollow" href="https://twitter.com/abc" class="link-block-2 w-inline-block">
  <img src="https://global.com/twitter.png" width="16" alt="Twitter link">
</a>
<a target="_blank" rel="nofollow" href="https://t.me/abc" class="link-block-2 w-inline-block">
  <img src="https://global.com/telegram.png" alt="Telegram link">
</a>
</div>
"""
soup = BeautifulSoup(content,"html5lib")
links = {item.get("alt").split(" ")[0]:link.get('href') for item,link in zip(soup.select(".align-center a img"),soup.select(".align-center a"))}
print(links)

输出:

{'Slack': None, 'Telegram': 'https://t.me/abc', 'Twitter': 'https://twitter.com/abc'}

或者你可以用稍微不同的方式做同样的事情:

soup = BeautifulSoup(content,"html5lib")
for item in soup.select(".align-center a img"):
    title = item.get("alt").split(" ")[0]
    link = item.findParent().get('href')
    print(title,link)

输出:

Slack None
Twitter https://twitter.com/abc
Telegram https://t.me/abc

【讨论】:

    【解决方案2】:

    使用BeatifulSoup 继续您的想法,您可以从每个标签中找到所有img 链接,然后检查该链接是否包含正确的alt 模式。

    如果模式正确,获取父链接。

    import re
    
    ...
    
    links = []
    tags = soup.find_all(class_="align-center")
    for tag in tags:
        # For each tag, get all the images
        for img in tag.find_all('img'):
            # Ensure the img has the correct `alt` pattern
            if re.match('(Twitter|Slack|Telegram) link', img.attrs.get('alt')):
                # Store the link found.
                links.append(img.findParent().attrs.get('href'))
    

    【讨论】:

      【解决方案3】:

      由于您要为子节点中每个对应的alt 属性提取href 值,您可以按照以下代码块使用Selenium:

      tags = driver.find_elements_by_xpath("//div[@class='align-center']/a/img")
      my_alt = []
      my_href= []
      for tag in tags:
          alt_text = tag.getAttribute("alt")
          my_alt.append(alt_text)
          my_href.append(driver.find_element_by_xpath("//div[@class='align-center']/a/img[.='" + alt_text + "']//preceding::a[1]").getAttribute("href"))
      for alt, href in zip(my_alt, my_href):
          print(alt, href)
      

      【讨论】:

        猜你喜欢
        • 2021-04-19
        • 2014-03-02
        • 2015-12-04
        • 1970-01-01
        • 1970-01-01
        • 2016-07-16
        • 2015-10-29
        • 1970-01-01
        • 2020-06-18
        相关资源
        最近更新 更多