【问题标题】:How to scrape content from a website with no class or id specified in attribute with BeautifulSoup4如何使用 BeautifulSoup4 从属性中未指定类或 id 的网站中抓取内容
【发布时间】:2021-10-13 10:13:47
【问题描述】:

我想在“a”标签(即只有名称-“42mm Architecture”)和“服务范围、建成项目类型、建成项目位置、工作风格、网站”中刮取单独的内容,如文本' 作为整个网页的 CSV 文件头及其内容。

元素没有与之关联的类或 ID。所以我有点纠结于如何正确提取这些细节,中间还有那些“br”和“b”标签。

在提供的代码块之前和之后有多个“p”标签。这是website。

<h2>
  <a href="http://www.dezeen.com/tag/design-by-42mm-architecture" rel="noopener noreferrer" target="_blank">
   42mm Architecture
  </a>
  |
  <span style="color: #808080;">
   Delhi | Top Architecture Firms/ Architects in India
  </span>
 </h2>
 <!-- /wp:paragraph -->
 <p>
  <b>
   Scope of services:
  </b>
  Architecture, Interiors, Urban Design.
  <br/>
  <b>
   Types of Built Projects:
  </b>
  Residential, commercial, hospitality, offices, retail, healthcare, housing, Institutional
  <br/>
  <b>
   Locations of Built Projects:
  </b>
  New Delhi and nearby states
  <b>
   <br/>
  </b>
  <b>
   Style of work
  </b>
  <span style="font-weight: 400;">
   : Contemporary
  </span>
  <br/>
  <b>
   Website
  </b>
  <span style="font-weight: 400;">
   :
   <a href="https://www.42mm.co.in/">
    42mm.co.in
   </a>
  </span>
 </p>

那么使用 BeautifulSoup4 是如何做到的呢?

【问题讨论】:

  • @Mooncrater 我要抓取的网站没有“class or id”属性。
  • 您仍然可以根据标签进行抓取。如果你知道它将是一个a 标签,你可以刮掉它。页面中是否还有其他a标签?
  • @Mooncrater 有 1 个“a”标签,后跟 span,p。但在整个页面中,它并不一致,也不遵循任何特定的模式。
  • 我不明白你所说的“不一致”是什么意思。请详细说明并显示您到目前为止编写的代码,以便我们了解您可能遇到的困难

标签: python web-scraping beautifulsoup python-requests export-to-csv


【解决方案1】:

您可以使用此示例作为如何从该页面抓取信息的基础:

import requests
import pandas as pd

url = "https://www.gov.uk/government/publications/endorsing-bodies-start-up/start-up"

soup = BeautifulSoup(requests.get(url).content, "html.parser")
parent = soup.select_one("div.govspeak")

mapping = {"sector": "sectors", "endorses businesses": "endorses businesses in"}

all_data = []
for h3 in parent.select("h3"):
    name = h3.text
    link = h3.a["href"] if h3.a else "-"

    ul = h3.find_next("ul")
    if ul and ul.find_previous("h3") == h3 and ul.parent == parent:
        li = [
            list(map(lambda x: mapping.get((i := x.strip()), i), v))
            for li in ul.select("li")
            if len(v := li.get_text(strip=True).split(":")) == 2
        ]
    else:
        li = []

    all_data.append({"name": name, "link": link, **dict(li)})


df = pd.DataFrame(all_data)
print(df)
df.to_csv("data.csv", index=False)

创建 data.csv(来自 LibreOffice 的屏幕截图):

【讨论】:

    【解决方案2】:

    这个有点费时间!网页不完整,标签和标识符较少。再补充一点,他们甚至没有对内容进行拼写检查例如。一个地方有一个标题Scope of Services,另一个地方有Scope of services,还有更多类似的!所以我所做的只是粗略的提取,如果您也有分页的想法,我相信它会对您有所帮助。

    import requests
    from bs4 import BeautifulSoup
    import csv
    
    page = requests.get('https://www.re-thinkingthefuture.com/top-architects/top-architecture-firms-in-india-part-1/')
    soup = BeautifulSoup(page.text, 'lxml')
    
    # there are many h2 tags but we want the one without any class name
    h2 = soup.find_all('h2', class_= '')
    
    headers = []
    contents = []
    header_len = []
    a_tags = []
    
    for i in h2:
        if i.find_next().name == 'a':             # to make sure we do not grab the wrong tag
            a_tags.append(i.find_next().text)
            p = i.find_next_sibling()
            contents.append(p.text)
            h =[j.text for j in  p.find_all('strong')]   #  some headings were bold in the website
            headers.append(h)
            header_len.append(len(h))
    
    # since only some headings were in bold the max number of bold would give all headers
    headers = headers[header_len.index(max(header_len))]
    
    # removing the : from headings
    headers = [i[:len(i)-1] for i in headers]
    
    # inserted a new heading
    headers.insert(0, 'Firm')
    
    # n for traversing through headers list
    # k for traversing through a_tags list
    n =1
    k =0
    
    # this is the difficult part where the content will have all the details in one value including the heading like this
    """
    Scope of services: Architecture, Interiors, Urban Design.Types of Built Projects: Residential, commercial, hospitality, offices, retail, healthcare, housing, InstitutionalLocations of Built Projects: New Delhi and nearby statesStyle of work: ContemporaryWebsite: 42mm.co.in
    """
    # thus I am splitting it using the ':' and then splicing it from the start of the each heading
    
    contents = [i.split(':') for i in contents]
    for i in contents:
        for j in i:
            h = headers[n][:5]
            if i.index(j) == 0:
                i[i.index(j)] = a_tags[k]
                n+=1
                k+=1
            elif h in j:
                i[i.index(j)] = j[:j.index(h)]
                j = j[:j.index(h)]
                if n < len(headers)-1:
                    n+=1
        n =1
    
        # merging those extra values in the list if any
        if len(i) == 7:
            i[3] = i[3] + ' ' + i[4]
            i.remove(i[4])
    
    # writing into csv file
    # if you don't want a line space between each row then add newline = '' argument in the open function below
    with open('output.csv', 'w') as f:   
        writer = csv.writer(f)
        writer.writerow(headers)
        writer.writerows(contents)
    

    这是输出:

    如果要分页,只需将页码添加到网址末尾即可!

    page_num = 1
    while page_num <13:
        page = requests.get(f'https://www.re-thinkingthefuture.com/top-architects/top-architecture-firms-in-india-part-1/{page_num}/')
    
        # paste the above code starting from soup = BeautifulSoup(page.text, 'lxml')
    
        page_num +=1
    

    希望对您有所帮助,如果有任何错误,请告诉我。

    编辑 1: 我忘了说最重要的部分抱歉,如果有一个带有no class 名称的标签,那么您仍然可以使用我在上面的代码中使用的标签来获取标签

    h2 = soup.find_all('h2', class_= '')
    

    这只是说给我所有没有类名的h2 标签。这本身有时可能是一个唯一标识符,因为我们使用此 no class value 来识别它。

    【讨论】:

    • 非常感谢。尝试对其进行分页,但网站制作不正确,其他页面上的 h2 标签存在很多不一致。
    • 看到那个来了。您可以尝试获取尽可能多的正确数据,然后使用re 模块对其进行操作。当您只想获取与特定模式匹配的字符串的子集时,该模块将对您有所帮助。
    • 嘿,我可以请求更多帮助吗? stackoverflow.com/q/68975727/16626943 谢谢。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2017-05-14
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-01-18
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多