【发布时间】:2021-10-13 10:13:47
【问题描述】:
我想在“a”标签(即只有名称-“42mm Architecture”)和“服务范围、建成项目类型、建成项目位置、工作风格、网站”中刮取单独的内容,如文本' 作为整个网页的 CSV 文件头及其内容。
元素没有与之关联的类或 ID。所以我有点纠结于如何正确提取这些细节,中间还有那些“br”和“b”标签。
在提供的代码块之前和之后有多个“p”标签。这是website。
<h2>
<a href="http://www.dezeen.com/tag/design-by-42mm-architecture" rel="noopener noreferrer" target="_blank">
42mm Architecture
</a>
|
<span style="color: #808080;">
Delhi | Top Architecture Firms/ Architects in India
</span>
</h2>
<!-- /wp:paragraph -->
<p>
<b>
Scope of services:
</b>
Architecture, Interiors, Urban Design.
<br/>
<b>
Types of Built Projects:
</b>
Residential, commercial, hospitality, offices, retail, healthcare, housing, Institutional
<br/>
<b>
Locations of Built Projects:
</b>
New Delhi and nearby states
<b>
<br/>
</b>
<b>
Style of work
</b>
<span style="font-weight: 400;">
: Contemporary
</span>
<br/>
<b>
Website
</b>
<span style="font-weight: 400;">
:
<a href="https://www.42mm.co.in/">
42mm.co.in
</a>
</span>
</p>
那么使用 BeautifulSoup4 是如何做到的呢?
【问题讨论】:
-
@Mooncrater 我要抓取的网站没有“class or id”属性。
-
您仍然可以根据标签进行抓取。如果你知道它将是一个
a标签,你可以刮掉它。页面中是否还有其他a标签? -
@Mooncrater 有 1 个“a”标签,后跟 span,p。但在整个页面中,它并不一致,也不遵循任何特定的模式。
-
我不明白你所说的“不一致”是什么意思。请详细说明并显示您到目前为止编写的代码,以便我们了解您可能遇到的困难
标签: python web-scraping beautifulsoup python-requests export-to-csv