【发布时间】:2022-01-08 01:51:18
【问题描述】:
我在 HTML 文件中有一些带有嵌套 <a> 标记的 <li> 标记,并且在列表和 <a> 标记中都有文本。但是,我想分别提取它们。我想让<li> 文本成为键tag 的值,而<a> 标记内的文本成为孩子tag 的键值。 (HTML sn-p 见下文)
我最终将其打印到 JSON 文件中,但得到了不需要的结果。主要的tag 应该只有“抽象可视化”......而不是所有其他东西。而且儿童标签后面应该只有“about”,而不是“/ Emotive and abstract”。 “情感和抽象”已经在标题中占有一席之地。您可以看到此索引示例的每个条目都显示相同的模式。如何将文本提取到正确的位置?我是 Beautiful Soup 的初学者。谢谢。
JSON 文件
{
"tag": "abstract visualization\n\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\nabout / Emotive and abstract\n\n",
"definition": "",
"source": [],
"children": [
{
"tag": "about / Emotive and abstract",
"definition": "",
"source": [
{
"title": "Emotive and abstract",
"href": "https://learning.oreilly.com/library/view/data-visualization-a/9781849693462/ch02s03.html"
}
]
}
]
},
{
"tag": "Adobe After Effects\n\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\nURL / Other specialist tools\n\n",
"definition": "",
"source": [],
"children": [
{
"tag": "URL / Other specialist tools",
"definition": "",
"source": [
{
"title": "Other specialist tools",
"href": "https://learning.oreilly.com/library/view/data-visualization-a/9781849693462/ch06.html"
}
]
}
]
},
HTML 文件 sn-p:
<ul id="letters">
<li>abstract visualization
<ul>
<li>about / <a href="ch02s03.html" title="Emotive and abstract" class="link">Emotive and abstract</a></li>
</ul>
</li>
<li>Adobe After Effects
<ul>
<li>URL / <a href="ch06.html" title="Other specialist tools" class="link">Other specialist tools</a></li>
</ul>
</li>
<li>Adobe Flash
<ul>
<li>about / <a href="ch06.html" title="Programming environments" class="link">Programming environments</a></li>
<li>URL / <a href="ch06.html" title="Programming environments" class="link">Programming environments</a></li>
</ul>
</li>
<li>Adobe Illustrator
<ul>
<li>about / <a href="ch06.html" title="Other specialist tools" class="link">Other specialist tools</a></li>
<li>URL / <a href="ch06.html" title="Other specialist tools" class="link">Other specialist tools</a></li>
</ul>
</li>
</ul>
相关代码:
# convert html to bs4 object
def bs4_convert(file):
with open(file, encoding='utf8') as fp:
html = BeautifulSoup(fp, 'html.parser')
return html
# create a tag
def li_parser(letter, link_prefix):
tags = []
for li in letter.find_all('li', recursive=False):
tag = {
'tag': li.text,
'definition': '',
'source': [{'title': link.text, 'href': link_prefix + link['href']} for link in li.find_all('a', recursive=False)]
}
if li.find('ul'):
tag['children'] = li_parser(li.find('ul'), link_prefix)
tags.append(tag)
return tags
# loop through all indices
def html_parser(html, link_prefix):
tags = []
# extract index
html.find(id='backindex')
# iterate over every indented letter in index
letters = html.find_all(attrs={'id': 'letters'})
for letter in letters:
tags += li_parser(letter, link_prefix)
return tags
tags = []
# parse the html
html = bs4_convert(course['file'])
# create tags
tags = html_parser(html, link_prefix)
# add course name as outermost tag
tags = add_course_tag(course['course'], tags)
【问题讨论】:
-
好吧,如果在 Beautiful Soup 中没有办法做到这一点,我可以在父标签上使用 Python hack:
.split('\n', 1)[0]。在上面的“抽象可视化”示例中,它去掉了字符\n之后的所有额外文本。
标签: python beautifulsoup