【问题标题】:Beautiful Soup - get outer tag text without getting inner tag textBeautiful Soup - 获取外部标签文本而不获取内部标签文本
【发布时间】:2022-01-08 01:51:18
【问题描述】:

我在 HTML 文件中有一些带有嵌套 <a> 标记的 <li> 标记,并且在列表和 <a> 标记中都有文本。但是,我想分别提取它们。我想让<li> 文本成为键tag 的值,而<a> 标记内的文本成为孩子tag 的键值。 (HTML sn-p 见下文)

我最终将其打印到 JSON 文件中,但得到了不需要的结果。主要的tag 应该只有“抽象可视化”......而不是所有其他东西。而且儿童标签后面应该只有“about”,而不是“/ Emotive and abstract”。 “情感和抽象”已经在标题中占有一席之地。您可以看到此索引示例的每个条目都显示相同的模式。如何将文本提取到正确的位置?我是 Beautiful Soup 的初学者。谢谢。

JSON 文件

{
    "tag": "abstract visualization\n\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\nabout / Emotive and abstract\n\n",
    "definition": "",
    "source": [],
    "children": [
        {
            "tag": "about / Emotive and abstract",
            "definition": "",
            "source": [
                {
                    "title": "Emotive and abstract",
                    "href": "https://learning.oreilly.com/library/view/data-visualization-a/9781849693462/ch02s03.html"
                }
            ]
        }
    ]
},
{
    "tag": "Adobe After Effects\n\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\t\nURL / Other specialist tools\n\n",
    "definition": "",
    "source": [],
    "children": [
        {
            "tag": "URL / Other specialist tools",
            "definition": "",
            "source": [
                {
                    "title": "Other specialist tools",
                    "href": "https://learning.oreilly.com/library/view/data-visualization-a/9781849693462/ch06.html"
                }
            ]
        }
    ]
},

HTML 文件 sn-p:

<ul id="letters">
    <li>abstract visualization
        <ul>
            <li>about / <a href="ch02s03.html" title="Emotive and abstract" class="link">Emotive and abstract</a></li>
        </ul>
    </li>
    <li>Adobe After Effects
        <ul>
            <li>URL / <a href="ch06.html" title="Other specialist tools" class="link">Other specialist tools</a></li>
        </ul>
    </li>
    <li>Adobe Flash
        <ul>
            <li>about / <a href="ch06.html" title="Programming environments" class="link">Programming environments</a></li>
            <li>URL / <a href="ch06.html" title="Programming environments" class="link">Programming environments</a></li>
        </ul>
    </li>
    <li>Adobe Illustrator
        <ul>
            <li>about / <a href="ch06.html" title="Other specialist tools" class="link">Other specialist tools</a></li>
            <li>URL / <a href="ch06.html" title="Other specialist tools" class="link">Other specialist tools</a></li>
        </ul>
    </li>
</ul>

相关代码:

# convert html to bs4 object
def bs4_convert(file):
    with open(file, encoding='utf8') as fp:
        html = BeautifulSoup(fp, 'html.parser')
    return html

# create a tag
def li_parser(letter, link_prefix):
    tags = []
    for li in letter.find_all('li', recursive=False):
        tag = {
            'tag': li.text,
            'definition': '',
            'source': [{'title': link.text, 'href': link_prefix + link['href']} for link in li.find_all('a', recursive=False)]
        }
        if li.find('ul'):
            tag['children'] = li_parser(li.find('ul'), link_prefix)
        tags.append(tag)

    return tags

# loop through all indices
def html_parser(html, link_prefix):
    tags = []
    # extract index
    html.find(id='backindex')
    # iterate over every indented letter in index
    letters = html.find_all(attrs={'id': 'letters'})
    for letter in letters:
        tags += li_parser(letter, link_prefix)

    return tags

tags = []
# parse the html
html = bs4_convert(course['file'])
# create tags
tags = html_parser(html, link_prefix)
# add course name as outermost tag
tags = add_course_tag(course['course'], tags)

【问题讨论】:

  • 好吧,如果在 Beautiful Soup 中没有办法做到这一点,我可以在父标签上使用 Python hack:.split('\n', 1)[0]。在上面的“抽象可视化”示例中,它去掉了字符 \n 之后的所有额外文本。

标签: python beautifulsoup


【解决方案1】:

可以通过名为contents 的列表访问标签的子代。在您的情况下,您正在搜索的文本只是 contents[0] 所以它比遍历所有孩子更容易。您只需使用strip() 删除不需要的标签和行

soup=BeautifulSoup(data, 'lxml')
lis=soup.select('#letters > li')
for li in lis:
    print(li.contents[0].strip())
    sub_li=li.select_one('ul li')
    print(sub_li.contents[0].strip()[:-2]) #get rid of the trailing slash

哪个输出

abstract visualization
about
Adobe After Effects
URL
Adobe Flash
about
Adobe Illustrator
about

【讨论】:

    【解决方案2】:

    要为您的标签获取正确的字符串,您可以在选择第一个元素时使用stripped_strings 与@diggusbickus 接近:

    'tag': list(li.stripped_strings)[0].strip(' /')
    

    示例

    def li_parser(letter, link_prefix):
        tags = []
        for li in letter.find_all('li', recursive=False):
            tag = {
                'tag': list(li.stripped_strings)[0].strip(' /'),
                'definition': '',
                'source': [{'title': link.text, 'href': link_prefix + link['href']} for link in li.find_all('a', recursive=False)]
            }
            if li.find('ul'):
                tag['children'] = li_parser(li.find('ul'), link_prefix)
            tags.append(tag)
    
        return tags
    

    输出

    [{"tag": "abstract visualization", "definition": "", "source": [], "children": [{"tag": "about", "definition": "", "source": [{"title": "Emotive and abstract", "href": "https://learning.oreilly.com/library/view/effective-data-storytelling/9781119615712/ch02s03.html"}]}]}, {"tag": "Adobe After Effects", "definition": "", "source": [], "children": [{"tag": "URL", "definition": "", "source": [{"title": "Other specialist tools", "href": "https://learning.oreilly.com/library/view/effective-data-storytelling/9781119615712/ch06.html"}]}]}, {"tag": "Adobe Flash", "definition": "", "source": [], "children": [{"tag": "about", "definition": "", "source": [{"title": "Programming environments", "href": "https://learning.oreilly.com/library/view/effective-data-storytelling/9781119615712/ch06.html"}]}, {"tag": "URL", "definition": "", "source": [{"title": "Programming environments", "href": "https://learning.oreilly.com/library/view/effective-data-storytelling/9781119615712/ch06.html"}]}]}, {"tag": "Adobe Illustrator", "definition": "", "source": [], "children": [{"tag": "about", "definition": "", "source": [{"title": "Other specialist tools", "href": "https://learning.oreilly.com/library/view/effective-data-storytelling/9781119615712/ch06.html"}]}, {"tag": "URL", "definition": "", "source": [{"title": "Other specialist tools", "href": "https://learning.oreilly.com/library/view/effective-data-storytelling/9781119615712/ch06.html"}]}]}]
    

    【讨论】:

      猜你喜欢
      • 2019-11-08
      • 2019-01-22
      • 1970-01-01
      • 1970-01-01
      • 2021-01-24
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-01-17
      相关资源
      最近更新 更多