【问题标题】:How to Convert HTML into List of Python Dictionaries如何将 HTML 转换为 Python 字典列表
【发布时间】:2021-04-01 06:42:49
【问题描述】:

我需要将 HTML 转换为 python 字典列表。示例 HTML 代码:

<div data-axite-container="1" data-axite-uuid="1f0c5634-9ff9-4942-861e-2c7e75d6f2ef" data-axite-id="a4d1a127-0fe8-4281-abd6-fe0653a8b519">
<h4>Question 1</h4>
<p>Some text here too</p> <p>And some text here</p>
<h4>Question 2</h4>
<p>Answer Here</p>
<ul><li>text1</li><li>text2</li>
<h4>Question 3</h4>
<p>Answer Here</p><table>...</table>
<h4>Question 4</h4>
<p>Answer</p>

预期结果:

[   
{
    "q": "Question 1",
    "a": ["Some text here too","And some text here"]
},
{
    "q": "Question 2",
    "a": ["Answer Here","<ul><li>text1</li><li>text2</li>"]
},
{
    "q": "Question 3",
    "a": ["Answer Here","<table>...</table>"]
},
{
    "q": "Question 4",
    "a": ["Answer"]
}]

我们将不胜感激任何帮助。在此先感谢:)

【问题讨论】:

标签: python html python-3.x


【解决方案1】:

试一试,它并不完美 - “可导航字符串”存在问题,我无法令人满意地解决。

from bs4 import Tag, NavigableString, BeautifulSoup


html_doc= """
<div data-axite-container="1" data-axite-uuid="1f0c5634-9ff9-4942-861e-2c7e75d6f2ef" data-axite-id="a4d1a127-0fe8-4281-abd6-fe0653a8b519">
<h4>Question 1</h4>
<p>Some text here too</p> <p>And some text here</p>
<h4>Question 2</h4>
<p> Answer Here</p>
<ul><li>text1</li><li>text2</li>
<h4>Question 3</h4>
<p>Answer Here</p><table>...</table>
<h4>Question 4</h4>
<p>Answer</p>
"""

soup = BeautifulSoup(html_doc, 'html.parser')

questions = soup.select('h4')


lst_questions = []
for tag in questions:
  lst = []
  for x in tag.next_siblings:
    if x.name == 'h4':
      break
    else:
      print(f'{str(x)}-{type(x)}')
      if isinstance(x, Tag):
        lst.append(x.string)
  dic = {'q': tag.string, 'a':lst}
  lst_questions.append(dic)
  
print(lst_questions)

【讨论】:

    【解决方案2】:

    复制自@Sayeed Hossain 非常清晰的解释性答案

    使用:Docs for the Python HTML parser module

    from html.parser import HTMLParser
    
    html_chunk =''' <div data-axite-container="1" data-axite-uuid="1f0c5634-9ff9-4942-861e-2c7e75d6f2ef" data-axite-id="a4d1a127-0fe8-4281-abd6-fe0653a8b519">
    <h4>Question 1</h4>
    <p>Some text here too</p> <p>And some text here</p>
    <h4>Question 2</h4>
    <p>Answer Here</p>
    <ul><li>text1</li><li>text2</li>
    <h4>Question 3</h4>
    <p>Answer Here</p><table>...</table>
    <h4>Question 4</h4>
    <p>Answer</p>'''
    
    
    class html_to_dict_parser(HTMLParser):
      def __init__(self):
        HTMLParser.__init__(self)
    
        self.in_word_label = False
        self.in_definition = False
        self.has_definition = False
    
      def handle_starttag(self, tag, attrs):
        if tag == 'h4':
          self.in_word_label  = True
        elif tag == 'p':
          self.in_definition  = True
    
    
      def handle_endtag(self, tag):
        if tag == 'h4':
          self.in_word_label = False
        elif tag == 'p':
          self.in_definition  = False
    
    
      def handle_data(self, data):
        if self.in_word_label:
          self.latest_word = data.lower()
          self.has_definition = True
        elif self.in_definition and self.has_definition:
          dictionary=dict()
          dictionary[ "q" ] = self.latest_word
          dictionary[ "a" ] = data
          list_of_dictionary.append(dictionary)
          self.has_definition = False
    
    
    # create empty list 
    list_of_dictionary = []
    
    # Run the parser!
    parser = html_to_dict_parser()
    
    parser.feed(html_chunk)
    
    parser.close()
    print(list_of_dictionary)
    
    

    输出:

    [{'q': 'question 1', 'a': 'Some text here too'}, {'q': 'question 2', 'a': 'Answer Here'}, {'q': 'question 3', 'a': 'Answer Here'}, {'q': 'question 4', 'a': 'Answer'}]
    
    

    【讨论】:

      猜你喜欢
      • 2017-05-30
      • 2019-10-16
      • 2019-01-10
      • 2015-07-23
      • 2012-07-12
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多