【问题标题】:python html table data parsingpython html表格数据解析
【发布时间】:2017-03-03 15:49:13
【问题描述】:

我正在寻找示例代码以从下面提到的 html 表代码中检索干净的输出。

<td width="40%" valign="top" colSpan="1" style="padding-left:5px;padding-top:4px"><input type="text" id="subscriberDeletionForm_OSILA_DisplayName" name="OSILA_DisplayName" class="roTextField" style="width:300px;" value="Rashmi HK" readonly="readonly" tabindex="-1"/>
</td>
<td colspan="1">&nbsp;</td>
</tr>
<tr>
<td class="firstLabelInRowCell"><label id="subscriberDeletionForm_OSILA_CountryCode_label" for="subscriberDeletionForm_OSILA_CountryCode" class="rwLabel">Country Code:</label>
</td>
<td width="40%" valign="top" colSpan="1" style="padding-left:5px;padding-top:4px"><input type="text" id="subscriberDeletionForm_OSILA_CountryCode" name="OSILA_CountryCode" class="roTextField" style="width:300px;" value="91" readonly="readonly" tabindex="-1"/>
</td>
<td class="labelCell"><label id="subscriberDeletionForm_OSILA_AreaCode_label" for="subscriberDeletionForm_OSILA_AreaCode" class="rwLabel">Area Code:</label>
</td>
<td width="40%" valign="top" colSpan="1" style="padding-left:5px;padding-top:4px"><input type="text" id="subscriberDeletionForm_OSILA_AreaCode" name="OSILA_AreaCode" class="roTextField" style="width:300px;" value="80" readonly="readonly" tabindex="-1"/>

我需要的输出。

OSILA_DisplayName  = Rashmi HK
OSILA_CountryCode  = 91                                                                               OSILA_AreaCode     = 80 

我正在使用以下代码并且能够检索它。但我需要以相同的方式提取大量字段,因此我正在寻找不同的方式来提取

    OSILA_DisplayName = 'id="subscriberDeletionForm_OSILA_DisplayName"'
    f22 = open('delsubinfo1', 'r')
    for line2 in f22:
        if OSILA_DisplayName in line2:
#            print line2
            line2 = line2.split('"')
#            print line2
            OSILA_DisplayName1 = line2[19].strip()
            print  OSILA_DisplayName1

    OSILA_CountryCode = 'name="OSILA_CountryCode"'
    f23 = open('delsubinfo1', 'r')
    for line3 in f23:
        if OSILA_CountryCode in line3:
#            print line3
            line3 = line3.split('"')
#            print line3
            OSILA_CountryCode1 = line3[19].strip()
            print OSILA_CountryCode1

【问题讨论】:

  • 你的代码在哪里?
  • 你试过什么了吗?
  • 已更新代码
  • BeautifulSoup 可能是您应该用来正确从 HTML 中抓取数据的代码库。

标签: python html


【解决方案1】:

我推荐使用BeautifulSoup来解析html文本。

您可以像这样从 html 代码中检索数据。

from bs4 import BeautifulSoup

txt = open('delsubinfo1', 'r').read()
soup = BeautifulSoup(txt, "html.parser")
for input_tag in soup.find_all("input"):
    if input_tag.get("name") in ('OSILA_DisplayName', 'OSILA_CountryCode'):
        print input_tag.get("name").ljust(18), '=', input_tag.get("value")
# OSILA_DisplayName  = Rashmi HK
# OSILA_CountryCode  = 91

【讨论】:

  • uehara,我们可以将名称和值作为变量传递以进行进一步处理
  • 是的,你可以。您可以获取更多信息here。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2021-09-08
  • 2017-05-18
  • 2013-01-02
  • 1970-01-01
  • 2014-09-23
  • 2018-08-31
  • 2021-11-15
相关资源
最近更新 更多