【问题标题】:Using Selenium and python to save a table使用 Selenium 和 python 保存表格
【发布时间】:2011-12-29 21:23:45
【问题描述】:

我正在尝试将 Selenium 与 Python 一起使用来存储表的内容。我的脚本如下:

import sys
import selenium
from selenium import webdriver
from selenium.webdriver.common.keys import Keys

driver = webdriver.Firefox()
driver.get("http://testsite.com")

value = selenium.getTable("table_id_10")

print value

driver.close()

这会打开我感兴趣的网页,然后应该保存我想要的表格内容。我在这个question 中看到了使用browser.get_table() 的语法,但是该程序的开头以browser=Selenium(...) 开头,我不明白。我不确定我应该使用什么语法作为 selenium.getTable("table_id_10") 不正确。


编辑:

我包含了我正在使用的表格的 html sn-p:

<table class="datatable" cellspacing="0" rules="all" border="1" id="table_id_10" style="width:70%;border-collapse:collapse;">
    <caption>
        <span class="captioninformation right"><a href="Services.aspx" class="functionlink">Return to Services</a></span>Data
    </caption><tr>
        <th scope="col">Read Date</th><th class="numericdataheader" scope="col">Days</th><th class="numericdataheader" scope="col">Values</th>

    </tr><tr>
        <td>10/15/2011</td><td class="numericdata">92</td><td class="numericdata">37</td>
    </tr><tr class="alternaterows">
        <td>7/15/2011</td><td class="numericdata">91</td><td class="numericdata">27</td>
    </tr><tr>
        <td>4/15/2011</td><td class="numericdata">90</td><td class="numericdata">25</td>    
</table>

【问题讨论】:

  • 如果您正在寻找网络抓取(我认为这就是您正在做的),您可能还想查看mechanize。我过去使用过它并且非常喜欢它,但不幸的是文档缺乏很多并且有点难以使用。只是一个想法,希望它不会离题。
  • @TKKocheran 我对抓取很感兴趣,尽管在这种情况下它只是一张桌子。我也可以保存 html 页面并稍后单独解析它。
  • 你可能想研究一下使用机械化。一旦你掌握了窍门,机械化在做事方面就会非常强大。我曾经写过一个脚本,它会登录到我的银行账户,回答一个安全问题,然后使用银行应用程序中可用的表格来获取我的财务数据。有趣的东西。

标签: python selenium


【解决方案1】:

旧的 Selenium RC API 包含一个 get_table 方法:

In [14]: sel=selenium.selenium("localhost",4444,"*firefox", "http://www.google.com/webhp")
In [19]: sel.get_table?
Type:       instancemethod
Base Class: <type 'instancemethod'>
String Form:    <bound method selenium.get_table of <selenium.selenium.selenium object at 0xb728304c>>
Namespace:  Interactive
File:       /usr/local/lib/python2.7/dist-packages/selenium/selenium.py
Definition: sel.get_table(self, tableCellAddress)
Docstring:
    Gets the text from a cell of a table. The cellAddress syntax
    tableLocator.row.column, where row and column start at 0.

    'tableCellAddress' is a cell address, e.g. "foo.1.4"

由于您使用的是较新的 Webdriver(a.k.a Selenium 2)API,因此该代码不适用。


也许可以试试这样的方法:

import selenium.webdriver as webdriver
import contextlib

@contextlib.contextmanager
def quitting(thing):
    yield thing
    thing.close()
    thing.quit()

with quitting(webdriver.Firefox()) as driver:
    driver.get(url)
    data = []
    for tr in driver.find_elements_by_xpath('//table[@id="table_id_10"]//tr'):
        tds = tr.find_elements_by_tag_name('td')
        if tds: 
            data.append([td.text for td in tds])
print(data)
# [[u'10/15/2011', u'92', u'37'], [u'7/15/2011', u'91', u'27'], [u'4/15/2011', u'90', u'25']]

【讨论】:

  • 感谢您解释为什么 get_table 不起作用。我添加了 html 表的 sn-p - 我正在尝试从单元格中提取所有值。
  • +1 包括contextlib。直到现在我才意识到这是一件事。我将不得不将我的所有脚本转换为以这种方式工作:D
  • @DuckPuncher:对于某些版本的 webdriver,你 may need to call both close and quit。我已经更新了上面的代码以显示一个上下文管理器,它可以同时进行。
  • @unutbu,这太棒了!
猜你喜欢
  • 2015-04-23
  • 2012-06-13
  • 2021-11-28
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-02-10
相关资源
最近更新 更多