【问题标题】:Python Loop Overwriting Last HTML Write [duplicate]Python循环覆盖最后的HTML写入[重复]
【发布时间】:2018-11-08 08:11:37
【问题描述】:

此脚本在“while True:”处循环,通过单击底部的下一步按钮从多个页面中抓取数据,但我无法弄清楚如何构造代码以在分页时继续写入 HTML。相反,它会覆盖之前编写的 html 结果。感谢您的帮助。谢谢!

while True:
    time.sleep(10)

    golds = driver.find_elements_by_css_selector(".widgetContainer #widgetContent > div.singleCell")
    print("found %d golds" % len(golds))  

    template = """\
        <tr class="border">
            <td class="image"><img src="{0}"></td>\
            <td class="title"><a href="{1}" target="_new">{2}</a></td>\
            <td class="price">{3}</td>
        </tr>"""

    lines = []

    for gold in golds:
        goldInfo = {}

        goldInfo['title'] = gold.find_element_by_css_selector('#dealTitle > span').text
        goldInfo['link'] = gold.find_element_by_css_selector('#dealTitle').get_attribute('href')
        goldInfo['image'] = gold.find_element_by_css_selector('#dealImage img').get_attribute('src')

        try:
            goldInfo['price'] = gold.find_element_by_css_selector('.priceBlock > span').text
        except NoSuchElementException:
            goldInfo['price'] = 'No price display'

        line = template.format(goldInfo['image'], goldInfo['link'], goldInfo['title'], goldInfo['price'])
        lines.append(line)

    try:
        #clicks next button
        driver.find_element_by_link_text("Next→").click()
    except NoSuchElementException:
        break

    time.sleep(10)

    html = """\
        <html>
            <body>
                <table>
                    <tr class='headers'>
                        <td class='image'></td>
                        <td class='title'>Product</td>
                        <td class='price'>Price / Deal</td>
                    </tr>
                </table>
                <table class='data'>
                    {0}
                </table>
            </body>
        </html>\
    """

    f = open('./result.html', 'w')
    f.write(html.format('\n'.join(lines)))
f.close()

【问题讨论】:

  • 尝试用f = open('./result.html', 'a')替换f = open('./result.html', 'w')
  • 谢谢安德森,完美。您想将其发布为我可以接受该答案的答案吗?
  • 我认为this answer 信息量更大

标签: python selenium


【解决方案1】:

查看在脚本末尾打开文件时的不同模式:https://docs.python.org/2/library/functions.html#open

mode 最常用的值是 'r' 用于读取,'w' 用于写入(如果文件已存在则截断文件)和 'a' 用于追加

还有更多

模式'r+'、'w+'和'a+'打开文件进行更新(读写);请注意,'w+' 会截断文件。在区分二进制文件和文本文件的系统上,将“b”附加到模式以二进制模式打开文件;在没有这种区别的系统上,添加“b”无效。

所以你有几个可用的选项。您可以使用a,因为您想将数据附加到它。

或者您可以将打开的文件移到循环之外,这样您就不会经常重新打开文件,这取决于您的需要。

f = open('./result.html', 'w')
while True:
  # do stuff
  f.write (...)
f.close()

【讨论】:

    【解决方案2】:

    您应该以追加模式打开文件

    f = open('./result.html', 'a')
    

    【讨论】:

      猜你喜欢
      • 2019-09-24
      • 1970-01-01
      • 2016-07-17
      • 2014-10-09
      • 1970-01-01
      • 1970-01-01
      • 2014-09-02
      • 2018-04-22
      相关资源
      最近更新 更多