【问题标题】:Web Scraping with Beautiful Soup in Python - JavaScript Table使用 Python 中的 Beautiful Soup 进行网页抓取 - JavaScript 表
【发布时间】:2017-10-05 19:34:04
【问题描述】:

我试图从网站上抓取一张表格,但我似乎无法使用 Python 中的 Beautifulsoup 来解决这个问题。我不确定是不是因为表格格式,但我基本上想把这张表格变成 CSV。

from bs4 import BeautifulSoup
import requests

page = requests.geenter code heret("https://spotwx.com/products/grib_index.php?model=hrrr_wrfprsf&lat=41.03399&lon=-73.76291&tz=America/New_York&display=table")
soup = BeautifulSoup(page.content, 'html.parser')
print(soup.prettify)

关于如何隔离此数据表的任何建议?我查看了很多 Beautifulsoup 教程,但 HTML 看起来与大多数参考资料不同。非常感谢您的帮助 -

【问题讨论】:

    标签: beautifulsoup python-requests prettify


    【解决方案1】:

    试试这个。该站点的表格是动态生成的,因此您无法仅使用 requests 获得结果。

    from selenium import webdriver
    from bs4 import BeautifulSoup
    import csv
    
    link = "https://spotwx.com/products/grib_index.php?model=hrrr_wrfprsf&lat=41.03399&lon=-73.76291&tz=America/New_York&display=table"
    
    with open("spotwx.csv", "w", newline='') as infile:
        writer = csv.writer(infile)
        writer.writerow(['DateTime','Tmp','Dpt','Rh','Wh','Wd','Wg','Apcp','Slp'])
        with webdriver.Chrome() as driver:
            driver.get(link)
            soup = BeautifulSoup(driver.page_source, 'lxml')
            for item in soup.select("table#example tbody tr"):
                data = [elem.text for elem in item.select('td')]
                print(data)
                writer.writerow(data)
    

    【讨论】:

    • 非常感谢您的回复。我不熟悉 Webdriver,但我不需要它来实时刷新(除非绝对必要,否则我宁愿不使用 Webdriver)。看来,简单地进行请求拉动会在 soup.prettify 代码中显示必要的数据,但我只是不知道如何将其提取到表中。再次感谢您的帮助!
    • 当我尝试上面的代码时,我得到错误 selenium.common.exceptions.WebDriverException: Message: 'chromedriver' executable needs to be in PATH。请看sites.google.com/a/chromium.org/chromedriver/home
    • 第一个应该可以。如果没有,那就去第二个。 1.driver = webdriver.Chrome('C:/path/to/chromedriver.exe') 2.driver = webdriver.Chrome('/path/to/chromedriver') .Btw,你必须根据你的系统尝试,我是说路径。谢谢。
    • Shahin 这很棒,完美运行。非常感谢您的帮助。我只是想确认没有在 chrome 中启动它就没有办法刮掉它,因为我确实看到了我试图在 requests.get pull 中隔离的数据。谢谢
    • 我也不是硒的忠实粉丝。然而,在处理启用 javascript 的网站时,selenium 是首屈一指的,requests 也无济于事。顺便说一句,你可能喜欢 Phantomjs 这样的无头浏览器,但这种无头浏览器自动化存在几个问题。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-09-04
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多