【发布时间】:2021-12-02 19:44:09
【问题描述】:
我一直在构建这个抓取工具(这里的用户提供了一些巨大的帮助)来获取数据,但最后我已经停止了。基本上我想使用抓取的信息生成一个字典并将其附加到一个列表中,但最终的列表最终只是我的数据的尾部而不是整个数据。到目前为止,这是我的代码:
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
import time
import math
#Defining important variables
path_driver = "C:/Users/CS330584/Documents/Documentos de Defesa da Concorrência/Automatização de Processos/chromedriver.exe"
website = "https://sat.sef.sc.gov.br/tax.NET/Sat.Dva.Web/ConsultaPublicaDevedores.aspx"
value_search = 300
final_table = []
#Opening the browser
driver = webdriver.Chrome(path_driver)
driver.get(website)
search_max = driver.find_element_by_id("Body_Main_Main_ctl00_txtTotalDevedores")
search_max.send_keys(value_search)
btn_consult = driver.find_element_by_id("Body_Main_Main_ctl00_btnBuscar")
btn_consult.click()
driver.implicitly_wait(10)
pages = math.ceil(value_search/50)
#Loop to scrape data, store it into a dict and append to list to generate a dataframe later
for i in range(2, pages+1):
try:
time.sleep(2)
cnpjs = driver.find_elements_by_xpath("//*[@id='Body_Main_Main_grpDevedores_gridView']/tbody/tr/td[1]")
empresas = driver.find_elements_by_xpath("//*[@id='Body_Main_Main_grpDevedores_gridView']/tbody/tr/td[2]")
dividas = driver.find_elements_by_xpath("//*[@id='Body_Main_Main_grpDevedores_gridView']/tbody/tr/td[3]")
for z in range(len(empresas)):
temp_data = {'CNPJ' : cnpjs[z].text,
'Empresas' : empresas[z].text,
'Divida' : dividas[z].text
}
final_table.append(temp_data)
driver.execute_script(f"javascript:GridView_ScrollToTop('Body_Main_Main_grpDevedores_gridView');__doPostBack('ctl00$ctl00$ctl00$Body$Main$Main$grpDevedores$gridView','Page${i}')")
except Exception as ex:
print(ex)
break
如何才能使 final_table 列表成为循环中生成的 dicts 的聚合?
【问题讨论】:
-
这里有很多代码可以解决听起来很简单的问题。你能把它减少到stackoverflow.com/help/minimal-reproducible-example
-
如果你按照@defladamouse 的建议去做,我想你会看到你的
final_table.append(temp_data)行只添加了最后一个,因为你没有在for z in range(len(empresas)):中这样做,我想?
标签: python selenium web-scraping