【发布时间】:2019-03-19 13:30:18
【问题描述】:
我正在尝试从网页中获取数据。这是链接https://www.cardekho.com/compare-cars。在此页面中,一旦我们在下拉菜单中提供了汽车型号及其变体的 URL,我们就需要抓取汽车数据表及其规格的比较。这是我的示例代码。
from bs4 import BeautifulSoup
import requests
import csv
def job():
url = 'https://www.cardekho.com/compare/maruti-gypsy-and-maruti-omni.htm'
headers = {'User-Agent': 'Mozilla/65.0'}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')
stat_table_1 = soup.find_all('table')
print(len(stat_table_1))
tab_1 = stat_table_1[0]
tab_2 = stat_table_1[1]
tab_3 = stat_table_1[2]
tab_4 = stat_table_1[3]
tab_5 = stat_table_1[4]
rows_tab_1 = tab_1.findAll('tr')
rows_tab_2 = tab_2.findAll('tr')
rows_tab_3 = tab_3.findAll('tr')
rows_tab_4 = tab_4.findAll('tr')
rows_tab_5 = tab_5.findAll('tr')
csv_file_1 = open("D:/CarDekho_Data/maruti/maruti_2/overview.csv", 'wt', encoding="utf-8", newline='')
csv_file_2 = open("D:/CarDekho_Data/maruti/maruti_2/engine.csv", 'wt', encoding="utf-8", newline='')
csv_file_3 = open("D:/CarDekho_Data/maruti/maruti_2/transmission.csv", 'wt', encoding="utf-8", newline='')
csv_file_4 = open("D:/CarDekho_Data/maruti/maruti_2/steering.csv", 'wt', encoding="utf-8", newline='')
csv_file_5 = open("D:/CarDekho_Data/maruti/maruti_2/brake_system.csv", 'wt', encoding="utf-8", newline='')
writer_1 = csv.writer(csv_file_1)
writer_2 = csv.writer(csv_file_2)
writer_3 = csv.writer(csv_file_3)
writer_4 = csv.writer(csv_file_4)
writer_5 = csv.writer(csv_file_5)
try:
for row in rows_tab_1:
csv_row = []
for cell in row.findAll(['td', 'th']):
csv_row.append(cell.get_text())
writer_1.writerow(csv_row)
finally:
csv_file_1.close()
try:
for row in rows_tab_2:
csv_row = []
for cell in row.findAll(['td', 'th']):
csv_row.append(cell.get_text())
writer_2.writerow(csv_row)
finally:
csv_file_2.close()
try:
for row in rows_tab_3:
csv_row = []
for cell in row.findAll(['td', 'th']):
csv_row.append(cell.get_text())
writer_3.writerow(csv_row)
finally:
csv_file_3.close()
try:
for row in rows_tab_4:
csv_row = []
for cell in row.findAll(['td', 'th']):
csv_row.append(cell.get_text())
writer_4.writerow(csv_row)
finally:
csv_file_4.close()
try:
for row in rows_tab_5:
csv_row = []
for cell in row.findAll(['td', 'th']):
csv_row.append(cell.get_text())
writer_5.writerow(csv_row)
finally:
csv_file_5.close()
但这里的问题是,由于 URL 的原因,我没有得到我需要的确切数据。这意味着如果我给出四种车型及其变体进行比较,它是从提到的下拉菜单中随机给出该车型的数据。
谁能解释我如何解决这个问题并从该 URL 获取我需要的确切数据。
任何帮助将不胜感激。
【问题讨论】:
标签: python-3.x web-scraping beautifulsoup request