【问题标题】:Scraped data to csv file beautifulsoup将数据抓取到 csv 文件 beautifulsoup
【发布时间】:2020-01-27 17:38:05
【问题描述】:

正如标题所说,我使用 beautifulsoup scraper 从网站抓取数据工作正常,但是当我尝试将数据加载到 csv 文件中时,它只保存 1 个数据区域,而不是 500 个数据区域,scraper 给出的是我的代码:

#from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup
import csv



#launch url
url = "https://www.canlii.org/en/#search/type=decision&jId=bc,ab,sk,mb,on,qc,nb,ns,pe,nl,yk,nt,nu&startDate=1990-01-01&endDate=1992-01-14&text=non-pecuniary%20award%20&resultIndex=1"

# create a new Chrome session
driver = webdriver.Chrome('C:\Program Files (x86)\Microsoft Visual Studio\Shared\Anaconda3_64\lib\site-packages\selenium\webdriver\common\chromedriver.exe')
driver.implicitly_wait(30)
driver.get(url)


#Selenium hands the page source to Beautiful Soup
soup=BeautifulSoup(driver.page_source, 'lxml')

csv_file = open('test.csv', 'w')

csv_writer = csv.writer(csv_file, quoting=csv.QUOTE_ALL)
csv_writer.writerow(['Reference', 'case', 'link', 'province', 'keywords','snippets'])

#Scrape all
for scrape in soup.find_all('li', class_='result '):
    print(scrape.text)    
    
#Reference Index
    Reference = scrape.find('span', class_='reference')
    print(Reference.text)

#Case Name Index
    case = scrape.find('span', class_='name')
    print(case.text)
    
#Canlii Keywords Index
    keywords = scrape.find('div', class_='keywords')
    print(keywords.text)
    
#Province Index
    province = scrape.find('div', class_='context')
    print(province.text)
               
#snippet Index
    snippet = scrape.find('div', class_='snippet')
    print(snippet.text)
 
# Extracting URLs from the attribute href in the <a> tags.
    link = scrape.find('a', href=True)
    print(link)        
            
csv_writer.writerow([Reference.text, case.text,link.href, province.text, keywords.text, snippet.text])
csv_file.close()

【问题讨论】:

  • 好吧,我解决了 csv 编写器的问题 enrique 是正确的,它已超出循环,然后出现编码错误,但我有一个新问题,如果我想向现有 csv 添加更多内容我去吧
  • 如果我想学习你的教程,我会在我的 linux 机器上遇到问题 -` Traceback(最近一次调用最后一次):文件“/home/martin/.atom/python/examples/bs_canlii. py",第 10 行,在 driver = webdriver.Chrome('C:\Program Files (x86)\Microsoft Visual Studio\Shared\Anaconda3_64\lib\site-packages\selenium\webdriver\common\chromedriver.exe' ) NameError: name 'webdriver' is not defined` - 我在 MX-Linux 上运行 ATOM - 我猜你是在一台 win 机器上

标签: python csv beautifulsoup


【解决方案1】:

你的 csv_writer.writerow() 函数在你的 for 循环之外。尝试缩进它,看看它是否有效。

【讨论】:

  • 当我在 csv_writer.writerow([Reference.text, case.text,link .href、province.text、keywords.text、sn-p.text]) ValueError: I/O operation on closed file.
猜你喜欢
  • 2017-06-18
  • 2018-06-06
  • 2016-08-08
  • 1970-01-01
  • 2020-12-08
  • 1970-01-01
  • 1970-01-01
  • 2021-08-22
  • 1970-01-01
相关资源
最近更新 更多