【问题标题】:How to save the scraped results into a CSV file from multiple website pages?如何将抓取的结果从多个网站页面保存到 CSV 文件中?
【发布时间】:2020-01-15 11:58:02
【问题描述】:

我正在尝试使用 selenium 和 beautifulsoup 从亚马逊网站(只是 ASIN)上抓取一些 ASIN(比如说 600 个 ASIN)。我的主要问题是如何将所有抓取的数据保存到 CSV 文件中?我尝试了一些方法,但它只保存了最后抓取的页面。

代码如下:

from time import sleep
import requests
import time
import json
import re
import sys
import numpy as np
from selenium import webdriver
import urllib.request
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.common.keys import Keys
import pandas as pd
from urllib.request import urlopen


i = 1
while(True):
    try:
        if i == 1:
            url = "https://www.amazon.es/s?k=doll&i=toys&rh=n%3A599385031&dc&page=1"
        else:
            url = "https://www.amazon.es/s?k=doll&i=toys&rh=n%3A599385031&dc&page={}".format(i)
        r = requests.get(url)
        soup = BeautifulSoup(r.content, 'html.parser')

        #print page url
        print(url)

        #rest of the scraping code
        driver = webdriver.Chrome()
        driver.get(url)

        HTML = driver.page_source
        HTML1=driver.page_source
        soup = BeautifulSoup(HTML1, "html.parser")
        styles = soup.find_all(name="div", attrs={"data-asin":True})
        res1 = [i.attrs["data-asin"] for i in soup.find_all("div") if i.has_attr("data-asin")]
        print(res1)
        data_record.append(res1)
        #driver.close()

        #don't overflow website
        sleep(1)

        #increase page number
        i += 1
        if i == 3:
            print("STOP!!!")
            break
    except:
        break



【问题讨论】:

  • 您只需要检查您的print (res1) 是否有您需要的asin 值,然后将它们存储在一个csv 文件中。您是否尝试复制此站点的功能:asintool.com?
  • print(res1) 给出第 1 页的 ASIN,然后显示第 2 页的 ASIN。我只想保存刮掉的页面中的所有 Asin。
  • 是的,我想从特定关键字中检索 ASIN。
  • 好的,首先在循环之前创建一个空数据帧,如df = pd.DataFrame([]),然后在循环中添加:df = df.append(res1),然后将数据帧导出到 csv,如下所示:df.to_csv。让我知道它是否有效。

标签: python selenium web-scraping


【解决方案1】:

删除目前似乎不使用的项目可能是一种可能的解决方案

import csv
import bs4
import requests
from selenium import webdriver
from time import sleep


def retrieve_asin_from(base_url, idx):
    url = base_url.format(idx)
    r = requests.get(url)
    soup = bs4.BeautifulSoup(r.content, 'html.parser')

    with webdriver.Chrome() as driver:
        driver.get(url)
        HTML1 = driver.page_source
        soup = bs4.BeautifulSoup(HTML1, "html.parser")
        res1 = [i.attrs["data-asin"]
                for i in soup.find_all("div") if i.has_attr("data-asin")]
    sleep(1)
    return res1


url = "https://www.amazon.es/s?k=doll&i=toys&rh=n%3A599385031&dc&page={}"
data_record = [retrieve_asin_from(url, i) for i in range(1, 4)]

combined_data_record = combine_records(data_record) # fcn to write

with open('asin_data.csv', 'w', newline='') as fd:
    csvfile = csv.writer(fd)
    csvfile.writerows(combined_data_record)

【讨论】:

  • 你让它看起来很简单 :) 它可以工作,但是我应该如何调整以便数据可以在同一列中?
  • 您可以为 data_record 添加一个函数 - 返回一个新数组 - 在其中组合必要的字段。然后可以在编写 csv 文件时使用新数组。
猜你喜欢
  • 2020-08-22
  • 1970-01-01
  • 1970-01-01
  • 2016-03-02
  • 2018-05-23
  • 1970-01-01
  • 2021-11-23
  • 2021-10-31
  • 1970-01-01
相关资源
最近更新 更多