【问题标题】:Python Web Scraping: Duplication and output display problemPython Web Scraping:重复和输出显示问题
【发布时间】:2020-05-06 23:50:01
【问题描述】:

我的代码有问题,我尝试过但无法识别。它与未正确显示和插入我的数据库的循环的输出有关。 我希望将每行刮取的数据打印为输出,然后插入到数据库表中。到目前为止,我得到的只是一个重复多次打印的结果(甚至没有合适的价格)。

实际电流输出:

Ford C-MAX 2019 1.1 Petrol 0
Ford C-MAX 2019 1.1 Petrol 0
Ford C-MAX 2019 1.1 Petrol 0
...

根据网页广告的期望输出(只是一个示例,因为它是动态的):

Ford C-MAX 2019 1.1 Petrol 15950
Ford C-MAX 2014 1.6 Diesel 12000
Ford C-MAX 2011 1.6 Diesel 9000
...

代码:

from __future__ import print_function
import requests
import re
import locale
import time
from time import sleep
from random import randint
from currency_converter import CurrencyConverter
c = CurrencyConverter()
from bs4 import BeautifulSoup
import pandas as pd
from datetime import date, datetime, timedelta
import mysql.connector
import numpy as np
import itertools

locale.setlocale( locale.LC_ALL, 'en_US.UTF-8' )

pages = np.arange(0, 210, 30)

entered = datetime.now()
make = "Ford"
model = "C-MAX"


def insertvariablesintotable(make, model, year, liter, fuel, price, entered):
    try:
        cnx = mysql.connector.connect(user='root', password='', database='FYP', host='127.0.0.2', port='8000')
        cursor = cnx.cursor()

        cursor.execute('CREATE TABLE IF NOT EXISTS ford_cmax ( make VARCHAR(15), model VARCHAR(20), '
                       'year INT(4), liter VARCHAR(3), fuel VARCHAR(6), price INT(6), entered TIMESTAMP) ')

        insert_query = """INSERT INTO ford_cmax (make, model, year, liter, fuel, price, entered) VALUES (%s,%s,%s,%s,%s,%s,%s)"""
        record = (make, model, year, liter, fuel, price, entered)

        cursor.execute(insert_query, record)

        cnx.commit()

    finally:
        if (cnx.is_connected()):
            cursor.close()
            cnx.close()

for response in pages:

    headers = {
        'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'}
    response = requests.get("https://www.donedeal.ie/cars/Ford/C-MAX?start=" + str(response), headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')

    cnx = mysql.connector.connect(user='root', password='', database='FYP', host='127.0.0.2', port='8000')
    cursor = cnx.cursor()

    for details in soup.findAll('ul', attrs={'class': 'card__body-keyinfo'}):

        details = details.text
        #print(details)
        year = details[:4]
        liter = details[4:7]
        fuel = details[8:14] #exludes electric which has 2 extra
        mileage = re.findall("[0-9]*,[0-9][0-9][0-9]..." , details)
        mileage = ''.join(mileage)
        mileage = mileage.replace(",", "")
        if "mi" in mileage:
            mileage = mileage.rstrip('mi')
            mileage = round(float(mileage) * 1.609)
        mileage = str(mileage)
        if "km" in mileage:
            mileage = mileage.rstrip('km')
        mileage = mileage.replace("123" or "1234" or "12345" or "123456", "0")

    for price in soup.findAll('p', attrs={'class': 'card__price'}):

        price = price.text
        price = price.replace("No Price", "0")
        price = price.replace("123" or "1234" or "12345" or "123456", "0")
        price = price.replace(",","")
        price = price.replace("€", "")
        if "p/m" in price:
            #price = price[:-3]
            price = price.rstrip('p/m')
            price = "0"
        if "£" in price:
            price = price.replace("£", "")
            price = c.convert(price, 'GBP', 'EUR')
            price = round(price)

    print(make, model, year, liter, fuel, price)

    #insertvariablesintotable(make, model, year, liter, fuel, price, entered) #same result as above

【问题讨论】:

    标签: python mysql beautifulsoup screen-scraping


    【解决方案1】:

    我查看了您的代码和您尝试从中获取数据的网站,看起来您正在检索该页面,然后使用 price 循环访问您从该页面获得的所有价格作为一个变量,但每次进入 for 循环时都会覆盖它。您的详细信息 for 循环也是如此。

    您可以尝试以下方法:

    make = "Ford"
    model = "C-MAX"
    price_list = [] # we will store prices here
    details_list = [] # and details like year, liter, mileage there
    for response in range(1,60,30): # I changed to a range loop for testing
    
        headers = {
            "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36"
        }
        response = requests.get(
            "https://www.donedeal.ie/cars/Ford/C-MAX?start=" + str(response),
            headers=headers,
        )
        soup = BeautifulSoup(response.text, "html.parser")
        count = 0
        for details in soup.findAll("ul", attrs={"class": "card__body-keyinfo"}):
            if count == 30:
                break # Takes us out of the for loop
            details = details.text
            # print(details)
            year = details[:4]
            liter = details[4:7]
            fuel = details[8:14]  # exludes electric which has 2 extra
            mileage = re.findall("[0-9]*,[0-9][0-9][0-9]...", details)
            mileage = "".join(mileage)
            mileage = mileage.replace(",", "")
            if "mi" in mileage:
                mileage = mileage.rstrip("mi")
                mileage = round(float(mileage) * 1.609)
            mileage = str(mileage)
            if "km" in mileage:
                mileage = mileage.rstrip("km")
            mileage = mileage.replace("123" or "1234" or "12345" or "123456", "0")
            details_list.append((year, liter, fuel, mileage)) # end of one loop go-through, we append
            count += 1 We update count value 
        count = 0
        for price in soup.findAll("p", attrs={"class": "card__price"}):
            if count == 30:
                break # Takes us out of the for loop
            price = price.text
            price = price.replace("No Price", "0")
            price = price.replace("123" or "1234" or "12345" or "123456", "0")
            price = price.replace(",", "")
            price = price.replace("€", "")
            if "£" in price:
                price = price.replace("£", "")
                price = c.convert(price, "GBP", "EUR")
                price = round(price)
            if "p/m" in price:
                # price = price[:-3]
                price = price.rstrip("p/m")
                price = "0"
            else:
                price_list.append(price) # end of loop go-through, we append but only if it is not a "p/m" price
                count += 1 # We update count value only when a value is appended to the list
    
    for i in range(len(price_list)):
        print(
        make,
        model,
        details_list[i][0],
        details_list[i][1],
        details_list[i][2],
        price_list[i],
    )
        #add your insertvariablesintotable(make,model,details_list[i][0], details_list[i][1],details_list[i][2],price_list[i]) there
    

    编辑:我没有将 p/m 价格添加到列表中,因为它们使 details_list 和 price_list 的长度不同。如果您还想添加 p/m 价格,则必须重新编写代码。此外,您不想让这 3 辆车出现在页面的最底部,因为它们可能不是福特 C-MAX,而是其他车型,甚至可能是其他制造商。

    【讨论】:

    • 谢谢先生,但是我试过了,我得到的输出是:` ... 15950 ('2011', '1.6', 'Diesel') 2750 ('2018', '1.0', 'Petrol') 1 ('2017', '1.2', 'Petrol') 7950 ('2019', '1.1', 'Petrol') Traceback(最近一次调用最后):文件“C:/xx.py”,第 122 行,在 print(price_list[i], details_list[i]) # 你可以在那里添加插入 IndexError: list index out of range` 而且我也不知道如何在下面添加插入,并且细节现在是一个列表而不是单个变量?
    • 问题是您实际上每页获得 33 个汽车详细信息(列表中的 30 个 + 页面底部的 3 个)和不同数量的价格,因为您还包括 p/m价格。我将编辑我的答案以帮助您实现它。价格和详细信息确实存储在列表中。
    • 谢谢先生!打印和插入工作正常,但结果仍然包括每个页面迭代的最后 3 个(不同模型)。我将前两页的 np.arange 设置为 (0,58,29),如果我更改此设置,则 url 不会从下一页刮掉,因为第 1 页是 ?start=0 而第 2 页是?start=29 等。如何排除最后 3 个结果?
    • 每次调用 API 时,都会返回 33 个结果,因此您无法通过调整 start 参数来更改此行为。但是,您可以做的是检查何时在您的价目表中添加了 30 个有效项目(不包括 p/m 价格)和在您的详细信息列表中添加了 30 个有效项目。查看我的更新答案
    • 几乎完美,实际上它会在 30 之后切断循环,但只占用前 30 并切断整个程序,并且不会刮掉第 2,3 页等。有什么解决办法吗?
    猜你喜欢
    • 1970-01-01
    • 2020-09-19
    • 2021-12-25
    • 1970-01-01
    • 2018-04-07
    • 2019-09-06
    • 2016-12-17
    • 1970-01-01
    • 2019-05-22
    相关资源
    最近更新 更多