【问题标题】:How to skip index out of range?如何跳过超出范围的索引?
【发布时间】:2023-03-07 07:58:01
【问题描述】:

我正在尝试进行学习练习以制作 ebay 列表刮擦,我的项目框长度为 48,但只有 26 个项目具有评级 div,我得到 IndexError: list index out of range,我该如何跳过这行或者如果 item_rating 是我该怎么写例如,空带“N/A”。我尝试继续,但无法修复。实际上,对于 item_shipping 等不同变量的这种情况,这是一个普遍问题。提前致谢。

更新

import requests
from bs4 import BeautifulSoup
import pandas as pd

URL='https://www.ebay.com/b/Makeup-Products/31786/bn_1865570' #'https://www.ebay.com/b/Makeup-Products/31786/bn_1865570' #https://www.ebay.com/b/Eye-Makeup/172020/bn_1880663
response=requests.get(URL)
soup= BeautifulSoup(response.content, 'html.parser')
columns=["Name","Price","Rating","Location"]
#Product features
main_table=soup.find('ul',attrs={'class':'b-list__items_nofooter'})
item_boxes=main_table.find_all('div',attrs={'class':'s-item__info clearfix'})
item = item_boxes[0]

df=pd.DataFrame(columns=columns)

for item in item_boxes:

    item_name = item.findAll('h3')
    try:
       item_name_row = item_name[0].text.replace('\n','')
    except:
       item_name = "N/A"


    item_price = item.find_all('span',{'class':'s-item__price'})
    try:
       item_price_row = item_price[0].text.replace('\n','')
    except:
       item_price_row = "N/A"      

    try:
       item_rating = item.findAll('div',{'class':'s-item__reviews'})[0].div
       item_rating_row = item_rating.text
    except:
       item_rating_row = None

    try:
       item_location = item_location = item.find_all('span',{'class':'s-item__location s-item__itemLocation'})[0]
       item_location_row = item_location.text
    except:
       item_location_row = None   

    row = [ item_name_row, item_price_row, item_rating_row, item_location_row ]
    df =df.append(pd.Series(row,index=columns),ignore_index=True)
    df.to_csv('ebay1.csv', index=False)


    if item_rating != None:

      row = [item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating.text.replace('\n','') for rating in item_rating] + [item_location_row[0].replace('\n','') for location in item_location]

    elif item_location != None:

      row = [item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating.text.replace('\n','') for rating in item_rating] + [item_location_row[0].replace('\n','') for location in item_location]
    else: 
      row = [ item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating] + [item_location_row]
    df =df.append(pd.Series(row,index=columns),ignore_index=True)

df.to_csv('ebay4.csv', index=False)

【问题讨论】:

  • Im getting IndexError: list index out of range 你从哪里得到错误?
  • in line item_rating = item.findAll('div',{'class':'s-item__reviews'})[0].div
  • 您尝试过什么来解决问题?
  • 我试过 if not rating 或 rating.startswith(''): rating='N/A' 但这会导致错误,因为它在定义变量之前。
  • 但这会导致错误,因为它在定义变量之前。 好的,我相信你可以弄清楚那个。您是否尝试过查看item.findAll('div',{'class':'s-item__reviews'}) 的结果是什么?然后item.findAll('div',{'class':'s-item__reviews'})[0]?

标签: python pandas web-scraping beautifulsoup


【解决方案1】:

给你,这是没有评级的列表:

import requests
from bs4 import BeautifulSoup
import pandas as pd

URL='https://www.ebay.com/b/Makeup-Products/31786/bn_1865570'
response=requests.get(URL)
soup= BeautifulSoup(response.content, 'html.parser')
columns=['name',"price","rating"]
#Product features
main_table=soup.find('ul',attrs={'class':'b-list__items_nofooter'})
item_boxes=main_table.find_all('div',attrs={'class':'s-item__info clearfix'})
item = item_boxes[0]

df=pd.DataFrame(columns=columns)

for item in item_boxes:

    item_name = item.findAll('h3')
    item_price = item.find_all('span',{'class':'s-item__price'})
    try:
       item_rating = item.findAll('div',{'class':'s-item__reviews'})[0].div
    except:
       item_rating = None
    if item_rating != None:
      row = [item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating.text.replace('\n','') for rating in item_rating]
    else: 
      row = [ item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating]
    df =df.append(pd.Series(row,index=columns),ignore_index=True)

df.to_csv('ebay1.csv', index=False)

这是我使用的一个,它是你的修改,用于刮取 recolorado 数据为邻居:

import requests
from bs4 import BeautifulSoup
import pandas as pd
URL='https://www.recolorado.com/find-real-estate/80817/1-pg/exclusive-dorder/price-dorder/photo-tab/'
response=requests.get(URL)
soup= BeautifulSoup(response.content, 'html.parser')
columns=['address',"price","active","bedrooms","bathrooms","sqft","courtesy"]
#Product features
main_table=soup.find('div',attrs={'class':'page--column', 'data-id':'listing-results'})
item_boxes=main_table.find_all('div',attrs={'class':'listing--information listing--information__photo'})
df=pd.DataFrame(columns=columns)

for item in item_boxes:

    price = item.find('li', attrs={'class': 'listing--detail listing--detail__photo listing--detail__price'})
    price_row = price.text.replace('\r','').replace('\n','').replace(' ', '')
    #print(price_row)

    address = item.find('h2', attrs={'class': 'listing--street listing--street__photo'})
    address_row = address.text.replace(', ', '')
    #print(address_row)

    active_listing = item.find('div', attrs={'class': 'listing--status listing--status__photo listing--status__Under Contract'})
    try:
       active_row = active_listing.text
    except:
       active_row = "N/A"
    #print(active_row)

    bedrooms = item.find('li', attrs={'class': 'listing--detail listing--detail__photo listing--detail__bedrooms'})
    try:
       bedrooms_row = bedrooms.text.replace('\r','').replace('\n','').replace(' ', '')
    except:
       bedrooms_row = "N/A"
    #print(bedrooms_row)

    bathrooms = item.find('li', attrs={'class': 'listing--detail listing--detail__photo listig--detail__bathrooms'})
    try:
       bathrooms_row = bathrooms.text.replace('\r','').replace('\n','').replace(' ', '')
    except:
       bathrooms_row = "N/A"
    #print(bathrooms_row)

    sqft = item.find('li', attrs={'class': 'listing--detail listing--detail__photo listing--detail__sqft'})
    try:
       sqft = item.find('li', attrs={'class': 'listing--detail listing--detail__photo listing--detail__sqft'})
       sqft_row = sqft.text.replace('\r','').replace('\n','').replace(' ', '')
    except:
       sqft_row = "N/A"
    #print(sqft_row)

    courtesy = item.find('div', attrs={'class': 'listing--courtesy listing--courtesy__photo show-mobile'})
    try:
       courtesy_row = courtesy.text.replace('\r','').replace('\n','').replace(' ', '')
    except:
       courtesy_row = "N/A"
    #print(courtesy_row)

    row = [ address_row, price_row, active_row, bedrooms_row, bathrooms_row, sqft_row, courtesy_row ]
    df =df.append(pd.Series(row,index=columns),ignore_index=True)

df
#                         address     price          active    bedrooms    bathrooms       sqft                                     courtesy
#0    6920 South US Highway 85-87  $699,000             N/A  5Bedrooms●  4Bathrooms●  3,978Sqft        CourtesyofColdwellBankerResidentialBK
#1               7095 Prado Drive  $414,900  Under Contract  9Bedrooms●  4Bathrooms●  3,000Sqft  CourtesyofKellerWilliamsClientsChoiceRealty
#2          7941 Whistlestop Lane  $399,500             N/A  3Bedrooms●  3Bathrooms●  2,577Sqft           CourtesyofRE/MAXRealEstateGroupInc
#3            7287 Van Wyhe Court  $389,900  Under Contract  4Bedrooms●  3Bathrooms●  2,750Sqft                         CourtesyofPinkRealty
#4   10737 Hidden Prairie Parkway  $369,900  Under Contract  4Bedrooms●  3Bathrooms●  2,761Sqft       CourtesyofKellerWilliamsPartnersRealty
#5            7327 Van Wyhe Court  $362,400             N/A  3Bedrooms●  2Bathrooms●  1,640Sqft                         CourtesyofPinkRealty
#6               7354 Chewy Court  $359,000             N/A  3Bedrooms●  2Bathrooms●  1,680Sqft      CourtesyofRedWhiteAndBlueRealtyGroupInc
#7           238 West Iowa Avenue  $355,000             N/A         N/A  4Bathrooms●  1,440Sqft                        CourtesyofAllenRealty
#8         8181 Wagon Spoke Trail  $350,000  Under Contract  4Bedrooms●  3Bathrooms●  2,848Sqft    CourtesyofKellerWilliamsPremierRealty,LLC
#9                     0 Missouri  $350,000             N/A         N/A          N/A        N/A                 CourtesyofRE/MAXNORTHWESTINC
#10  10817 Hidden Prairie Parkway  $340,000  Under Contract  3Bedrooms●  3Bathrooms●  2,761Sqft       CourtesyofKellerWilliamsPartnersRealty
#11         8244 Campground Drive  $335,000  Under Contract  4Bedrooms●  3Bathrooms●  2,018Sqft                         CourtesyofPinkRealty

我很快会尝试在这里重新抓取一个 ebay 网站,如果你有另一个链接,请将它留在 cmets 中,我很乐意看看我是否可以抓取它

更新:

在另一个页面上尝试过,它成功了

import requests
from bs4 import BeautifulSoup
import pandas as pd

URL='https://www.ebay.com/b/Eye-Makeup/172020/bn_1880663' #'https://www.ebay.com/b/Makeup-Products/31786/bn_1865570'
response=requests.get(URL)
soup= BeautifulSoup(response.content, 'html.parser')
columns=['name',"price","rating"]
#Product features
main_table=soup.find('ul',attrs={'class':'b-list__items_nofooter'})
item_boxes=main_table.find_all('div',attrs={'class':'s-item__info clearfix'})
item = item_boxes[0]

df=pd.DataFrame(columns=columns)

for item in item_boxes:

    item_name = item.findAll('h3')
    try:
       item_name_row = item_name[0].text.replace('\n','')
    except:
       item_name = "N/A"

    
    item_price = item.find_all('span',{'class':'s-item__price'})
    try:
       item_price_row = item_price[0].text.replace('\n','')
    except:
       item_price_row = "N/A"      
 
    try:
       item_rating = item.findAll('div',{'class':'s-item__reviews'})[0].div
       item_rating_row = item_rating.text
    except:
       item_rating_row = None

    row = [ item_name_row, item_price_row, item_rating_row ]
    df =df.append(pd.Series(row,index=columns),ignore_index=True)
    df.to_csv('ebay1.csv', index=False)


    if item_rating != None:
      row = [item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating.text.replace('\n','') for rating in item_rating]
    else: 
      row = [ item_name[0].text.replace('\n','') for name in item_name] + [item_price[0].text.replace('\n','') for price in item_price] + [item_rating]
    df =df.append(pd.Series(row,index=columns),ignore_index=True)

df.to_csv('ebay1.csv', index=False)


【讨论】:

  • 成功了,非常感谢,我想我可以对每个项目使用相同的方法
  • 艾伦,也谢谢你,我正在使用 Beautiful Soup 并学习如何解析,所以看到这些问题对我的想法也有帮助。我正在尝试从哈利波特维基中提取咒语,所以这个问题给了我一些想法 +1
  • 您好,再次感谢您的帮助,在那之后,我尝试添加另一个项目 item_location,首先我尝试了“或”,当然它工作错了,所以我从未尝试过和'。然后我又写了一个尝试,但是现在我遇到了值错误,你怎么看?
  • 艾伦,我刚刚修改了你的页面,所以我可以删除另一个页面,我会在他们的主要帖子下面发布它,(给我几分钟它会出现)也许它会给你一些想法,如果不是,也许你可以问一个新问题,因为我正在使用 BS4 来报废许多房地产数据,所以我可能会在 ebay 报废方面提供更多帮助。我可能会重试您的帖子并将我的帖子从房地产更改为另一个 ebay 报废以及我正在积极进行报废
  • 艾伦,我尝试了另一页的底部页面,它有效,有帮助吗?如果不打开另一个问题,因为我认为这个问题正在填满。并在此处发布链接,我一定会尝试一下。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2021-01-27
  • 2019-02-19
  • 2018-03-01
  • 1970-01-01
  • 1970-01-01
  • 2013-09-28
相关资源
最近更新 更多