【发布时间】:2021-05-01 13:19:08
【问题描述】:
我正在尝试从网站上抓取交互式图表,这个脚本可以工作并从图表中收集文本,但速度很慢。为了使文本出现,光标必须悬停在图形上的某些位置。有人对如何提高效率有任何建议吗?
现在,对于每个偏移动作,它都会在前进到下一个动作之前遍历所有先前的动作。有谁知道如何绕过它? (例如,它从 0 到 5,然后不是从 5 到 10,而是再次回到 0,然后是 5,然后是 10)
# set the pace at which the cursor will move and the limit to which it will move to
# (which should be the current date or the x-axis limit of the image)
# set the limit to be current date
limit = full_length
pace = 5
count = 0
while count <= limit:
value = driver.find_element_by_class_name('highcharts-tooltip').text
date_price = value.split("\n")
date = date_price[0]
price = date_price[1].split(": ")
price = price[1]
# take values at current point and add to dictionary
dp = {'date': date,
'price': price }
archived_prices.append(dp)
# move to the next date
action.move_by_offset(pace, 0).perform()
# set up a counter to figure out when we will reach the limit
count = count + pace
print(count)
新代码:
# adding new column with complete url for api call
full_urls = []
for value in dataframe['urlKeys']:
full = 'https://stockx.com/api/products/'+value+'?includes=market,360¤cy=EUR&country=IT'
full_urls.append(full)
dataframe['urlFull'] = full_urls
def get_shoe_info(url_list):
for url in url_list:
headers = {
"accept-encoding": "gzip, deflate, br",
"sec-fetch-mode": "cors",
"sec=fetch-site": "same-origin",
"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36",
"x-requested-with": "XMLHttpRequest"
}
response = requests.get(url, headers=headers)
response.raise_for_status()
product = response.json()["Product"]
for p in product:
id_num = p["id"]
brand = p["brand"]
colorway = p["colorway"]
release_date = p["releaseDate"]
retail_price = p["retailPrice"]
shoe_name = p["shoe"]
volatility = p["market"]["volatility"]
change_percentage = p["market"]["changePercentage"]
gender = p["gender"]
print(f'ID: {id_num}\n'
f'Brand: {brand}\n'
f'Colorway: {colorway}\n'
f'Release Date: {release_date}\n'
f'Retail Price: {retail_price}\n'
f'Shoe Name: {shoe_name}\n'
f'Volatility: {volatility}\n'
f'Change Percentage: {change_percentage}\n'
f'Gender: {gender}')
#print(p)
return 0
if __name__ == "__main__":
import sys
sys.exit(get_shoe_info(full_urls))
我仍然难以理解如何传递变量,所以我不确定我是否正确地做到了这一点。第一部分是我获取所有鞋子的 url 键并创建一个 url 列表以进行迭代。然后我试图将列表传递给 get_shoe_info 函数。我的错误弹出为“TypeError:字符串索引必须是整数”,我调查并看到当我尝试 print(p) 查看路径时,我只得到了关键部分的字符串。我不确定如何获得我想要的值。
我已将所有内容添加到 (my github),以防您需要查看其他内容。
【问题讨论】:
-
可以分享一下网页的网址吗?您可能根本不必使用 Selenium。
-
@PaulM。我正在尝试在底部stockx.com/adidas-yeezy-boost-700-bright-blue 上抓取过去价格的交互式图表
标签: python selenium web-scraping while-loop