【问题标题】:Scraping Interactive Chart using Selenium Python More Efficiently?使用 Selenium Python 更有效地抓取交互式图表?
【发布时间】:2021-05-01 13:19:08
【问题描述】:

我正在尝试从网站上抓取交互式图表,这个脚本可以工作并从图表中收集文本,但速度很慢。为了使文本出现,光标必须悬停在图形上的某些位置。有人对如何提高效率有任何建议吗?

现在,对于每个偏移动作,它都会在前进到下一个动作之前遍历所有先前的动作。有谁知道如何绕过它? (例如,它从 0 到 5,然后不是从 5 到 10,而是再次回到 0,然后是 5,然后是 10)

# set the pace at which the cursor will move and the limit to which it will move to 
# (which should be the current date or the x-axis limit of the image)

# set the limit to be current date
limit = full_length
pace = 5
count = 0

while count <= limit:
    value = driver.find_element_by_class_name('highcharts-tooltip').text
    date_price = value.split("\n")
    date = date_price[0]
    price = date_price[1].split(": ")
    price = price[1]
    # take values at current point and add to dictionary
    dp = {'date': date,
         'price': price }
    archived_prices.append(dp)
    # move to the next date 
    action.move_by_offset(pace, 0).perform()
    # set up a counter to figure out when we will reach the limit
    count = count + pace
    print(count)

新代码:

# adding new column with complete url for api call
full_urls = []

for value in dataframe['urlKeys']:
    full = 'https://stockx.com/api/products/'+value+'?includes=market,360&currency=EUR&country=IT'
    full_urls.append(full)
    
dataframe['urlFull'] = full_urls


def get_shoe_info(url_list):
    
    for url in url_list:

        headers = {
            "accept-encoding": "gzip, deflate, br",
            "sec-fetch-mode": "cors",
            "sec=fetch-site": "same-origin",
            "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36",
            "x-requested-with": "XMLHttpRequest"
        }

        response = requests.get(url, headers=headers)
        response.raise_for_status()
        
        product = response.json()["Product"]

        for p in product:
            id_num = p["id"]
            brand = p["brand"]
            colorway = p["colorway"]
            release_date = p["releaseDate"]
            retail_price = p["retailPrice"]
            shoe_name = p["shoe"]
            volatility = p["market"]["volatility"]
            change_percentage = p["market"]["changePercentage"]
            gender = p["gender"]
            print(f'ID: {id_num}\n'
              f'Brand: {brand}\n'
              f'Colorway: {colorway}\n'
              f'Release Date: {release_date}\n'
              f'Retail Price: {retail_price}\n'
              f'Shoe Name: {shoe_name}\n'
              f'Volatility: {volatility}\n'
              f'Change Percentage: {change_percentage}\n'
              f'Gender: {gender}')
            #print(p)
        return 0
    

if __name__ == "__main__":
    import sys
    sys.exit(get_shoe_info(full_urls))

我仍然难以理解如何传递变量,所以我不确定我是否正确地做到了这一点。第一部分是我获取所有鞋子的 url 键并创建一个 url 列表以进行迭代。然后我试图将列表传递给 get_shoe_info 函数。我的错误弹出为“TypeError:字符串索引必须是整数”,我调查并看到当我尝试 print(p) 查看路径时,我只得到了关键部分的字符串。我不确定如何获得我想要的值。

我已将所有内容添加到 (my github),以防您需要查看其他内容。

【问题讨论】:

标签: python selenium web-scraping while-loop


【解决方案1】:

Selenium 对此太过分了。当我在浏览器中访问该页面时,我记录了我的网络流量,我看到我的浏览器向 REST API 发出了几个 XHR (XmlHttpRequest) HTTP GET 请求。其中一个具有端点api/products/.../chart,它返回包含您尝试抓取的所有图表信息的JSON。您需要做的就是模仿那个 HTTP GET 请求。不需要硒。我只是将所有请求标头和查询字符串参数复制到一些字典(headers 和params)中。在 API 抱怨并说我的请求格式错误之前,我将请求标头缩减到最低要求。

我更改了 accept-encoding 标头以仅接受 requests 库本身支持的那些编码格式,因为默认情况下 API 希望返回 Brotli 编码的 JSON。

我目前位于德国,这就是查询字符串参数显示"currency": "EUR" 和"country": "DE" 的原因。结果,响应将包含以欧元为单位的价格信息,但您应该能够更改这些键值对以满足您的需求(我猜“USD”和“US”应该可以工作)。此外,重要的是要注意响应包含一个“对”列表,一个用于图表上的每个 XY 坐标/数据点。 X 分量是时间(表示为以毫秒为单位的 unix 时间戳),Y 分量是以欧元为单位的价格(同样,由于我的查询字符串参数)。

下面,我定义了一个生成器get_price_history,它向 API 发出请求,然后生成所有数据点。每个 unix 时间戳首先转换为 datetime.datetime 对象:

def get_price_history():

    import requests
    from datetime import datetime

    url = "https://stockx.com/api/products/f27be8fd-2e05-4caa-a70e-fe787aa6283e/chart"

    params = {
        "start_date": "all",
        "end_date": "2021-05-01",
        "intervals": "100",
        "format": "highstock",
        "currency": "EUR",
        "country": "DE"
    }

    headers = {
        "accept-encoding": "gzip, deflate",
        "sec-fetch-mode": "cors",
        "sec-fetch-site": "same-origin",
        "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36",
        "x-requested-with": "XMLHttpRequest"
    }

    response = requests.get(url, params=params, headers=headers)
    response.raise_for_status()

    for timestamp, price in response.json()["series"][0]["data"]:
        date = datetime.utcfromtimestamp(int(timestamp) // 1000)
        yield date, price


def main():

    for date, price in get_price_history():
        print(f"[{date}]: €{price}")
    
    return 0


if __name__ == "__main__":
    import sys
    sys.exit(main())

输出:

[2021-04-01 13:36:27]: €358
[2021-04-01 20:48:41]: €358
[2021-04-02 04:00:55]: €358
[2021-04-02 11:13:10]: €401
[2021-04-02 18:25:24]: €401
[2021-04-03 01:37:38]: €401
[2021-04-03 08:49:53]: €410
[2021-04-03 16:02:07]: €442
[2021-04-03 23:14:22]: €422
[2021-04-04 06:26:36]: €422
[2021-04-04 13:38:50]: €422
[2021-04-04 20:51:05]: €414
[2021-04-05 04:03:19]: €414
[2021-04-05 11:15:33]: €462
[2021-04-05 18:27:48]: €500
...

如果您想了解更多关于我如何记录网络流量并找到 API URL、请求标头和查询字符串参数的信息,Take a look at this other answer 我已发布,我将在此处进行更深入的介绍。


编辑 - 感谢您分享您的代码。问题出在你的 get_shoe_info 函数中,你在哪里做:

product = response.json()["Product"]

for p in product:
    id_num = p["id"]
    brand = p["brand"]
    ...

问题在于product 是一个字典,所以当您执行for p in product: 时,您正在迭代该字典的键。键当然是字符串,因此每个p 将是一个字符串,p["id"] 将引发TypeError。实际上,您编写的内容等同于"id"["id"]、"brand"["brand"] 等。for 循环是不必要的——因此解决方案是简单地删除 for 循环。你所调用的p 实际上应该是product 字典:

product = response.json()["Product"]

id_num = product["id"]
brand = product["brand"]
colorway = product["colorway"]
release_date = product["releaseDate"]
retail_price = product["retailPrice"]
shoe_name = product["shoe"]
volatility = product["market"]["volatility"]
change_percentage = product["market"]["changePercentage"]
gender = product["gender"]
print(f'ID: {id_num}\n'
  f'Brand: {brand}\n'
  f'Colorway: {colorway}\n'
  f'Release Date: {release_date}\n'
  f'Retail Price: {retail_price}\n'
  f'Shoe Name: {shoe_name}\n'
  f'Volatility: {volatility}\n'
  f'Change Percentage: {change_percentage}\n'
  f'Gender: {gender}')

输出(对于单个 URL):

ID: cfa4ef16-7dec-4ef5-9318-3793c2c8546d
Brand: Louis Vuitton
Colorway: Red
Release Date: 2009-07-01
Retail Price: 870
Shoe Name: Louis Vuitton Don
Volatility: 0
Change Percentage: 0
Gender: men
>>> 

【讨论】:

  • 太棒了。也非常感谢您的解释。你为我节省了大量时间。
  • @GabbyV。很高兴我能帮上忙。如果您有任何后续问题,请告诉我。
  • 您好,我还有一个问题。所以我已经从网站上为每双鞋抓取了 url 键。我现在正在尝试调整您编写的内容以使用 url 列表进行迭代,而不仅仅是一个 url。我试图将列表传递给函数,但我收到一条错误消息,提示“IndexError:字符串索引超出范围”。这是我应该这样做的方式吗?如果您有时间,我很乐意与您分享代码。
  • @GabbyV。当然,我认为最简单的方法是使用新的最新代码编辑原始帖子。我很乐意看看。
  • 听起来不错!我已经更新了它,并附上了我的 github 和笔记本,以防你想查看更多
猜你喜欢
  • 1970-01-01
  • 2022-08-14
  • 2021-07-01
  • 2019-04-26
  • 2020-07-07
  • 2015-08-12
  • 1970-01-01
相关资源
最近更新 更多