【问题标题】:Unable to download NSE data using python无法使用 python 下载 NSE 数据
【发布时间】:2021-10-01 09:24:27
【问题描述】:

我正在尝试下载 NSE 股票数据(“https://www.nseindia.com/market-data/live-equity-market?symbol=NIFTY%2050”)

当我将以下 URL 粘贴到浏览器中时,文件将被下载。 https://www.nseindia.com/api/equity-stockIndices?csv=true&index=SECURITIES%20IN%20F%26O

当我尝试使用 python 的 requests 包下载相同的文件时,它会一直循环。

这是我用来下载文件的代码:

def download_url(url, save_path, chunk_size=1024):
    """
    r = requests.get(url, stream=True, verify=False)
    if r.status_code == 200:
        with open(save_path, 'wb') as fd:
            for chunk in r.iter_content(chunk_size=chunk_size):
                fd.write(chunk)
    """
    try:
        r = requests.get(url, stream=True, verify=False)
    except HTTPError as http_err:
        print(f'Adjacent Error occurred while accessing URL:{http_err}')
    except Exception as err:
        print(f'Adjacent error occurred while accessing URL:{err}')
    else:
        with open(save_path, 'wb') as fd:
            for chunk in r.iter_content(chunk_size=chunk_size):
                fd.write(chunk)

【问题讨论】:

  • 当我将 URL 粘贴到浏览器中时,我得到“找不到资源”。
  • 您不能仅使用 URL 下载。在浏览器中,有一些 cookie 被设置并使用 URL 发送。您必须提供所有正确的标题才能下载工作。

标签: python python-requests


【解决方案1】:

您需要加载页面的 cookie 才能执行请求。您可以先访问主页,然后尝试加载 API 请求:

from requests import Session

# Session to hold cookies
s = Session()
# Emulate browser
s.headers.update({"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/94.0.4606.61 Safari/537.36"})

# Get the cookies from the main page (will update automatically in headers)
s.get("https://www.nseindia.com/")

# Get the API data
data = s.get("https://www.nseindia.com/api/equity-stockIndices?csv=true&index=SECURITIES IN F%26O").text

# Write to file
with open("securities.csv", "w") as f:
    f.write(data)

【讨论】:

  • 非常感谢,一切正常。我正在学习网站报废,但找不到好的学习资料,您能帮帮我吗?
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-04-22
  • 2018-04-12
  • 2020-05-07
  • 2023-03-14
  • 2017-09-21
  • 2020-11-02
相关资源
最近更新 更多