【问题标题】:how to make Python3 web scraping program deal with local cookies?如何让 Python3 网页抓取程序处理本地 cookie?
【发布时间】:2018-10-12 03:51:35
【问题描述】:

我尝试编写一个可以自动下载文件的程序(带有php链接)。但是,我现在有两个问题

首先,我的目标网站需要注册才能首次访问。然后,每次我点击下载链接时,它都会自动下载我想要的文件。看起来像是搜索了一些保存在我电脑上的 cookie 以确定我是谁。如何让我的 python 程序处理我的本地 cookie?如果是倍数?

其次,谁能给我一个关于如何处理php下载链接文件的示例代码?我想以特定名称将所有这些文件保存在特定位置。我应该如何在 python3 中做到这一点?

【问题讨论】:

  • 添加你想要解决的代码..??

标签: php python web-scraping web-crawler session-cookies


【解决方案1】:

获取 cookie:

试试:

import urllib.request
cookier = urllib.request.HTTPCookieProcessor()
# create the cookie handler
opener = urllib.request.build_opener(cookier)
urllib.request.install_opener(opener)

HTTPCookieProcessor 将返回包含这些 cookie 的 cookielib.CookieJar 对象。你可以遍历它来找到你想要的cookie。

for c in cookier.cookiejar: 
    if c.domain == '.stackoverflow.com': 
        # do something

阅读链接中的内容:

试试:

url = 'YOUR_URL'
req = urllib.request.Request(url, headers=_headers) # where headers is the header setting you can find in your brwoser
f = urllib.request.urlopen(req)
contents = f.read().decode('utf-8')
# contents is the content inside your file
# You can add the code here to write contents to other file to save it

【讨论】:

  • 非常感谢。但我不太了解 cookie 部分。我应该如何使用这个饼干?还有,这背后的机制是什么?你能给我解释一下吗?
  • for c in cookier.cookiejar: if c.domain == '.yahoo.com': # do something 你可以这样做。
  • 一旦你找到你想要的cookie,你可以打电话c.value取回它并在某处使用
  • @yorkevin 你现在能拿到cookie吗?
猜你喜欢
  • 2017-11-20
  • 1970-01-01
  • 2021-12-11
  • 1970-01-01
  • 1970-01-01
  • 2013-06-15
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多