【问题标题】:Downloading all hypelinked URLs from a Tumblr blog?从 Tumblr 博客下载所有超链接 URL?
【发布时间】:2017-04-09 23:31:23
【问题描述】:

从 Tumblr 博客下载所有图片/webms/mp4 的最佳方式是什么?

我希望从一些 Tumblr 博客下载所有帖子/图像/视频,它们在帖子正文中超链接 gfycat / webm 版本,Tumblripper / BulkImageDownloader / 其他 Tumblr 图像下载器无法捕获。我认为这是一个问题,因为它们在体内是超链接的,而不是实际上“在”Tumblr 上。

有人知道从 Tumblr 博客下载所有内容的好方法吗?我也尝试过 wget 和 httrack,但它们似乎不起作用。

我更喜欢使用带有 GUI 的程序来做我需要做的事情,而不是基于命令行的程序,因为我几乎不知道如何使用它们。我花了很长时间才弄明白 wget,而且我没有时间学习另一个下载 Tumblr 博客。

【问题讨论】:

    标签: download tumblr


    【解决方案1】:

    我知道您不喜欢命令行工具,但是我个人会使用 curl 将页面源写入文件:

    curl www.tumblr.com/something > outfile.html
    

    然后,您可以用您熟悉的任何语言解析文件。 这个答案有一些关于如何用 grep 做到这一点的极好的建议: https://unix.stackexchange.com/questions/181254/how-to-use-grep-and-cut-in-script-to-obtain-website-urls-from-an-html-file

    比如这个:

    $ curl -sL https://www.google.com | grep -Po '(?<=href=")[^"]*(?=")'
    /search?
    

    这给了你:

    https://www.google.co.in/imghp?hl=en&tab=wi
    https://maps.google.co.in/maps?hl=en&tab=wl
    https://play.google.com/?hl=en&tab=w8
    https://www.youtube.com/?gl=IN&tab=w1
    https://news.google.co.in/nwshp?hl=en&tab=wn
    ...
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-06-22
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多