【问题标题】:Wget complete webpage from list of urls从 url 列表中获取完整的网页
【发布时间】:2014-07-06 03:58:17
【问题描述】:

我正在寻找一些关于如何使用我的单个 url wget 脚本并从文本文件中实现 url 列表的提示。我不确定如何编写脚本 - 在循环中或以某种方式枚举它?这是我用来从一个页面收集所有内容的代码:

wget \
    --recursive \
    --no-clobber \
    --page-requisites \
    --html-extension \
    --convert-links \
    --restrict-file-names=windows \
    --domains example.com \
    --no-parent \
        http://www.example.com/folder1/folder/

它工作得非常好——我只是迷失了如何使用带有以下网址的list.txt:

http://www.example.com/folder1/folder/
http://www.example.com/sports1/events/
http://www.example.com/milfs21/delete/
...

我想这是相当简单的,但又一次永远不知道,谢谢。

【问题讨论】:

    标签: macos bash shell wget


    【解决方案1】:

    根据wget --help:

       -i file
       --input-file=file
           Read URLs from a local or external file.  If - is specified as
           file, URLs are read from the standard input.  (Use ./- to read from
           a file literally named -.)
    

    另一种方法是在从文件中读取列表时使用循环:

    readarray -t LIST < list.txt
    
    for URL in "${LIST[@]}"; do
        wget \
            --recursive \
            --no-clobber \
            --page-requisites \
            --html-extension \
            --convert-links \
            --restrict-file-names=windows \
            --domains example.com \
            --no-parent \
            "$URL"
    done
    

    同样可以使用while read 循环。

    【讨论】:

    • 哇,这很容易(对你来说)。我在 osx 上,所以我不得不使用不同的循环方式,因为 readarray 不存在。我使用了一些非常规的东西,也许是(similar to this answer),虽然它可以工作,所以谢谢你让我走上了正确的道路:) 内置选项-i file 也可以正常工作,消除了循环的需要,但它很棒了解两者。 另一个问题:我将如何指定数据的保存位置?
    • @ctfd -P 可能会有所帮助:-P, --directory-prefix=PREFIX save files to PREFIX/...。欢迎:)
    猜你喜欢
    • 2020-02-25
    • 2020-11-06
    • 2010-10-24
    • 2011-04-08
    • 2011-05-13
    • 1970-01-01
    • 2017-04-22
    • 2014-07-21
    • 2015-09-14
    相关资源
    最近更新 更多