【问题标题】:produce list of urls in python using regular expressions使用正则表达式在 python 中生成 url 列表
【发布时间】:2015-09-24 05:24:32
【问题描述】:

我的目标: 提取此 url 中的所有文字记录,并清理它们以供我特殊使用。
我需要递归地提取遵循某种模式的链接。我是一个新手,在编写完整的代码时遇到了麻烦。

以下是 URL 外观的一些示例:

http://tvmegasite.net/transcripts/oltl/main/1998transcripts.shtml
http://tvmegasite.net/transcripts/oltl/older/2004/oltl-trans-01-20-04.htm
http://tvmegasite.net/transcripts/amc/main/2003transcripts.shtml
http://tvmegasite.net/transcripts/amc/older/2002/amc-trans-01-08-02.shtml

所以都以http://tvmegasite.net/transcripts 开头,然后是show 缩写,然后是main 或更早,等等。

到目前为止我所尝试的: 使用 BeautifulSoup 从特定页面获取 url 很容易,但我还没有弄清楚如何递归地做到这一点。我正在考虑使用像 Scrapy 这样的刮板来获取从 tvmegasite.net/transcripts 开始的所有 url,然后使用 re 包来搜索与模式匹配的那些。我仍然不确定如何将其变成完整的代码。 据我猜测,这些可能是可以工作的正则表达式:

http://tvmegasite.net/transcripts\w+\/main/\d+\w+\.shtml
http://tvmegasite.net/transcripts\w+\/older/\d+/\w+\-\w+\-\d+\-\d+\.shtml

【问题讨论】:

  • tvmegasite.net/transcripts/passions/older/2001/…Timmy: It's like Timmy and Tabby are in a giant bath tub and someone pulled out the drain. What's going to happen to Timmy and Tabby?
  • 我知道,阅读这些内容很有趣,但我远离。 @maxymoo

标签: python regex parsing web-scraping scrapy


【解决方案1】:

如果你使用 Scrapy,你就不需要正则表达式——或者至少你可以将它们限制在最低限度。例如,使用LxmlLinkExtractor,您可以设置要遵循的 URL (allow) 和 XPath 分支 (restrict_xpaths)。

您可以在allow 限制中使用您的正则表达式(乍一看对我来说很好)——对于这个站点,您不需要对 XPath 进行限制。

【讨论】:

    猜你喜欢
    • 2012-03-16
    • 2015-08-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-11-24
    • 2019-04-01
    • 2019-05-16
    • 1970-01-01
    相关资源
    最近更新 更多