【问题标题】:Regex - get URL protocol, host, path, but not filename - PCRE正则表达式 - 获取 URL 协议、主机、路径,但不是文件名 - PCRE
【发布时间】:2017-10-05 22:27:44
【问题描述】:

目标

替换主机和路径(位置),但保留文件名(它们不变)。

没有子域的 URL - 不起作用

这适用于具有至少一个子域的主机(域)(例如“www.somedomain.com”),但无法获取仅包含域 + TLD 的路径(例如“somedomain.com”)

(http[s]?:\/\/([^:\/\s]+)(\/\w+)*\/)+

在下面的 HTML sn-p

junk before tag <img src="https://somedomain.com/wp-content/uploads/2017/10/someimage.jpg" alt="" />Random text after

PCRE 引擎只会捕获:

https://somedomain.com/

URL 带有子域 - 有效

在下面的 HTML sn-p(域有一个子域)

junk before tag <img src="https://www.somedomain.com/wp-content/uploads/2017/10/someimage.jpg" alt="" />Random text after

PCRE 引擎捕获整个 URL(保存为文件):

https://www.somedomain.com/wp-content/uploads/2017/10/

问题

如何调整正则表达式以捕获具有子域的img src="" URL 的完整协议、域和路径(但不是文件名)以及那些没有子域的?

【问题讨论】:

  • 所以在第二个例子中你想返回www.somedomain.com?我不太清楚想要的输出到底是什么。
  • 在第一个例子中,我想要https://somedomain/wp-content/uploads/2017/10/,但我只得到https://somedomain/。第二个示例按预期工作。

标签: regex pcre url-scheme


【解决方案1】:
https?:\/\/(?:[^\/ ]*\/)*

演示here.

说明

http      //Should start with http
s?        // s is optional
:\/\/     // should follow up with ://
(?:       //START Non capturing group
[^\/ ]*   //Any character but a / or a space
\/        //Ends with /
)         //END Non capturing group
*         //Repeat non-capturing group

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2012-03-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多