【问题标题】:regex for all urls in string not capturing urls with question mark字符串中所有 url 的正则表达式不捕获带问号的 url
【发布时间】:2018-09-12 20:40:23
【问题描述】:

我从这个 PHP 字符串开始。

$bodyString = '
    another 1 body
    reg http://www.regularurl.com/home
    secure https://facebook.com/anothergreat.
    a subdomain http://info.craig.org/
    dynamic; http://www.spring1.com/link.asp?id=100408
    www domain; at www.wideweb.com
    single no subdomain; simple.com';

需要将所有域、urls 转换为anchor(<a>) 元素。

preg_replace('#[-a-zA-Z0-9@:%_\+.~\#?&//=]{2,256}\.[a-z]{2,4}\b(\/[-a-zA-Z0-9@:%_\+.~\#?&//=]*)?#si', '<a href="$0">$0</a>', $bodyString)

$bodyString 结果:

'another 1 body
    reg <ahref="http://www.regularurl.com/home">http://www.regularurl.com/home</a>
    secure <a href="https://facebook.com/anothergreat.">https://facebook.com/anothergreat.</a>
    a subdomain <a href="http://info.craig.org/">http://info.craig.org/</a>
    dynamic; <a href="http://www.spring1.com/link.asp">http://www.spring1.com/link.asp</a>?id=100408
    www domain; at <a href="www.wideweb.com">www.wideweb.com</a>
    single no subdomain; <a href="simple.com">simple.com</a>';

结果:所有的url、域名都变成&lt;a&gt;除了http://www.spring1.com/link.asp?id=100408

正则表达式中缺少什么来完成这项工作?

【问题讨论】:

  • 可能是this will help。
  • 我认为如果您可以像问题中那样编写这么多的正则表达式,您可以自己找到解决方案。在从互联网复制粘贴代码之前,请务必从头开始尝试并发布您的研究。

标签: php regex url


【解决方案1】:
$bodyString = '
    another 1 body
    reg http://www.regularurl.com/home
    secure https://facebook.com/anothergreat.
    a subdomain http://info.craig.org/
    dynamic; http://www.spring1.com/link.asp?id=100408
    www domain; at www.wideweb.com
    single no subdomain; simple.com';

$regex = '@(http)?(s)?(://)?(([a-zA-Z])([-\w]+\.)+([^\s\.]+[^\s]*)+[^,.\s])@'; 
$converted_string = preg_replace($regex, '<a href="$0">$0</a>', $bodyString);
echo $converted_string;

Demo

正则表达式解释here

【讨论】:

    【解决方案2】:

    基于@WiktorStribiżew 的评论,您可以试试这个:

    [^\s]{2,256}\.[a-z]{2,4}\b(?:[?/][^\s]*)*
    

    Trial over here

    注意-虽然到目前为止已经有2个答案,但这似乎更简洁,使用[^\s]

    解释-

    [^\s]{2,256} 匹配 2 到 256 个字符,即 https://facebook 和 https://www.randomdomain 部分,
    \. 匹配之后的点,
    [a-z]{2,4} 是域扩展名,例如:com , in 等
    \b 是单词边界,
    (?:[?/][^\s]*)* 是一个非捕获组,它匹配斜杠 / 或问号 ? 以及更多的 url,所有其中可以重复0次或多次,表示URL的子页面。

    为了更好地理解正则表达式语法,您应该try this website: rexegg.com

    【讨论】:

      【解决方案3】:

      [-\w@:%+.\~#?&amp;/=]{2,256}\.[a-z]{2,4}\b[^\s]*

      [^\s]* 会将任何非空格字符添加到 url。当有空格时,它不是 URL 的一部分。简单易行。

      工作网址here

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2015-04-21
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多