【问题标题】:Regex to match each subdomain of a URL正则表达式匹配 URL 的每个子域
【发布时间】:2017-08-19 03:18:34
【问题描述】:

我正在尝试编写一个正则表达式,将 URL 的子域/域部分提取为单独的字符串。

我试过了:

/^[^:]+:\/\/([^\.\/]+)(\.[^\.\/]+)+(?:\/|$)/

它应该适用于这些 URL:

http;//www.mail.yahoo.co.uk/blah/blah

http;//test.test.again.mail.yahoo.com/blah/blah

我想把它分解成这样的部分:

["http://", "www", ".mail", ".yahoo", ".co", ".uk"]

["http://", "test", ".test", ".again", ".mail", ".yahoo", ".com"]

现在我只能将它们捕获为:

["http://", "www", ".uk"]

["http://", "test", ".com"]

有人知道我可以如何修复我的正则表达式吗?

【问题讨论】:

  • 你能发布你的代码吗?
  • 只是regex.exec(url)。

标签: javascript regex subdomain


【解决方案1】:

您可以使用/(http[s]?:\/\/|\w+(?=\.)|\.\w+)/g。 Test it online

【讨论】:

  • 你能解释一下为什么使用/^[^:]+:\/\/([^\.\/]+)(\.[^\.\/]+)+(?:\/|$)/来匹配http://test.test.again.mail.yahoo.com/blah/blah只有两个组。 ()+不会扩大,只有一组。如果我想拥有多个具有相同模式的组,如何编写?
  • 让我尝试一个简单的例子。考虑一个带有数字98997088567的手机号码,这三种情况:([9])表示匹配并返回所有出现的数字9,([9]+)表示匹配并返回所有连续出现的9s,([9])+表示所有出现连续的9s,但只返回最后一个匹配的组。
  • 谢谢!我明白。 :)
【解决方案2】:

你可以使用正则表达式

(^\w+:\/\/)([^.]+)

匹配第一部分,然后使用

\.\w+

匹配第二部分

检查代码sn-p

function getSubDomains(str){
    let result = str.match(/(^\w+:\/\/)([^.]+)/);
    result.splice(0, 1);
    result = result.concat(str.match(/\.\w+/g));
    console.log(result);
    return result;
}

getSubDomains('http://www.mail.yahoo.co.uk/blah/blah');
getSubDomains('http://test.test.again.mail.yahoo.com/blah/blah');

【讨论】:

    【解决方案3】:

    如何使用sticky flag y开始链接匹配

    var str = 'http://test.test.again.mail.yahoo.com/blah/blah';
    
    var res = str.match(/^[a-z]+:\/\/|\.?[^/.\s]+/yig);
    
    console.log(res);
    • ^[a-z]+:\/\/ 匹配协议:开始,一个或多个 a-z,后跟冒号和双斜杠。
    • |\.?[^/.\s]+ 或可选的点,后跟一个或多个字符 that are not 斜杠、点、空格。

    See Regex101 demo for more explanation

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2018-11-09
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多