【问题标题】:Matching RegEx for Url Encoded links C#为 URL 编码链接 C# 匹配正则表达式
【发布时间】:2017-08-03 15:29:14
【问题描述】:

我有一个包含一些链接的 XML 文件

<SupportingDocs>
<LinkedFile>http://llcorp/ll/lljomet.dll/open/864606</LinkedFile>
<LinkedFile>http://llcorp/ll/lljomet.dll/open/1860632</LinkedFile>
<LinkedFile>%20http%3A%2F%2Fllenglish%2Fll%2Fll.exe%2Fopen%2F927515</LinkedFile>
<LinkedFile>%20http%3A%2F%2Fllenglish%2Fll%2Fll.exe%2Fopen%2F973783</LinkedFile>
</SupportingDocs>

我正在使用正则表达式 "\]+>(?:https?://|www.)[^\]+\[^\]+>"并使用 c# var matches = MyParser.Matches(FormXml); 但它匹配前两个链接,但不匹配编码的链接。

我们如何使用 RegEx 匹配 URL 编码的链接?

【问题讨论】:

  • 您在 https 之后匹配两个斜杠。这些存在于前两个中,但没有出现在第二个中。可能还有其他问题,但这是我第一次看到。

标签: c# regex pattern-matching urlencode


【解决方案1】:

这是一个可能有用的 sn-p。我真的怀疑你是否使用了最好的方法,所以我做了一些假设(也许你只是没有提供足够的细节)。

我将 xml 解析为 XmlDocument 以在代码中使用它。相关标签(“LinkedFile”)被拉出。每个标签都被解析为Uri。如果失败,它会被转义并再次尝试解析。最后将是一个包含正确解析的 url 的字符串列表。如果你真的需要,你可以在这个集合上使用你的正则表达式。

// this is for the interactive console
#r "System.Xml.Linq"
using System.Xml;
using System.Xml.Linq;

// sample data, as provided in the post.
string rawXml = "<SupportingDocs><LinkedFile>http://llcorp/ll/lljomet.dll/open/864606</LinkedFile><LinkedFile>http://llcorp/ll/lljomet.dll/open/1860632</LinkedFile><LinkedFile>%20http%3A%2F%2Fllenglish%2Fll%2Fll.exe%2Fopen%2F927515</LinkedFile><LinkedFile>%20http%3A%2F%2Fllenglish%2Fll%2Fll.exe%2Fopen%2F973783</LinkedFile></SupportingDocs>";
var xdoc = new XmlDocument();
xdoc.LoadXml(rawXml)

// will store urls that parse correctly
var foundUrls = new List<String>();

// temp object used to parse urls
Uri uriResult;

foreach (XmlElement node in xdoc.GetElementsByTagName("LinkedFile"))
{
    var text = node.InnerText;

    // first parse attempt
    var result = Uri.TryCreate(text, UriKind.Absolute, out uriResult);

    // any valid Uri will parse here, so limit to http and https protocols
    // see https://stackoverflow.com/a/7581824/1462295
    if (result && (uriResult.Scheme == Uri.UriSchemeHttp || uriResult.Scheme == Uri.UriSchemeHttps))
    {
        foundUrls.Add(uriResult.ToString());
    }
    else
    {
        // The above didn't parse, so check if this is an encoded string.
        // There might be leading/trailing whitespace, so fix that too
        result = Uri.TryCreate(Uri.UnescapeDataString(text).Trim(), UriKind.Absolute, out uriResult);

        // see comments above
        if (result && (uriResult.Scheme == Uri.UriSchemeHttp || uriResult.Scheme == Uri.UriSchemeHttps))
        {
            foundUrls.Add(uriResult.ToString());
        }
    }
}

// interactive output:
> foundUrls
List<string>(4) { "http://llcorp/ll/lljomet.dll/open/864606", "http://llcorp/ll/lljomet.dll/open/1860632", "http://llenglish/ll/ll.exe/open/927515", "http://llenglish/ll/ll.exe/open/973783" }

【讨论】:

  • xml 文件在不同的部分包含多种类型的 url。实际上,代码采用所有匹配类型的 url,然后处理每种类型。但是你的回答给了我一些思考的选择。谢谢
猜你喜欢
  • 1970-01-01
  • 2012-08-05
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多