【问题标题】:Improving a parsing solution改进解析解决方案
【发布时间】:2016-04-05 14:18:40
【问题描述】:

所以问题的基本前提是我们有一个文本文件,其中包含可能是或可能不是 Web 服务的数据列表。从文本文件中存在的 Web 服务列表中,我想解析每个 Web 服务可用的 Web 方法并将这些数据发布到 Excel 工作表。

我会给你一个测试数据是什么样子的例子:

<Resource Name="APP1">
    <Uri UriType="PAGE" ResourceUri="http://exampleurl/default.aspx" />
</Resource>
<Resource Name="App2">
    <Uri UriType="PAGE" ResourceUri="http://exampleurl2/example.aspx" />
</Resource>
<Resource Name="App3">
    <Uri UriType="PAGE" ResourceUri="http://exampleurl3/exampleapp.asmx" />
</Resource>

基本上,最后一行是我想使用的那一行。另一个可用行的例子是

<Resource Name="Example" WSDL="http://example.wsdl">
    <Uri UriType="ASMX" ResourceUri="http://example.asmx" />
</Resource>

所以,我实际上是在寻找 .asmx 和 .wsdl 文件。我考虑这个问题的方法是标准化我的输入,只为每个 Web 服务查找 WSDL,因此对于具有 .asmx 的 URL,我将添加 ?wsdl。

现在,下面是我实施的解决方案。由于源文件中有数千个 Web 服务,并且可能有 n 个 Web 方法,因此我看到执行时间长达 1-2 小时。我想知道是否可以进一步改进此解决方案以加快运行时间。

using System;
using System.Collections.Generic;
using System.Linq;
using System.Text;
using System.IO;
using System.Text.RegularExpressions;
using System.Xml;
using System.Net;
using System.Data;
using ClosedXML.Excel;

namespace ParseWebservices
{
    class Program
    {
        static void Main(string[] args)
        {

            var lines = File.ReadAllText(@"PATH\SourceFIle.xml");
            int count = 0;
            string text = "";
            DataTable Webservices= new DataTable();
            Webservices.Columns.Add("Wsdl URL");
            Webservices.Columns.Add("Webservice Name");
            Webservices.Columns.Add("WebMethod");

            Regex r = new Regex("(?<=ResourceUri=\")(.*)(.asmx)(?=\")", RegexOptions.IgnoreCase);
            Match m = r.Match(lines.ToString());
            while (m.Success)
            {


                try
                {

                    string[] test = m.ToString().Split('/');
                    string webservicename = test[test.Length - 1].Replace(".asmx", "");
                    string wsdlurl="";

                    var webClient = new WebClient();
                    string readHtml="";
                    try
                    {
                        readHtml = webClient.DownloadString(wsdlurl);
                    }
                    catch (Exception excxx)
                    {
                        wsdlurl = m.ToString().Replace(".asmx", ".wsdl");
                        readHtml = webClient.DownloadString(wsdlurl);
                    }

                    int count2 = 0;
                    string text2 = "";
                    Regex r2 = new Regex(@"(?<=s:element name\=\"")(.*)(?=Response"")", RegexOptions.IgnoreCase);
                    Match m2 = r2.Match(readHtml);
                    while (m2.Success)
                    {
                        DataRow dr = Webservices.NewRow();

                        dr[0] = wsdlurl;
                        dr[1] = webservicename;
                        dr[2] = m2.ToString();
                        Console.WriteLine(wsdlurl + "\n" + webservicename + "\n" + m2.ToString());
                        Webservices.Rows.Add(dr);
                        count2++;
                        m2 = m2.NextMatch();
                    }
                    count++;
                    m = m.NextMatch();
                }
                catch (Exception ex)
                {
                    m = m.NextMatch();
                }
            }

            XLWorkbook wb = new XLWorkbook();
            wb.Worksheets.Add(Webservices, "Example");
            wb.SaveAs(@"PATH\example.xlsx");
        }
    }
}

我不喜欢这个解决方案的一点是它依赖于异常。因为正则表达式匹配.asmx 字符串,我意识到它将无法找到.wsdl 的字符串。但我也注意到,在包含.wsdl 的源文本中,.asmx 前缀完全相同。所以我为这些测试用例添加了错误处理,但绝对不理想。

无论如何,如果有任何关于如何改进和使其更快(更好!)的建议,我将不胜感激。

【问题讨论】:

  • 该文档似乎是完全有效的 XML,您是否尝试过使用 XDocument 或 XmlDocument 来解析数据?它比使用正则表达式解析一个非常大的文件要快得多。
  • 您可能在这里采取了错误的方法。如果您的输入文件是 XML,您应该查看 XML 解析而不是正则表达式。
  • 另外,不是在另一个线程中连续检查所有 url 的创建,而是使用这些 url 并在 pararlell 上进行这些测试的队列
  • 我尝试将其加载到 XMLdocument 中,但我看到的行为是节点未正确提供名称(我看到附加到标签的值为 null)。但是,由于这是我第一次使用 lib,所以我可能没有找对地方。
  • 另外,您能否提供有关为什么 XML 库比正则表达式更快的见解?我只是想明白其中的道理。不过,平行的东西绝对是有道理的,我没有想到这一点(也从未尝试过,但这似乎是一个好机会)。

标签: c# regex string web-services text


【解决方案1】:

这很慢,因为这一切都是在一个线程上完成的! (无论是 xml 还是 regex 都与速度关系不大:真正拖慢速度的是所有内联 Web 请求)

没有你的源文件很难做一个有效的例子,所以我写了一个辅助扩展来异步加载一个 URL 列表 - 你显然需要在它周围填充你的代码。

using System;
using System.Collections.Generic;
using System.Linq;
using System.Text;
using System.IO;
using System.Text.RegularExpressions;
using System.Xml;
using System.Net;
using System.Data;
using System.Collections.Concurrent;

using System.Threading.Tasks;

namespace ParseWebservices
{
    static class UrlLoaderExtension
    {
    public static async Task<ConcurrentDictionary<string, string>> LoadUrls(this IEnumerable<string> urls)
    {
        var result = new ConcurrentDictionary<string,string>();                
        Task[] tasks = urls.Select(url => {
            return Task.Run(async () =>
            {
                using (WebClient wc = new WebClient())
                {
                    // Console.WriteLine("Thread: " + System.Threading.Thread.CurrentThread.ManagedThreadId);
                    try
                    {
                        var r = await wc.DownloadStringTaskAsync(url);
                        result[url] = r;
                    }
                    catch (Exception err)
                    {
                        result[url] = err.Message;
                    }
                }
            });
        }).ToArray();                
        await Task.WhenAll(tasks);
        return result;
    }
    }

    class Program
    {
        static void Main(string[] args)
        {
            var requests = new ConcurrentDictionary<string,string>();

            // load desired urls into the structure
            requests["http://www.microsoft.com"] = null;
            requests["http://www.google.com"] = null;
            requests["http://www.google.com/asdfdsaf"] = null;

            try
            {
                Task.Run(async () =>
                {
                    requests = await requests.Keys.LoadUrls();
                }).GetAwaiter().GetResult();
            }
            catch (Exception ex)
            {
                Console.WriteLine("Error: " + ex.Message);
                Console.ReadLine();
                return;
            }

            Console.WriteLine("Finished loading data concurrently");
            Console.ReadLine();

            // this part is synchronous (it's not waiting for IO)
            foreach(var url in requests.Keys)
            {
                var response = requests[url];
                Console.WriteLine(response); // 
                Console.WriteLine("Response from " + url);
                Console.ReadLine();
            }



            Console.Write("DONE");
            Console.ReadLine();
        }
    }
}

我建议您将您的网址放入此演示中,以了解加载数据的速度有多快:它告诉您加载完成的时间点是它收集了所有响应。

然后,在您确定了(非常!)速度有多快之后,您就会有动力在它周围填充您的其他逻辑:)

希望对你有帮助!

【讨论】:

  • @user2044754 你有机会尝试一下吗? :)
【解决方案2】:

就像 cmets 建议的那样,如果您的示例是有效的 XML,我怀疑 XML 解析解决方案可能比 Regex 更易于使用且速度更快。您可以尝试以下方法:

var files = XElement.Parse(xmlString)
    .Descendants("Resource").SelectMany(resource =>
    {
        XAttribute wsdlAttribute = resource.Attribute("WSDL");
        XAttribute resourceUriAttribute = resource.Element("Uri").Attribute("ResourceUri");
        if (wsdlAttribute != null)
            return new[] { wsdlAttribute.Value, resourceUriAttribute.Value };
        else
            return new[] { resourceUriAttribute.Value };
    }).Select(uri => Path.GetFileName(uri));

返回:

  • default.aspx
  • example.aspx
  • exampleapp.asmx
  • example.wsdl
  • example.asmx

使用我从您的帖子中创建的测试 xml 字符串:

        string xmlString = 
@"<Root>
    <Resource Name=""APP1"">
        <Uri UriType=""PAGE"" ResourceUri=""http://exampleurl/default.aspx"" />
    </Resource>
    <Resource Name=""App2"">
        <Uri UriType=""PAGE"" ResourceUri=""http://exampleurl2/example.aspx"" />
    </Resource>
    <Resource Name=""App3"">
        <Uri UriType=""PAGE"" ResourceUri=""http://exampleurl3/exampleapp.asmx"" />
    </Resource>
    <Resource Name=""Example"" WSDL=""http://example.wsdl"">
        <Uri UriType=""ASMX"" ResourceUri=""http://example.asmx"" />
    </Resource>
</Root>";

我不能保证它会比您的解决方案更快,但我们非常欢迎您测试它!如果你有多个文件要处理,你也可以线程化。

【讨论】:

  • 我要试一试,然后报告结果!
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2010-10-16
  • 1970-01-01
  • 2021-05-06
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多