【发布时间】:2016-04-05 14:18:40
【问题描述】:
所以问题的基本前提是我们有一个文本文件,其中包含可能是或可能不是 Web 服务的数据列表。从文本文件中存在的 Web 服务列表中,我想解析每个 Web 服务可用的 Web 方法并将这些数据发布到 Excel 工作表。
我会给你一个测试数据是什么样子的例子:
<Resource Name="APP1">
<Uri UriType="PAGE" ResourceUri="http://exampleurl/default.aspx" />
</Resource>
<Resource Name="App2">
<Uri UriType="PAGE" ResourceUri="http://exampleurl2/example.aspx" />
</Resource>
<Resource Name="App3">
<Uri UriType="PAGE" ResourceUri="http://exampleurl3/exampleapp.asmx" />
</Resource>
基本上,最后一行是我想使用的那一行。另一个可用行的例子是
<Resource Name="Example" WSDL="http://example.wsdl">
<Uri UriType="ASMX" ResourceUri="http://example.asmx" />
</Resource>
所以,我实际上是在寻找 .asmx 和 .wsdl 文件。我考虑这个问题的方法是标准化我的输入,只为每个 Web 服务查找 WSDL,因此对于具有 .asmx 的 URL,我将添加 ?wsdl。
现在,下面是我实施的解决方案。由于源文件中有数千个 Web 服务,并且可能有 n 个 Web 方法,因此我看到执行时间长达 1-2 小时。我想知道是否可以进一步改进此解决方案以加快运行时间。
using System;
using System.Collections.Generic;
using System.Linq;
using System.Text;
using System.IO;
using System.Text.RegularExpressions;
using System.Xml;
using System.Net;
using System.Data;
using ClosedXML.Excel;
namespace ParseWebservices
{
class Program
{
static void Main(string[] args)
{
var lines = File.ReadAllText(@"PATH\SourceFIle.xml");
int count = 0;
string text = "";
DataTable Webservices= new DataTable();
Webservices.Columns.Add("Wsdl URL");
Webservices.Columns.Add("Webservice Name");
Webservices.Columns.Add("WebMethod");
Regex r = new Regex("(?<=ResourceUri=\")(.*)(.asmx)(?=\")", RegexOptions.IgnoreCase);
Match m = r.Match(lines.ToString());
while (m.Success)
{
try
{
string[] test = m.ToString().Split('/');
string webservicename = test[test.Length - 1].Replace(".asmx", "");
string wsdlurl="";
var webClient = new WebClient();
string readHtml="";
try
{
readHtml = webClient.DownloadString(wsdlurl);
}
catch (Exception excxx)
{
wsdlurl = m.ToString().Replace(".asmx", ".wsdl");
readHtml = webClient.DownloadString(wsdlurl);
}
int count2 = 0;
string text2 = "";
Regex r2 = new Regex(@"(?<=s:element name\=\"")(.*)(?=Response"")", RegexOptions.IgnoreCase);
Match m2 = r2.Match(readHtml);
while (m2.Success)
{
DataRow dr = Webservices.NewRow();
dr[0] = wsdlurl;
dr[1] = webservicename;
dr[2] = m2.ToString();
Console.WriteLine(wsdlurl + "\n" + webservicename + "\n" + m2.ToString());
Webservices.Rows.Add(dr);
count2++;
m2 = m2.NextMatch();
}
count++;
m = m.NextMatch();
}
catch (Exception ex)
{
m = m.NextMatch();
}
}
XLWorkbook wb = new XLWorkbook();
wb.Worksheets.Add(Webservices, "Example");
wb.SaveAs(@"PATH\example.xlsx");
}
}
}
我不喜欢这个解决方案的一点是它依赖于异常。因为正则表达式匹配.asmx 字符串,我意识到它将无法找到.wsdl 的字符串。但我也注意到,在包含.wsdl 的源文本中,.asmx 前缀完全相同。所以我为这些测试用例添加了错误处理,但绝对不理想。
无论如何,如果有任何关于如何改进和使其更快(更好!)的建议,我将不胜感激。
【问题讨论】:
-
该文档似乎是完全有效的 XML,您是否尝试过使用 XDocument 或 XmlDocument 来解析数据?它比使用正则表达式解析一个非常大的文件要快得多。
-
您可能在这里采取了错误的方法。如果您的输入文件是 XML,您应该查看 XML 解析而不是正则表达式。
-
另外,不是在另一个线程中连续检查所有 url 的创建,而是使用这些 url 并在 pararlell 上进行这些测试的队列
-
我尝试将其加载到 XMLdocument 中,但我看到的行为是节点未正确提供名称(我看到附加到标签的值为 null)。但是,由于这是我第一次使用 lib,所以我可能没有找对地方。
-
另外,您能否提供有关为什么 XML 库比正则表达式更快的见解?我只是想明白其中的道理。不过,平行的东西绝对是有道理的,我没有想到这一点(也从未尝试过,但这似乎是一个好机会)。
标签: c# regex string web-services text