【问题标题】:Finding certain text and filtering the rest out查找特定文本并过滤掉其余文本
【发布时间】:2012-10-17 20:07:57
【问题描述】:

假设我有这个字符串(巨大的),我想过滤掉除我要查找的内容之外的所有内容。这是我想要的一个例子:

<strong>You</strong></font> <font size="3" color="#05ABF8">
<strong>Shook</strong></font> Me All <font size="3" color="#05ABF8">
<strong>Night</strong></font> <font size="3" color="#05ABF8">
<strong>Long</strong></font> mp3</a></div>

如您所见,所有这些之间都有文字。我想得到“You Shook Me All Night Long”,然后把剩下的拿出来。我将如何完成这项工作?

【问题讨论】:

标签: c# regex http


【解决方案1】:

假设您在您发布的 xml/html 结尾处有结束 &lt;/a&gt;&lt;/div&gt; 的有效开始标签。

string value = XElement.Parse(string.Format("<root>{0}</root>", yourstring)).Value;

或者剥离Html的方法:

public static string StripHTML(this string HTMLText)
{
    var reg = new Regex("<[^>]+>", RegexOptions.IgnoreCase);
    return reg.Replace(HTMLText, "").Replace("&nbsp;", " ");
}

【讨论】:

  • 你不应该使用 LINQ2XML 来解析 html
【解决方案2】:

您可以使用以下正则表达式:&gt;([\s|\w]+)&lt;

var input = @"
<strong>You</strong></font> <font size='3' color='#05ABF8'>
<strong>Shook</strong></font> Me All <font size='3' color='#05ABF8'>
<strong>Night</strong></font> <font size='3' color='#05ABF8'>
<strong>Long</strong></font> mp3</a></div>";

var regex = new Regex(@">(?<match>[\s|\w]+)<");

var matches = regex.Matches(input).Cast<Match>()
   // Get only the values from the group 'match'
   // So, we ignore '<' and '>' characters
   .Select(p => p.Groups["match"].Value);

// Concatenate the captures to one string
var result = string.Join(string.Empty, matches)
    // Remove unnecessary carriage return characters if needed
    .Replace("\r\n", string.Empty);

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2012-12-10
    • 2019-01-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多