【问题标题】:HtmlAgilityPack treats everything after < (less than sign) as attributesHtmlAgilityPack 将 <(小于号)之后的所有内容视为属性
【发布时间】:2016-06-15 19:37:39
【问题描述】:

我有一些通过 textarea 获得的输入,然后我将该输入转换为 html 文档,然后再将其解析为 PDF 文档。

当我的用户输入小于号 (

在这个字符数据块中,我可以尽可能多地使用双破折号(以及

如果我只是添加它会更好一点

htmlDocument.OptionOutputOptimizeAttributeValues = true;

这给了我:

在这个字符数据块中,我可以尽可能多地使用双破折号(以及

我已经尝试了 htmldocument 上的所有选项,但没有一个可以让我指定解析器不应该是严格的。另一方面,我也许可以忍受它去掉

void Main()
{
    var input = @"Within this Character Data block I can use double dashes as much as I want (along with <, &, ', and ') *and * % MyParamEntity; will be expanded to the text 'Has been expanded'...however, I can't use the CEND sequence(if I need to use it I must escape one of the brackets or the greater-than sign).";

    var htmlDoc = WrapContentInHtml(input);

    htmlDoc.DocumentNode.OuterHtml.ToString().Dump();
}

private HtmlDocument WrapContentInHtml(string content)
{
    var htmlBuilder = new StringBuilder();
    htmlBuilder.AppendLine("<!DOCTYPE html>");
    htmlBuilder.AppendLine("<html>");
    htmlBuilder.AppendLine("<head>");
    htmlBuilder.AppendLine("<title></title>");
    htmlBuilder.AppendLine("</head>");
    htmlBuilder.AppendLine("<body><div id='sagsfremstillingContainer'>");
    htmlBuilder.AppendLine(content); 
    htmlBuilder.AppendLine("</div></body></html>");

    var htmlDocument = new HtmlDocument();
    htmlDocument.OptionOutputOptimizeAttributeValues = true;
    var htmlDoc = htmlBuilder.ToString();

    htmlDocument.LoadHtml(htmlDoc);

    return htmlDocument;
}

有没有人知道我该如何解决这个问题。

我能找到的最接近的问题是: Losing the 'less than' sign in HtmlAgilityPack loadhtml

他实际上抱怨

编辑: 我正在使用 HtmlAgilityPack 1.4.9

【问题讨论】:

  • HTML-转义您的内容。如果它包含标记字符,它肯定会中断。​​
  • 我猜这不仅会破坏 HtmlAgilityPack。尝试在浏览器中查看它

标签: c# html-agility-pack


【解决方案1】:

您的内容明显错误。这与“严格性”无关,它实际上是关于您假装一段文本是有效 HTML 的事实。事实上,你得到的结果正是因为解析器不严格。

当你需要将纯文本插入HTML时,你需要先对其进行编码,以便所有的各种HTML控制字符都正确地转换为HTML——例如&amp;lt;必须改为&amp;lt;和@987654323 @到&amp;amp;。

解决此问题的一种方法是使用 DOM - 在目标 div 上使用 InnerText,而不是将字符串拼接在一起并假装它们是 HTML。另一种是使用一些显式的编码方法——例如HttpUtility.HtmlEncode。

【讨论】:

  • 你是对的。我的问题是我已经尝试对 html 进行编码,而我创建 PDF 的引擎没有对它们进行解码,但我当然可以自己这样做,然后它就可以工作了。在另一个节点上,我无法使用 InnerText,因为它是只读的,并且从上面的输入创建新节点不起作用。
【解决方案2】:

您可以使用System.Net.WebUtility.HtmlEncode,即使没有引用System.Web.dll,它也有HttpServerUtility.HtmlEncode

var input = @"Within this Character Data block I can use double dashes as much as I want (along with <, &, ', and ') *and * % MyParamEntity; will be expanded to the text 'Has been expanded'...however, I can't use the CEND sequence(if I need to use it I must escape one of the brackets or the greater-than sign).";
var htmlDocument = new HtmlDocument();
htmlDocument.LoadHtml(System.Net.WebUtility.HtmlEncode(input));
Debug.Assert(!htmlDocument.ParseErrors.Any());

结果:

Within this Character Data block I can use double dashes as much as I want (along with &lt;, &amp;, &#39;, and &#39;) *and * % MyParamEntity; will be expanded to the text &#39;Has been expanded&#39;...however, I can&#39;t use the CEND sequence(if I need to use it I must escape one of the brackets or the greater-than sign).

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2022-01-22
    • 1970-01-01
    • 1970-01-01
    • 2019-07-25
    • 2012-06-10
    • 2010-10-15
    • 1970-01-01
    • 2012-05-22
    相关资源
    最近更新 更多