【问题标题】:Jsoup - extract html from elementJsoup - 从元素中提取 html
【发布时间】:2016-05-17 19:40:16
【问题描述】:

我想使用 jsoup HTML 解析器库从 div 元素中提取 HTML 代码。

HTML 代码:

<div class="entry-content">
   <div class="entry-body">
      <p><strong>Text 1</strong></p>
      <p><strong> <a class="asset-img-link" href="http://example.com" style="display: inline;"><img alt="IMG_7519" class="asset  asset-image at-xid-6a00d8341c648253ef01b7c8114e72970b img-responsive" src="http://example.com" style="width: 500px;" title="IMG_7519" /></a><br /></strong></p>
      <p><em>Text 2</em> </p>
   </div>
</div>

提取部分:

String content = ... the content of the HTML from above
Document doc = Jsoup.parse(content);
Element el = doc.select("div.entry-body").first();

我希望结果 el.html() 是来自 div 选项卡 entry-body 的整个 HTML:

<p><strong>Text 1</strong></p>
  <p><strong> <a class="asset-img-link" href="http://example.com" style="display: inline;"><img alt="IMG_7519" class="asset  asset-image at-xid-6a00d8341c648253ef01b7c8114e72970b img-responsive" src="http://example.com" style="width: 500px;" title="IMG_7519" /></a><br /></strong></p>
  <p><em>Text 2</em> </p>

但我只得到第一个&lt;p&gt; 标签:

<p><strong>Text 1</strong></p>

【问题讨论】:

  • 这个问题对我来说是不可重现的。如果我完全按照你所说的那样做,我会得到所有内部 HTML 就好了。您使用的是哪个版本的 JSoup?我的检查是用 1.8.3 版完成的
  • 我也在使用 1.8.3 - 最后一个版本,但它不起作用...你得到了整个 div 内容的结果?

标签: android html-parsing jsoup


【解决方案1】:

试试这个:

Elements el = doc.select("div.entry-body");

而不是这个:

Element el = doc.select("div.entry-body").first();

然后:

for(Element e : el){
    e.html();
}

编辑

如果你这样做,也许你会得到结果: 我已经尝试做到这一点,它给出了正确的结果。 Elements el = doc.select("a.asset-img-link");

【讨论】:

  • 它不起作用.. el.size() 是 1 并且只打印 &lt;p&gt;&lt;strong&gt;Text 1&lt;/strong&gt;&lt;/p&gt;
【解决方案2】:

正如 OP 的 cmets 中所述,我不明白。这是我对问题的重现,它完全符合您的要求:

String html = ""
    +"<div class=\"entry-content\">"
    +"   <div class=\"entry-body\">"
    +"      <p><strong>Text 1</strong></p>"
    +"      <p><strong> <a class=\"asset-img-link\" href=\"http://example.com\" style=\"display: inline;\"><img alt=\"IMG_7519\" class=\"asset  asset-image at-xid-6a00d8341c648253ef01b7c8114e72970b img-responsive\" src=\"http://example.com\" style=\"width: 500px;\" title=\"IMG_7519\" /></a><br /></strong></p>"
    +"      <p><em>Text 2</em> </p>"
    +"   </div>"
    +"</div>"
    ;
Document doc = Jsoup.parse(html);
Element el = doc.select("div.entry-body").first();
System.out.println(el.html());

这会产生以下输出:

<p><strong>Text 1</strong></p> 
<p><strong> <a class="asset-img-link" href="http://example.com" style="display: inline;"><img alt="IMG_7519" class="asset  asset-image at-xid-6a00d8341c648253ef01b7c8114e72970b img-responsive" src="http://example.com" style="width: 500px;" title="IMG_7519"></a><br></strong></p> 
<p><em>Text 2</em> </p>

【讨论】:

  • 似乎这是我的愚蠢错误......我在 logcat 中有一些过滤器,这就是为什么我不能只看到第一行......对于这种混乱,我深表歉意,非常感谢。跨度>
【解决方案3】:

在你的情况下,你会使用

  doc.select("div[name=entry-body]") to select that specific <div>

据此cookbook

【讨论】:

  • 它不起作用 - 它返回 0 个元素。根据您所说的文档,应该使用el.class: elements with class, e.g. div.masthead,但这并不能让我了解整个div 内容..
  • EDIT 如果你这样做,也许你会得到结果:我已经尝试这样做并且它给出了正确的结果。 Elements el = doc.select("a.asset-img-link");
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-01-21
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-12-08
相关资源
最近更新 更多