【发布时间】:2013-03-16 11:14:33
【问题描述】:
HTML 结构如下所示:
<div class="Parent">
<div id="A">more tags and text</div>
<div id="B">more tags and text</div>
more tags
<p> and text </p>
</div>
我想只从父级和标签中提取文本,而不是 A 和 B 子级。 我努力了 /div[@class='Parent']//text()
它从所有后代节点中提取文本,因此 a 做了一个约束,例如 /div[@class='Parent']//text()[not(self::div)]
但它并没有改变任何事情。
感谢您的建议
【问题讨论】:
标签: xpath html-parsing