【问题标题】:web harvest - scraping an url网络收获 - 抓取一个网址
【发布时间】:2013-03-15 00:09:11
【问题描述】:

我正在使用网络收获。但是,我想从 URL 中抓取数据:

http://derstandard.at/anzeiger/immoweb/Suchergebnis.aspx?Regionen=9&Bezirke=&Arten=&AngebotTyp=&timestamp=1363305908912

我的代码是:

<?xml version="1.0" encoding="UTF-8"?>

<config>
    <var-def name="google">
    <html-to-xml>
    <http url="http://derstandard.at/anzeiger/immoweb/Suchergebnis.aspx?Regionen=9&Bezirke=&Arten=&AngebotTyp=&timestamp=1363305908912"></http>
    </html-to-xml>
    </var-def>
</config>

但是我得到:

对实体 Bezirke 的引用必须以 ';' 结尾

我不明白网络收获是什么意思,带有';'?

【问题讨论】:

  • 我不确定你将如何收获网络,但我会推荐你​​使用 Jsoup。这真的很简单也很有用。

标签: java eclipse web web-scraping webharvest


【解决方案1】:

我不太了解网络收获,但他们的例子是这样的:

<xpath expression="//a[@shape='rect']/@href">
    <html-to-xml>
        <http url="http://www.somesite.com/"/>
    </html-to-xml>
</xpath>

<http url =".." />

而你的代码有

<http url = ".."></http> 

也许这是你的问题?不需要结束标签

【讨论】:

    【解决方案2】:

    您应该在您的 url 中编码与符号,即。将每个&amp;amp; 更改为&amp;amp;。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2015-05-25
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多