【问题标题】:HtmlUnit to click on specific link from link's reference with same nameHtmlUnit 以单击具有相同名称的链接引用中的特定链接
【发布时间】:2013-08-27 06:48:56
【问题描述】:

我今天开始使用 HtmlUnit,所以我当时有点菜鸟。

我设法去 IMDB 搜索了 1996 年的电影《沉睡者》,我得到了一堆同名的结果:

Here are the results from that search

我想从列表中选择第一个“睡眠者”,这是正确的,但我不知道如何使用 HtmlUnit 获取该信息。我查看了代码并找到了链接,但我不知道如何提取它。

我想我可以使用一些正则表达式,但这会破坏使用 HtmlUnit 的目的。

这是我的代码(它有一些来自 HtmlUnit 的教程和在这里找到的一些代码):

public IMdB() {
    try {
        //final WebClient webClient = new WebClient();

        final WebClient webClient = new WebClient(BrowserVersion.INTERNET_EXPLORER_8, "10.255.10.34", 8080);

        //set proxy username and password 
        final DefaultCredentialsProvider credentialsProvider = (DefaultCredentialsProvider) webClient.getCredentialsProvider();
        credentialsProvider.addCredentials("xxxx", "xxxx");

        // Get the first page
        final HtmlPage page1 = webClient.getPage("http://www.imdb.com");

        // Get the form that we are dealing with and within that form, 
        // find the submit button and the field that we want to change.
        //final HtmlForm form = page1.getFormByName("navbar-form");
        HtmlForm form = page1.getFirstByXPath("//form[@id='navbar-form']");

        //
        HtmlButton button = form.getFirstByXPath("/html/body//form//button[@id='navbar-submit-button']");            
        HtmlTextInput textField = form.getFirstByXPath("/html/body//form//input[@id='navbar-query']");

        // Change the value of the text field
        textField.setValueAttribute("Sleepers");

        // Now submit the form by clicking the button and get back the second page.
        HtmlPage page2 = button.click();

       // form = page2.getElementByName("s");

        //page2 = page2.getFirstByXPath("/html/body//form//div//tr[@href]");

        System.out.println("content: " + page2.asText());

        webClient.closeAllWindows();
    } catch (IOException ex) {
        Logger.getLogger(IMdB.class.getName()).log(Level.SEVERE, null, ex);
    }

    System.out.println("END");
}

【问题讨论】:

  • 你能使用这个建议吗?
  • 不,因为这不是我想要的,但无论如何谢谢。我终于使用了一些正则表达式来提取一些特定的数据。

标签: java html regex htmlunit


【解决方案1】:

你应该这样做:

HtmlPage htmlPage = new WebClient().getPage("http://imdb.com/blah");
HtmlAnchor anchor = htmlPage.getFirstByXPath("//td[@class='primary_photo']//a")
System.out.println(anchor.getHrefAttribute());

【讨论】:

  • 谢谢。我会试试这个。我已经设法使用正则表达式提取了一些特定的数据,但我认为 HtmlUnit 有一些用于此类事情的工具。
  • 我将如何从中提取“睡眠者”部分:<td class="title"> <span class="wlb_wrapper" data-tconst="tt0117665" data-size="small" data-caller-name="search"></span> <a href="/title/tt0117665/">Sleepers</a> <span class="year_type">(1996)</span><br> <div class="user_rating">?我的程序运行良好。它通过正则表达式查找电影、演员、提名、评级等。
  • 正则表达式显然不是要走的路。检查此question 和第一个答案。我建议您使用 XPath。你会在 google 上找到很多教程。
  • 谢谢,我会调查一下,但我认为正则表达式(尽管不是最好的方法)适用于简单的事情。我已经让我的程序工作并使用正则表达式获取信息(直到我学会如何使用 xpath9)。
【解决方案2】:

我建议你宁愿使用IMDB api 然后做所有这些

IMDb 目前有两个公共 API,虽然没有记录,但非常快速和可靠(通过 AJAX 在他们自己的网站上使用)。

  1. 静态缓存的搜索建议 API:

  2. 更高级的搜索

【讨论】:

  • 我认为他的目的是学习使用 HtmlUnit,而不是从 IMDB 中提取数据。
  • @MostyMostacho 也许好吧,我推断不是这样。
  • 是的,就是这个主意,但还是谢谢你。很高兴知道它。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2014-11-04
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-12-26
相关资源
最近更新 更多