【问题标题】:Guzzle Http Client ambiguous responseGuzzle Http 客户端模棱两可的响应
【发布时间】:2015-01-15 17:37:32
【问题描述】:

这是我在不同文件中的代码

use Symfony\Component\DomCrawler\Crawler;
$guzzle = new GuzzleHttp\Client(['base_url' => 'http://pricematch.pk/mobile/samsung-galaxy-s6-80514-price-in-Pakistan']);
$response = $guzzle->get();
$crawler = new Crawler((string)$response->getBody());
echo $crawler->filter('.product-shop-wrapper .price')->text()."\r\n";

这个网址是硬编码的,这个网址成功地回显了过滤后的文本。当下面代码中每个循环中的相同 url/任何 url 来自变量时

$guzzle = new GuzzleHttp\Client(['base_url' => 'pricematch.pk/mobile-phone-prices-in-pakistan']);

$response = $guzzle->get();

$crawler = new Crawler((string)$response->getBody());
$crawler->filter('.product-name')->each(function ($node,$counter) {
    echo $counter." ".$node->text()."\r\n";
    $url=$node->filter('a')->extract(array('href'))[0]."\r\n";
    echo $url."\r\n";
    $url='http://pricematch.pk'.$url;
    echo $url;
    $guzzle = new GuzzleHttp\Client(['base_url' => $url]);
    $response = $guzzle->get();
    $crawler = new Crawler((string)$response->getBody());

爬虫抛出异常,说当前节点列表为空。 href 返回一个相对 url,我在上面的代码中附加了根 url。我已经打印了很多次生成的网址。即使 url 与 code#1 中的 url 相同,过滤器也会抛出异常。 我究竟做错了什么?
更新 2:我刚刚发现 code2 爬虫中的数据来自

pricematch.pk/mobile-phone-prices-in-pakistan


它应该来自哪里

$url
这是怎么回事?

【问题讨论】:

    标签: php symfony web-scraping guzzle


    【解决方案1】:

    我太傻了。在上面的代码中提取 URL 时,我连接了可能将 URL 编码为 __ 在 URL 末尾的换行符,这基本上改变了 URL,因此改变了响应。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2023-03-14
      • 2015-10-25
      • 2012-08-17
      • 1970-01-01
      • 2011-06-13
      • 2016-04-10
      • 2020-10-05
      • 1970-01-01
      相关资源
      最近更新 更多