【发布时间】:2015-05-26 05:07:30
【问题描述】:
嗨,我正在尝试抓取 google 搜索结果,只是为了我自己的学习,同时也想看看我能否加快访问直接 URL 的速度(我知道他们的 API,但我只是想我现在试试这个)。
它工作正常,但似乎已经停止,它现在根本没有返回,我不确定我是否做了什么,但我可以说我在 for 循环中有这个,以允许 start 参数增加和我想知道这可能会导致问题。
Google 是否可以阻止 IP 抓取?
谢谢..
$url = "https://www.google.ie/search?q=adrian+de+cleir&start=1&ie=utf-8&oe=utf-8&rls=org.mozilla:en-US:official&client=firefox-a&channel=fflb&gws_rd=cr&ei=D730U7KgGfDT7AbNpoBY#channel=fflb&q=adrian+de+cleir&rls=org.mozilla:en-US:official";
$ch = curl_init();
$timeout = 5;
curl_setopt($ch, CURLOPT_SSL_VERIFYHOST, 0);
curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, 0);
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, $timeout);
$html = curl_exec($ch);
curl_close($ch);
# Create a DOM parser object
$dom = new DOMDocument();
# Parse the HTML from Google.
# The @ before the method call suppresses any warnings that
# loadHTML might throw because of invalid HTML in the page.
@$dom->loadHTML($html);
# Iterate over all the <a> tags
foreach($dom->getElementsByTagName('h3') as $link) {
$actual_link = $link->getElementsbyTagName('a');
foreach ($actual_link as $single_link) {
# Show the <a href>
echo '<pre>';
print_r($single_link->getAttribute('href'));
echo '</pre>';
}
}
【问题讨论】:
-
是的,如果您请求太频繁,他们当然能够并允许阻止您的 IP。您是否尝试在另一台机器上运行您的脚本?
-
还没有,但我会回家试试看,我也尝试过包含一个带'curl'的代理,但没有成功
-
只是为了扩展上述内容,我将
h3更改为div只是作为测试,它向我显示了一些数据(但显然搜索结果中没有任何内容 -
只需打印出整个 html 并粘贴到这里
-
是的,它是一个任何浏览器都能够理解的重定向,但是像你的脚本这样的 curl 太微不足道了,这就是谷歌如何防止 curl bots 的方式
标签: php curl google-search