【发布时间】:2012-02-11 18:18:00
【问题描述】:
我正在制作 Torrent PHP Crawler,但我遇到了问题,这是我的代码:
// ... the cURL codes (they're working) ...
// Contents of the Page
$contents = curl_exec($crawler->curl);
// Find the Title
$pattern = "/<title>(.*?)<\/title>/s";
preg_match($pattern, $contents, $titlematches);
echo "Title - ".$titlematches[1]."<br/>";
// Find the Category
$pattern = "/Тип<\/td><td(?>[^>]+)>((?>[^<]+))<\/td>/s";
preg_match($pattern, $contents, $categorymatches);
echo "Category - ".$categorymatches[1]."<br/>";
HTML 页面(“Тип”表示类别,“Филми”表示电影):
<title>The Matrix</title>
<!--Some Codes Here--!>
<tr><td>Тип</td><td valign="top" align=left>Филми</td></tr>
<!--Some Codes Here--!>
结果:
Title - The Matrix
Notice: Undefined offset: 1 in /var/www/spider.php on line 117
它显示的是标题而不是类别..这是为什么呢?
我尝试回显$categorymatches[0]、$categorymatches[2]、$categorymatches[3],但没有任何运气。
【问题讨论】:
-
这意味着
contents不会为categorymatches创建匹配项。此外,cmets 以-->关闭,而不是--!> -
$contents不包含正确的 HTML 数据。尝试在curl_exec()之后立即回应它,看看会出现什么。我使用您提供的 HTML 在本地进行了尝试,效果很好,完美匹配。
标签: php curl preg-match web-crawler