【问题标题】:PHP script supposed to take 6 hours but stops after 30 minutesPHP 脚本需要 6 小时,但在 30 分钟后停止
【发布时间】:2012-01-01 07:24:52
【问题描述】:

我已经制作了一个基本的网络爬虫来从网站上抓取信息,我估计大约需要 6 个小时(将页面数乘以抓取信息所需的时间),但大约需要 30-40 分钟循环通过我的函数,它停止工作,我只有我想要的信息的一小部分。当它运行时,页面看起来像是正在加载,它会在屏幕上输出它的位置,但是当它停止时,页面停止加载并且输入停止显示。

无论如何我可以保持页面加载,这样我就不必每 30 分钟重新启动一次?

编辑:这是我的代码

function scrape_ingredients($recipe_url, $recipe_title, $recipe_number, $this_count) {
    $page   = file_get_contents($recipe_url);

    $edited = str_replace("<h2 class=\"ingredients\">", "<h2 class=\"ingredients\"><h2>", $page);

    $split  = explode("<h2 class=\"ingredients\">", $edited);
    preg_match("/<div[^>]*class=\"module-content\">(.*?)<\\/div>/si", $split[1], $ingredients);

    $ingred = str_replace("<ul>", "", $ingredients[1]);
    $ingred = str_replace("</ul>", "", $ingred);
    $ingred = str_replace("<li>", "", $ingred);
    $ingred = str_replace("</li>", ", ", $ingred);

    echo $ingred;
    mysql_query("INSERT INTO food_tags (title, link, ingredients) VALUES ('$recipe_title', '$recipe_url', '$ingred')");

    echo "<br><br>Recipes indexed: $recipe_number<hr><br><br>";

}

$get_urls   = mysql_query("SELECT * FROM food_recipes WHERE id>3091");
while($row  = mysql_fetch_array($get_urls)) {
    $count++;
    $thiscount++;
    scrape_ingredients($row['link'], $row['title'], $count, $thiscount);

    sleep(1);
}

【问题讨论】:

    标签: php web-crawler


    【解决方案1】:

    你的 php.ini 的 set_time_limit 选项值是多少?必须设置为 0 才能使脚本能够无限工作

    【讨论】:

    • 我将它设置为零并修复它!谢谢你:)
    【解决方案2】:

    尝试添加

    set_time_limit(0);
    

    在脚本的顶部。

    【讨论】:

      猜你喜欢
      • 2021-07-25
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-08-11
      • 1970-01-01
      • 1970-01-01
      • 2022-10-20
      • 2014-04-28
      相关资源
      最近更新 更多