【问题标题】:scraping images from url using php使用 php 从 url 抓取图像
【发布时间】:2017-09-21 18:23:11
【问题描述】:

我正在尝试制作一个允许我从另一个链接抓取和保存图像的页面,所以这是我想在我的页面上添加的内容:

  1. 文本框(输入我想从中获取图像的网址)。
  2. 保存对话框指定保存图片的路径。

但我想在这里做的只是从该 url 和特定元素内部保存图像。

例如,在我的代码中,我说去 example.com 并从元素 class="images" 内部抓取所有图像。

注意:并非页面中的所有图像,仅来自元素内部 元素中有 3 张图片还是 50 张或 100 张我不在乎。

这是我使用 php 尝试和工作的内容

<?php
$html = file_get_contents('http://www.tgo-tv.net');
preg_match_all( '|<img.*?src=[\'"](.*?)[\'"].*?>|i',$html, $matches ); 
echo $matches[ 1 ][ 0 ];
?>

这会获取图像名称和路径,但我要做的是一个保存对话框,代码必须将图像直接保存到该路径而不是回显它

希望你能理解

编辑 2

没有保存对话框也没关系。我必须从代​​码中指定保存路径

【问题讨论】:

  • 你必须抓取网站,获取网址,然后看看这个答案:stackoverflow.com/questions/3938534/… 不想粗鲁,但我也不知道该怎么做,但是一个简单的谷歌搜索我找到了答案。对您下一个问题的小建议
  • @Nytrix 我真的需要它。如果您成功并在此处发布作为答案,我将不胜感激。我对php了解不多。是的,你可能会说先学习 php,我正在学习,但现在我急需它
  • 不。如果您需要超出您的知识范围,您可以聘请能够做到的人。 SO 不是免费的编码服务,紧急绝不是在这里发布内容的好理由......我给了你一个帖子,你的问题是我>回答了。
  • @Nytrix 你不会永远为我做这件事,这只是我要求的简单帮助
  • 我已经给你解答了如何从 URL 保存文件,你还需要什么?

标签: php


【解决方案1】:

如果你想要通用的东西,你可以使用:

<?php
    $the_site = "http://somesite.com";
    $the_tag = "div"; #
    $the_class = "images";

    $html = file_get_contents($the_site);
    libxml_use_internal_errors(true);
    $dom = new DOMDocument();
    $dom->loadHTML($html);
    $xpath = new DOMXPath($dom);

    foreach ($xpath->query('//'.$the_tag.'[contains(@class,"'.$the_class.'")]/img') as $item) {

        $img_src =  $item->getAttribute('src');
        print $img_src."\n";

    }

用法:

更改site、tag,可以是div、span、a等,也可以更改class名称。

例如,将值更改为:

$the_site = "https://stackoverflow.com/questions/23674744/what-is-the-equivalent-of-python-any-and-all-functions-in-javascript";
$the_tag = "div"; #
$the_class = "gravatar-wrapper-32";

输出:

https://www.gravatar.com/avatar/67d8ca039ee1ffd5c6db0d29aeb4b168?s=32&d=identicon&r=PG
https://www.gravatar.com/avatar/24da669dda96b6f17a802bdb7f6d429f?s=32&d=identicon&r=PG
https://www.gravatar.com/avatar/24780fb6df85a943c7aea0402c843737?s=32&d=identicon&r=PG

【讨论】:

  • 感谢您的时间和帮助,这是我收到的错误 Parse error: syntax error, unexpected 'libxml_use_internal_errors' (T_STRING) in C:\xampp\htdocs\grabIMG\index.php on line 16,这是第 16 行 libxml_use_internal_errors(true);
  • 缺少分号。
  • 我已经测试了代码,它按预期工作,检查更新的答案。
  • 但我不明白你的这部分代码,你是如何指定路径以及如何保存图像的? foreach ($xpath-&gt;query('//'.$the_tag.'[contains(@class,"'.$the_class.'")]/img') as $item)
  • 不保存图片,只显示url,你可以使用copy()来保存图片。
【解决方案2】:

也许你应该试试HTML DOM Parser for PHP。我最近发现了这个工具,老实说它工作得很好。正如您在网站上看到的那样,它是类似 JQuery 的选择器。我建议您看一下并尝试以下方法:

<?php
require_once("./simple_html_dom.php");
foreach ($html->find("<tag>") as $<tag>) //Start from the root (<html></html>) find the the parent tag you want to search in instead of <tag> (e.g "div" if you want to search in all divs)
{
    foreach ($<tag>->find("img") as $img) //Start searching for img tag in all (divs) you found
    {
        echo $img->src . "<br>"; //Output the information from the img's src attribute (if the found tag is <img src="www.example.com/cat.png"> you will get www.example.com/cat.png as result)
    }
}
?>

希望我对你的帮助更少或更多。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-11-30
    • 1970-01-01
    • 2014-02-15
    • 2012-04-14
    • 2015-06-08
    相关资源
    最近更新 更多