【问题标题】:How to Extract Particular String from the HTML Source code using PHP如何使用 PHP 从 HTML 源代码中提取特定字符串
【发布时间】:2017-05-23 20:23:13
【问题描述】:

我正在尝试从整个 HTML 源代码中提取特定的字符串。

HTML来源:view-source:https://www.instagram.com/p/BUbZXXMjnxY/?taken-by=narentrigger&hl=en

需要提取字符串:https://instagram.fmaa1-2.fna.fbcdn.net/t51.2885-15/e35/18645014_163619900839441_7821159798480568320_n.jpg 来自“og:image”元属性。

我尝试了一些方法,但一切都出错了。有没有办法从源代码的 og:image 元属性中获取图像链接。提取后需要将图像 url 存储在特定变量上。需要专家帮助。 Url that need to extract

【问题讨论】:

  • 你可以使用 PHP DomDocument 来构建一个爬虫。 php.net/manual/en/class.domdocument.php
  • 你为什么不分享那些“一些方法”以及“一切都出错”是什么意思,即你得到了什么具体的错误?
  • 所以要提取og:image元的content属性?
  • 是的,我需要从整个源代码@BenM 中提取 og:image meta 的 content 属性
  • @Narendhiranvignesh 请看我的回答。

标签: php html string url substring


【解决方案1】:

如果您只获取一个子字符串,请不要使用preg_match_all()。加载 DOMDocument 对于这项任务来说似乎有点过头了。

通过使用\K,您可以减少结果数组膨胀。

示例输入:

$input='<meta property="og:title" content="Instagram post by Narendiran blah blah" />
<meta property="og:image" content="https://instagram.fmma1-2.blah.jpg" />
<meta property="og:description" content="8 Likes, 1 Comments - blah" />';

方法(Demo):

$url=preg_match('/"og:image"[^"]+"\K[^"]+/',$input,$out)?$out[0]:null;
echo $url;

输出:

https://instagram.fmma1-2.blah.jpg

通过使用否定字符类,正则表达式引擎将更有效地运行。 [^"]。 (Pattern Demo)

【讨论】:

    【解决方案2】:

    假设您使用 PHP 在字符串中包含标记,RegEx 有什么问题?

    preg_match_all('/<meta.*property="og:image".*content="(.*)".*\/>/', $string, $matches);
    echo $matches[1][0];
    

    Demo

    免责声明:可能会提供更高效的正则表达式。

    【讨论】:

    • 上述方法效果很好。如何在不假设标记的情况下从整个源代码中提取上述字符串(因为它可能因不同的 html 源而异)。是否可以从整个 HTML 源代码中提取“og:meta”元的“内容”属性。如果可能的话,您能否提供方法..
    • 以上内容适用于此。 $string 可以包含您拥有的任何源代码。
    • 是的,我已经按照我的要求执行了代码,上面的代码运行良好。谢谢@BenM
    【解决方案3】:

    在这段代码 sn-p 中,我使用 DOMDocument 从元标记中删除属性内容。它将它存储在一个数组中,以防有更多并返回它。 希望它有效。

       function get_img_url($url) { 
    
            // Create a new DOM object 
            $html = new DOMDocument(); 
    
            // load the HTML page 
            $html->loadHTMLFile($url); 
    
            // create a empty array object 
            $imageArray = array(); 
    
            //Loop through each meta tag
            foreach($html->getElementsByTagName('meta') as $meta) { 
                $imageArray[] = array('url' => $meta->getAttribute('content')); 
            } 
    
            //Return the list 
            return $imageArray; 
        } 
    

    【讨论】:

      【解决方案4】:

      试试这个代码来报废网页。 我使用了 simple_html_dom_parser。 你可以从https://sourceforge.net/projects/simplehtmldom/files/下载它

      include_once("simple_html_dom.php");
      
      $output_filename = "example_homepage.html";
      $fp = fopen($output_filename, 'w');
      $url = 'https://www.instagram.com/p/BUbZXXMjnxY/?taken-by=narentrigger&hl=en';
      $curl = curl_init();
      
      curl_setopt($curl, CURLOPT_URL, $url);
      curl_setopt($curl, CURLOPT_RETURNTRANSFER, false);
      curl_setopt ($curl, CURLOPT_FILE, $fp);
      $result = curl_exec($curl);
      
      curl_close($curl);
      fclose($fp);
      
      $html = file_get_html('example_homepage.html');
      
      foreach($html->find('meta[property=og:image]') as $element) 
         echo $element->content . '<br>';
      

      【讨论】:

        猜你喜欢
        • 2020-02-16
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多