【问题标题】:PHP: regex search a pattern in a file and pick it upPHP:正则表达式在文件中搜索模式并选择它
【发布时间】:2011-02-11 20:17:43
【问题描述】:

我真的对 PHP 的正则表达式感到困惑。

无论如何,我现在无法阅读整个教程,因为我有一堆 html 文件,我必须尽快在其中找到链接。我想出了用我知道的语言的 php 代码来自动化它的想法。

所以我想我可以使用这个脚本:

$address = "file.txt"; 
$input = @file_get_contents($address) or die("Could not access file: $address");
$regexp = "??????????"; 
if(preg_match_all("/$regexp/siU", $input, $matches)) { 
    // $matches[2] = array of link addresses 
   // $matches[3] = array of link text - including HTML code 
} 

我的问题是$regexp

我需要的模式是这样的:

href="/content/r807215r37l86637/fulltext.pdf" title="Download PDF

我想从上面的行中搜索并获取/content/r807215r37l86637/fulltext.pdf,我在文件中有很多。

有什么帮助吗?

===================

编辑

标题属性对我来说很重要,我想要的所有属性都有标题

title="下载 PDF"

【问题讨论】:

    标签: php regex search


    【解决方案1】:

    再一次,正则表达式是bad for parsing html。

    保存您的理智并使用内置的 DOM 库。

    $dom = new DOMDocument();
    @$dom->loadHTML($html);
    $x = new DOMXPath($dom);
        $data = array();
    foreach($x->query("//a[@title='Download PDF']") as $node)
    {
        $data[] = $node->getAttribute("href");
    }
    

    编辑 根据 ircmaxell 注释更新代码。

    【讨论】:

    • 呃。如果您只想进行节点名搜索,为什么要使用 xpath?为什么不只是$dom->getElementsByTagName('a');?如果您执行$x->query('//a[contains(@title, "Download Pdf")]');,我可以理解xpath,这将返回完全匹配... ;-)
    • @ircmaxell,你完全正确。getElementsByTagName() 可能是一种更有效的方法..
    • @safaali 在查询中,将@title='Download Pdf' 更改为@class='nameOfClass' 或使用contains(@title, 'Download PDF')。即使里面有多余的东西,包含也会抓住它们。
    • 谢谢!我应该安装 DOMDocument 类吗?我在本地主机上使用 xampp 1.7.3。
    • @safaali,所有的dom库都是内置的,不需要安装任何东西。
    【解决方案2】:

    使用 phpQuery 或 QueryPath 会更容易:

    foreach (qp($html)->find("a") as $a) { 
        if ($a->attr("title") == "PDF") {
            print $a->attr("href");
            print $a->innerHTML();
        }
    }
    

    对于正则表达式,它取决于源的某种一致性:

    preg_match_all('#<a[^>]+href="([^>"]+)"[^>]+title="Download PDF"[^>]*>(.*?)</a>#sim', $input, $m);
    

    寻找固定的title="..." 属性是可行的,但更困难,因为它取决于右括号之前的位置。

    【讨论】:

    • @Byron:有些人讨厌不必要的繁琐 API。
    • @mario 您是否真的尝试过使用内置库进行 dom 解析?我承认,dom 上的 php 站点文档很麻烦。起初我也很抗拒,直到我看到了光。这真的很容易。如果你知道 xquery,DOMXPath::xquery 就是你所需要的。
    • @Byron:尝试并使用。但很像原始的 Javascript DOM 方法,我正在避免它。
    • @mario 很公平。只是好奇,这些库中的任何一个都使用内置的 php dom 吗?
    • @Byron:据我所知,实际上他们都这样做(phpQuery、QueryPath、FluentDom)。虽然 QP 带有自己的替代解析器,以实现更古怪的 HTML。
    【解决方案3】:

    试试这样的。如果它不起作用,请显示一些您要解析的链接示例。

    <?php
    $address = "file.txt"; 
    $input = @file_get_contents($address) or die("Could not access file: $address");
    $regexp = '#<a[^>]*href="([^"]*)"[^>]*title="Download PDF"#'; 
    
    if(preg_match_all($regexp, $input, $matches, PREG_SET_ORDER)) { 
      foreach ($matches as $match) {
        printf("Url: %s<br/>", $match[1]);
      }
    } 
    

    编辑:已更新,因此仅搜索下载“PDF 条目”

    【讨论】:

    【解决方案4】:

    最好的方法是使用DomXPath一步完成搜索:

    $dom = new DomDocument();
    $dom->loadHTML($html);
    $xpath = new DomXPath($dom);
    
    $links = array();
    foreach($xpath->query('//a[contains(@title, "Download PDF")]') as $node) {
        $links[] = $node->getAttribute("href");
    }
    

    甚至:

    $links = array();
    $query = '//a[contains(@title, "Download PDF")]/@href';
    foreach($xpath->evaluate($query) as $attr) {
        $links[] = $attr->value;
    }
    

    【讨论】:

      【解决方案5】:

      href="([^]+)" 将为您提供该表单的所有链接。

      【讨论】:

      • 谢谢,但是文件中有很多herfs,我想要标题为“下载PDF”的链接
      猜你喜欢
      • 2023-01-30
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2023-03-17
      相关资源
      最近更新 更多