【问题标题】:Replace only the second html tag [closed]仅替换第二个 html 标记 [关闭]
【发布时间】:2012-11-19 18:23:54
【问题描述】:

我想将第二个h2 标记替换为h3,希望有人可以帮助我替换正则表达式,或者preg_split - 我不太确定。

例如,这个:

<h2>My text one</h2>
<h2>My text two</h2>
text …
<h2>My text three</h2>

应该变成这样:

<h2>My text one</h2>
<h3>My text two</h3>
text …
<h2>My text three</h2>

【问题讨论】:

  • 为什么要用 PHP 替换 html 元素我认为是这里的问题。你确定你说的不是 Javascript 吗?
  • 它们是否像您的示例中那样被新行分隔?
  • 你应该使用 DOM 解析器。
  • 您需要让我们了解上下文。这需要在哪里发生 - 客户端或服务器端?如果是服务器端,打印标题标签的代码在哪里?
  • 不要使用正则表达式解析PHP。 htmlparsing.com/regexes.html 解释了原因。使用适当的 DOM 解析器。 htmlparsing.com/php.html 给出了一些例子。

标签: php regex html-parsing preg-replace preg-split


【解决方案1】:

我同意其他评论,这应该通过 dom 解析器来完成。但是这里有一个可行的 php 解决方案。

<?php 
     // Fill $str with the html;

     preg_replace("/[h]{1}[2]/i", "h3", $str);
?>

或

<?php
     // Fill $str with the html;

     str_replace("h2", "h3", $str);      
?>

这应该可以正常工作。将 $matches 参数添加到 preg_replace 也将跟踪所做的更改数量。 现在,使用循环您可以控制需要替换哪个元素,但是,上面编写的函数将检测 h2 的 所有 出现。

另外,我将正则表达式过度复杂化,以便您能够换出数字,从而使用它创建更有用的功能。只需使用“/(h2)/i”也可以解决问题。

因此,您的代码应该以正确的方式实现循环以防止替换所有标签,并且您应该决定该函数是仅处理 h2 还是应该更灵活。

最后,str_replace 比 preg_replace 快,所以如果这是您唯一需要进行的编辑,我建议您使用 str_replace。

【讨论】:

    【解决方案2】:

    您可以使用 Javascript 轻松做到这一点,真的需要使用 PHP 吗?

    获取第二个&lt;h2&gt;值

    $text = $("h2:eq(1)").html();
    

    摧毁它。

    $("h2:eq(1)").remove();
    

    在第一个&lt;h2&gt; 之后创建一个&lt;h3&gt;,其中包含$text

    $("h2:eq(0)").after("<h3>" + $text + "</h3>");
    

    【讨论】:

    • OP 没有说它不可能是 javascript。
    • 在服务器端做某事和在客户端做某事之间有很大的区别。他标记了 php,询问了 php,但从未提及 javascript。不过,他对他实际尝试做什么以及他正在使用什么非常模糊,所以,第二个,他可能指的是 javascript。
    【解决方案3】:

    您不需要为此使用服务器端 HTML 解析器,这在 imo 中完全是多余的。以下是一个显然可能被某些 HTML 结构破坏的示例,但对于大多数标记来说,它不会有任何问题 - 并且比服务器端 HTML 解析器更优化。

    $html = '
    <h2>My text one</h2>
    <h2>My text two</h2>
    text ...
    <h2>My text three</h2>
    ';
    

    preg_match_all

    /// the following preg match will find all <h2> mark-up, even if 
    /// the content of the h2 splits over new lines - due to the `s` switch
    /// It is a non-greedy match too - thanks to the `.+?` so it shouldn't 
    /// have problems with spanning over more than one h2 tag. It will only
    /// really break down if you have a h2 as a descendant of a h2 - which
    /// would be illegal html - or if you have a `>` in one of your h2's
    /// attributes i.e. <h2 title="this>will break">Text</h2> which again
    /// is illegal as they should be encoded.
    
    preg_match_all(
      '#(<)h2([^>]*>.+?</)h2(>)#is',
      $html,
      $matches,
      PREG_OFFSET_CAPTURE|PREG_SET_ORDER
    );
    

    替换和重建

    /// Because you wanted to only replace the 2nd item use the following. 
    /// You could however make this code as general or as specific as you wanted.
    /// The following works because the surrounding content for the found 
    /// $matches was stored using the grouping brackets in the regular 
    /// expression. This means you could easily change the regexp, and the 
    /// following code would still work.
    
    /// to get a better understanding of what is going on it would be best
    /// to `echo '<xmp>';print_r( $matches );echo '/<xmp>';`
    
    if ( isset($matches[1][0]) ) {
      $html = substr( $html, 0, $matches[1][0][1] ) . 
              $matches[1][1][0] . 'h3' . 
              $matches[1][2][0] . 'h3' . 
              $matches[1][3][0] .
              substr( $html, $matches[1][0][1] + strlen($matches[1][0][0]) );
    }
    

    我不知道为什么很多人说要使用客户端 JavaScript 来进行这种更改,PHP 代表 PHP: Hypertext Preprocessor,它旨在预处理超文本。 OP 只提到过 PHP 函数,并用 PHP 标记了这篇文章,所以没有任何东西指向客户端。

    确实,尽管可以并且应该尽可能使用客户端来减轻服务器端的处理,但不建议将其用于诸如标题之类的核心结构标签 - 屏幕阅读器和搜索引擎机器人将依赖它。最好使用客户端 JavaScript 来增强用户体验。如果您使用它来极大地增强您网站的功能,您最好确保您的整个用户群都支持它。

    但是,如果你们中的任何人提到 Node.js 和 JSDOM,我会很高兴地同意。

    【讨论】:

    • 这还有待观察。浏览器不需要为它做很多额外的工作。它已经有了dom-tree。重点是,请求是更改硬编码的 html。在我看来,这应该尽可能地留给客户端。
    • @Digitalis 抱歉,我的答案中没有包括客户端,因为 OP 没有要求它。关于更改标题标签,这对于 SEO 而言不是 JS 的好主意。
    • 为什么在客户端更改它会对 SEO 产生负面影响?此外,没有造成任何伤害。您提供的示例也完成了它需要做的事情,只要它富有成效,我总是愿意讨论。
    • @Digital 是肯定的,不用担心 :) 它会影响 SEO/屏幕阅读器,因为许多机器人/阅读器不会执行 JavaScript - 所以在这个例子中,标题仍然是 H2,而不是给出文档大纲它是正确的层次结构,即 H1、H2、H3。虽然是的,对 SEO 的影响可能很小,但这会对屏幕阅读器产生很大的影响...... (所有标题都将以同样的重要性被处理和读出 - 如果你看不到,相当混乱和烦人) 因此,核心标记结构最好由服务器端处理。
    • 加一个,好点。没想到。老实说,我更喜欢在服务器端切换类,但这是你在这里提出的一个很好的观点。这可能会在很长一段时间内被忽视,但实际上会影响您的 SEO。谢谢老哥!
    猜你喜欢
    • 1970-01-01
    • 2021-02-24
    • 1970-01-01
    • 2021-03-25
    • 2017-11-03
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-04-09
    相关资源
    最近更新 更多