【问题标题】:Highlight keywords in a paragraph突出段落中的关键字
【发布时间】:2011-05-04 03:23:25
【问题描述】:

我需要在段落中突出显示关键字,就像 google 在其搜索结果中所做的那样。假设我有一个带有博客文章的 MySQL 数据库。当用户搜索某个关键字时,我希望返回包含这些关键字的帖子,但只显示部分帖子(包含搜索关键字的段落)并突出显示这些关键字。

我的计划是这样的:

  • 找到内容中包含搜索关键字的帖子ID;
  • 再次阅读该帖子的内容并将每个单词放入一个固定的缓冲区数组(50 个单词),直到找到关键字。

你能帮我一些逻辑,或者至少告诉我我的逻辑是否正常吗?我处于 PHP 学习阶段。

【问题讨论】:

  • 您不需要使用数组来存储每个单词,而且可能也不应该这样做。
  • 您的数据是如何存储的?是纯文本还是 HTML?
  • 哦,你想如何匹配和突出匹配?匹配整个词或子词?突出显示整个词或子词?
  • @Gumbo:它以纯文本形式存储。 TY 为您提供宝贵的回放和 cmets。

标签: php search string


【解决方案1】:

当你连接到数据库时,也许你可以做这样的事情:

$keyword = $_REQUEST["keyword"]; //fetch the keyword from the request
$result = mysql_query("SELECT * FROM `posts` WHERE `content` LIKE '%".
        mysql_real_escape_string($keyword)."%'"); //ask the database for the posttexts
while ($row = mysql_fetch_array($result)) {//do the following for each result:
  $text = $row["content"];//we're only interested in the content at the moment
  $text=substr ($text, strrpos($text, $keyword)-150, 300); //cut out
  $text=str_replace($keyword, '<strong>'.$keyword.'</strong>', $text); //highlight
  echo htmlentities($text); //print it
  echo "<hr>";//draw a line under it
}

【讨论】:

    【解决方案2】:

    如果您是初学者,这不会像有人想象的那样超级简单......

    我认为您应该执行以下步骤:

    1. 根据用户搜索的内容构建查询(谨防 sql 注入)
    2. 获取结果并组织它们(数组应该没问题)
    3. 从前一个数组构建 html 代码

    在第三步中,您可以使用一些正则表达式将用户搜索的关键字替换为粗体等效项。 str_replace 也可以工作...

    我希望这会有所帮助... 如果你能提供你的数据库结构,也许我可以给你一些更准确的提示......

    【讨论】:

      【解决方案3】:

      如果你想剪掉相关的段落,在做完上面提到的str_replace函数之后,你可以使用stripos()来找到这些强段的位置,然后用那个位置的偏移量和substr()来剪掉段落的一部分,例如:

      $搜索词; foreach($searchterms 作为 $search) { $paragraph = str_replace($search, "$search", $paragraph); } $pos = 0; for($i = 0; $i ", $pos); $section[$i] = substr($paragraph, $pos - 100, 200); }

      这将为您提供一系列小句子(每个 200 个字符),以按照您的意愿使用。从剪切位置搜索最近的空格,并从那里剪切以防止出现半字,这也可能是有益的。哦,你还需要检查错误,但我会留给你。

      【讨论】:

        【解决方案4】:

        如果它包含 html(请注意,这是一个非常强大的解决方案):

        $string = '<p>foo<b>bar</b></p>';
        $keyword = 'foo';
        $dom = new DomDocument();
        $dom->loadHtml($string);
        $xpath = new DomXpath($dom);
        $elements = $xpath->query('//*[contains(.,"'.$keyword.'")]');
        foreach ($elements as $element) {
            foreach ($element->childNodes as $child) {
                if (!$child instanceof DomText) continue;
                $fragment = $dom->createDocumentFragment();
                $text = $child->textContent;
                $stubs = array();
                while (($pos = stripos($text, $keyword)) !== false) {
                    $fragment->appendChild(new DomText(substr($text, 0, $pos)));
                    $word = substr($text, $pos, strlen($keyword));
                    $highlight = $dom->createElement('span');
                    $highlight->appendChild(new DomText($word));
                    $highlight->setAttribute('class', 'highlight');
                    $fragment->appendChild($highlight);
                    $text = substr($text, $pos + strlen($keyword));
                }
                if (!empty($text)) $fragment->appendChild(new DomText($text));
                $element->replaceChild($fragment, $child);
            }
        }
        $string = $dom->saveXml($dom->getElementsByTagName('body')->item(0)->firstChild);
        

        结果:

        <p><span class="highlight">foo</span><b>bar</b></p>
        

        还有:

        $string = '<body><p>foobarbaz<b>bar</b></p></body>';
        $keyword = 'bar';
        

        你得到(为了可读性分成多行):

        <p>foo
            <span class="highlight">bar</span>
            baz
            <b>
                <span class="highlight">bar</span>
            </b>
        </p>
        

        提防非 dom 解决方案(例如 regex 或 str_replace),因为突出显示“div”之类的内容会完全破坏您的 HTML...这只会“突出显示”正文中的字符串,永远不会在标签内...


        编辑既然您想要 Google 风格的结果,这里有一种方法:

        function getKeywordStubs($string, array $keywords, $maxStubSize = 10) {
            $dom = new DomDocument();
            $dom->loadHtml($string);
            $xpath = new DomXpath($dom);
            $results = array();
            $maxStubHalf = ceil($maxStubSize / 2);
            foreach ($keywords as $keyword) {
                $elements = $xpath->query('//*[contains(.,"'.$keyword.'")]');
                $replace = '<span class="highlight">'.$keyword.'</span>';
                foreach ($elements as $element) {
                    $stub = $element->textContent;
                    $regex = '#^.*?((\w*\W*){'.
                         $maxStubHalf.'})('.
                         preg_quote($keyword, '#').
                         ')((\w*\W*){'.
                         $maxStubHalf.'}).*?$#ims';
                    preg_match($regex, $stub, $match);
                    var_dump($regex, $match);
                    $stub = preg_replace($regex, '\\1\\3\\4', $stub);
                    $stub = str_ireplace($keyword, $replace, $stub);
                    $results[] = $stub;
                }
            }
            $results = array_unique($results);
            return $results;
        }
        

        好的,那么它的作用是返回一个匹配数组,其中包含 $maxStubSize 单词(即之前最多一半,之后一半)...

        所以,给定一个字符串:

        <p>a whole 
            <b>bunch of</b> text 
            <a>here for</a> 
            us to foo bar baz replace out from this string
            <b>bar</b>
        </p>
        

        调用getKeywordStubs($string, array('bar', 'bunch')) 将导致:

        array(4) {
          [0]=>
          string(75) "here for us to foo <span class="highlight">bar</span> baz replace out from "
          [3]=>
          string(34) "<span class="highlight">bar</span>"
          [4]=>
          string(62) "a whole <span class="highlight">bunch</span> of text here for "
          [7]=>
          string(39) "<span class="highlight">bunch</span> of"
        }
        

        因此,您可以通过按strlen 对列表进行排序然后选择两个最长的匹配项来构建您的结果简介...(假设 php 5.3+):

        usort($results, function($str1, $str2) { 
            return strlen($str2) - strlen($str1);
        });
        $description = implode('...', array_slice($results, 0, 2));
        

        结果:

        here for us to foo <span class="highlight">bar</span> baz replace out...a whole <span class="highlight">bunch</span> of text here for 
        

        我希望这会有所帮助...(我确实觉得这有点...臃肿...我确信有更好的方法可以做到这一点,但这里有一种方法)...

        【讨论】:

        • 虽然是一个很好的突出显示解决方案,但它不能解决 OP 的问题(至少如果我理解正确的话)。 OP 希望突出显示并返回部分环境,例如 Google 搜索结果摘录。
        • @Gordon:以一种非常不优雅的方式编辑。
        • @ircmaxell 我收到以下警告:警告:DOMDocument::loadHTML() [domdocument.loadhtml]: htmlParseEntityRef: expecting ';'在 Entity 中,line: 1 in /../ 在第 118 行。通常使用 hetmlentities() 解决它会破坏 DOM 突出显示的优点
        • 不错的解决方案,尽管值得注意的是它区分大小写,尽管使用了stipos,因为xpath 函数contains 区分大小写。我发现的唯一方法是测试几个匹配项,例如'//*[contains(.,"'.$keyword.'") or contains(.,"'.strtolower($keyword).'") or contains(.,"'.strtoupper($keyword).'") or contains(.,"'.ucfirst($keyword).'")]')
        【解决方案5】:

        您可以尝试使用explode 将您的数据库搜索结果集分解为一个数组,然后在每个搜索结果上使用array_search()。将下面示例中的 $distance 变量设置为您希望在 $keyword 的第一个匹配项的任一侧出现的单词数。

        在示例中,我包含了 lorum ipsum 文本作为示例数据库结果段落,并将 $keyword 设置为“scelerisque”。你显然会在你的代码中替换这些。

        //example paragraph text
        $lorum = 'Nunc nec magna at nibh imperdiet dignissim quis eu velit. 
        vel mattis odio rutrum nec. Etiam sit amet tortor nibh, molestie 
        vestibulum tortor. Integer condimentum magna dictum purus vehicula 
        et scelerisque mauris viverra. Nullam in lorem erat. Ut dolor libero, 
        tristique et pellentesque sed, mattis eget dui. Cum sociis natoque 
        penatibus et magnis dis parturient montes, nascetur ridiculus mus. 
        .';
        
        //turn paragraph into array
        $ipsum = explode(' ',$lorum);
        //set keyword
        $keyword = 'scelerisque';
        //set excerpt distance
        $distance = 10;
        
        //look for keyword in paragraph array, return array key of first match
        $match_key = array_search($keyword,$ipsum);
        
        if(!empty($match_key)){
        
            foreach($ipsum as $key=>$value){
                //if paragraph array key inside excerpt distance
                if($key > $match_key-$distance and $key< $match_key+$distance){ 
                    //if array key matches keyword key, bold the word
                    if($key == $match_key){
                        $word = '<b>'.$value.'</b>';
                        }
                    else{
                        $word = $value;
                        }
                    //create excerpt array to hold words within distance
                    $excerpt[] = $word;
                    }
        
                }
            //turn excerpt array into a string
            $excerpt = implode(' ',$excerpt);
            }
        //print the string
        echo $excerpt;
        

        $excerpt 返回: "前庭侵权。整数 condimentum magna dictum purus vehicula et scelerisque mauris viverra. Nullam in lorem erat. Ut dolor libero,"

        【讨论】:

          【解决方案6】:

          这是纯文本的解决方案:

          $str = 'Lorem ipsum dolor sit amet, consectetur adipisicing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum.';
          $keywords = array('co');
          $wordspan = 5;
          $keywordsPattern = implode('|', array_map(function($val) { return preg_quote($val, '/'); }, $keywords));
          $matches = preg_split("/($keywordsPattern)/ui", $str, -1, PREG_SPLIT_DELIM_CAPTURE);
          for ($i = 0, $n = count($matches); $i < $n; ++$i) {
              if ($i % 2 == 0) {
                  $words = preg_split('/(\s+)/u', $matches[$i], -1, PREG_SPLIT_DELIM_CAPTURE);
                  if (count($words) > ($wordspan+1)*2) {
                      $matches[$i] = '…';
                      if ($i > 0) {
                          $matches[$i] = implode('', array_slice($words, 0, ($wordspan+1)*2)) . $matches[$i];
                      }
                      if ($i < $n-1) {
                          $matches[$i] .= implode('', array_slice($words, -($wordspan+1)*2));
                      }
                  }
              } else {
                  $matches[$i] = '<b>'.$matches[$i].'</b>';
              }
          }
          echo implode('', $matches);
          

          与当前模式"/($keywordsPattern)/ui" 匹配并突出显示子词。但如果你愿意,你可以改变它:

          • 如果您只想匹配整个单词而不是子单词,请使用单词边界\b:

            "/\b($keywordsPattern)\b/ui"
            
          • 如果要匹配子词但要突出显示整个单词,请在关键字前后使用可选单词字符\w:

            "/(\w*?(?:$keywordsPattern)\w*)/ui"
            

          【讨论】:

            【解决方案7】:

            我在搜索如何突出显示关键字搜索结果时发现了这篇文章。我的要求是:

            • 必须是完整的单词
            • 必须适用于多个关键字
            • 只能是 PHP

            我正在从一个不包含元素的 MySQL 数据库中获取我的数据,这是根据存储数据的表单设计的。

            这是我发现最有用的代码:

            $keywords = array("fox","jump","quick");
            $string = "The quick brown fox jumps over the lazy dog";
            $test = "The quick brown fox jumps over the lazy dog"; // used to compare values at the end.
            
            if(isset($keywords)) // For keyword search this will highlight all keywords in the results.
                {
                foreach($keywords as $word)
                    {
                    $pattern = "/\b".$word."\b/i";
                    $string = preg_replace($pattern,"<span class=\"highlight\">".$word."</span>", $string);
                    }
                }
             // We must compare the original string to the string altered in the loop to avoid having a string printed with no matches.
            if($string === $test)
                {
                echo "No match";
                }
            else
                {
                echo $string;
                }
            

            输出:

            The <span class="highlight">quick</span> brown <span class="highlight">fox</span> jumps over the lazy dog.
            

            我希望这对某人有所帮助。

            【讨论】:

            • 据我了解,这对于要求 1 而言将失败。如果您碰巧在关键字列表中包含“row”,那么部分匹配项(例如“brown”)也会突出显示。这取决于您对“整个单词”的概念。在关键字中?在您正在查看的文本中?在这两个地方?
            猜你喜欢
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 2017-05-24
            • 2016-09-18
            • 2012-06-28
            • 1970-01-01
            相关资源
            最近更新 更多