【问题标题】:twitter regex breaking, thinking syntax issue推特正则表达式破坏,思考语法问题
【发布时间】:2012-01-14 05:20:33
【问题描述】:

我的正则表达式语法以某种方式破坏了“rel=”行周围的 链接。

这里是:

<?php   
function parseTweet($text) {
   $pattern_url = '~(?>[a-z+]{2,}://|www\.)(?:[a-z0-9]+(?:\.[a-z0-9]+)?@)?(?:(?:[a-z](?:[a-z0-9]|(?<!-)-)*[a-z0-9])(?:\.[a-z](?:[a-z0-9]|(?<!-)-)*[a-z0-9])+|(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))(?:/[^\\/:?*"|\n]*[a-z0-9])*/?(?:\?[a-z0-9_.%]+(?:=[a-z0-9_.%:/+-]*)?(?:&[a-z0-9_.%]+(?:=[a-z0-9_.%:/+-]*)?)*)?(?:#[a-z0-9_%.]+)?~i';
    '@([A-Za-z0-9_]+)';

   $tweet = preg_replace('/(^|\s)#(\w+)/', '\1#<a href="http://search.twitter.com/search?q=%23\2? rel="nofollow">\2</a>', $text);
   $tweet = preg_replace('/(^|\s)@(\w+)/', '\1@<a href="http://www.twitter.com/\2? rel="nofollow">\2</a>', $tweet);
   $tweet = preg_replace('#(^|[\n ])(([\w]+?://[\w\#$%&~.\-;:=,?@\[\]+]*)(/[\w\#$%&~/.\-;:=,?@\[\]+]*)?)#is', '\\1
                      <a href=\"\\2\" title=\"\\2\" rel=\"nofollow\">[link]</a>', $tweet);
   return $tweet;
}

$username='stephenfry'; // set user name
$format='json'; // set format
$tweet=json_decode(file_get_contents("http://api.twitter.com/1/statuses/user_timeline/{$username}.{$format}")); // get tweets and decode them into a variable

$theTweet = parseTweet($tweet[0]->text);

echo $theTweet; 
?>   

链接解析的 HTML:

Great deal: Jot by Adonit, a precise capacitive touch stylus, today 15% off with coupon code: 'Jot' -
<a rel="\"nofollow\"" title="\"http://t.co/QvFi6CKK\"" href="\"http://t.co/QvFi6CKK\"">[link]</a>

哈希标签解析 HTML:

I'm so sorry - that last #
<a nofollow"="" href="http://search.twitter.com/search?q=%23GameOfShadowsUK? rel=">GameOfShadowsUK</a>
tweet should hav 3been sent at 2:21 - my f****d up arsing w**k-mess of a life disallowed it :-( 

将不可靠的代码装箱并采用更好的方法。见答案。

【问题讨论】:

  • 你能发布一些输入/输出与预期输出的对比吗?
  • 当然,发布了一些示例推文的输出。
  • 你能把原件也发一下吗?
  • 您是否考虑过使用 DOMDocument 而不是使用正则表达式对其进行攻击?
  • @Evildonald。这是一条真实的、大手笔的推文。所以,除非你反对“hav 3been”的糟糕用法,否则你的牛肉是什么? twitter.com/#!/stephenfry/status/144073211276050435

标签: php regex json twitter


【解决方案1】:
            <?php

            function getLastXTwitterStatus($userid,$x){
            $url = "http://twitter.com/statuses/user_timeline/$userid.xml?count=$x";

            $xml = simplexml_load_file($url) or die('could not connect');
                echo '<ul>';
                   foreach($xml->status as $status){
                   $text = twitterify( $status->text );
                   echo '<li>'.utf8_decode($text).'</li>';
                   }
                echo '</ul>';
             }

             function twitterify($ret) {
              $ret = preg_replace("#(^|[\n ])([\w]+?://[\w]+[^ \"\n\r\t< ]*)#", "\\1<a href=\"\\2\" >\\2</a>", $ret);
              $ret = preg_replace("#(^|[\n ])((www|ftp)\.[^ \"\t\n\r< ]*)#", "\\1<a href=\"http://\\2\" >\\2</a>", $ret);
              $ret = preg_replace("/@(\w+)/", "<a href=\"http://www.twitter.com/\\1\" >@\\1</a>", $ret);
              $ret = preg_replace("/#(\w+)/", "<a href=\"http://search.twitter.com/search?q=\\1\" >#\\1</a>", $ret);
            return $ret;
            }

            //my user id kenrick1991
            getLastXTwitterStatus('simonpegg',1);

            ?> 

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2016-08-31
    • 1970-01-01
    • 2012-08-11
    • 1970-01-01
    • 1970-01-01
    • 2017-02-20
    • 2012-02-24
    相关资源
    最近更新 更多