【问题标题】:Python's equivalent to PHP's strip_tags?Python 相当于 PHP 的 strip_tags?
【发布时间】:2010-02-19 11:38:14
【问题描述】:

Python 相当于 PHP 的 strip_tags?

http://php.net/manual/en/function.strip-tags.php

【问题讨论】:

标签: php python strip


【解决方案1】:
from bleach import clean
print clean("<strong>My Text</strong>", tags=[], strip=True, strip_comments=True)

【讨论】:

    【解决方案2】:

    Python 标准库中没有这样的东西。这是因为 Python 是一种通用语言,而 PHP 最初是一种面向 Web 的语言。

    不过,您有 3 个解决方案:

    • 你很着急:自己做吧。 re.sub(r'&lt;[^&gt;]*?&gt;', '', value) 可能是一个快速而肮脏的解决方案。
    • 使用第三方库(推荐因为更防弹):beautiful soup 是一个非常好的库,无需安装,只需复制 lib 目录并导入。 Full tuto with beautiful soup。
    • 使用框架。大多数 Web Python 开发人员从不从头开始编写代码,他们使用诸如 django 之类的框架来自动为您完成这些工作。 Full tuto with django。

    【讨论】:

    • @Hank 对任何混淆感到抱歉。遗憾的是,这不再是一个好方法。
    【解决方案3】:

    使用BeautifulSoup

    from BeautifulSoup import BeautifulSoup
    soup = BeautifulSoup(htmltext)
    ''.join([e for e in soup.recursiveChildGenerator() if isinstance(e,unicode)])
    

    【讨论】:

    • 你可能想让他知道这是第三方库。
    【解决方案4】:

    您不会找到很多内置 Python 等价物用于内置 PHP HTML 函数,因为 Python 更像是一种通用脚本语言,而不是一种 Web 开发语言。对于 HTML 处理,一般推荐BeautifulSoup。

    【讨论】:

      【解决方案5】:

      Python 没有内置的,但有一个 ungodly number of implementations。

      【讨论】:

      • 实际上,我用谷歌搜索了你建议的相同查询,但这些并没有让我满意。
      【解决方案6】:

      我使用 HTMLParser 类为 Python 3 构建了一个。它比 PHP 更冗长。我把它叫做 HTMLCleaner 类,你可以找到源代码here,你可以找到例子here。

      【讨论】:

        【解决方案7】:

        为此有一个活动状态配方,

        http://code.activestate.com/recipes/52281/

        这是旧代码,因此您必须将 sgml 解析器更改为 HTMLparser,如 cmets 中所述

        这是修改后的代码,

        import HTMLParser, string
        
        class StrippingParser(HTMLParser.HTMLParser):
        
            # These are the HTML tags that we will leave intact
            valid_tags = ('b', 'a', 'i', 'br', 'p', 'img')
        
            from htmlentitydefs import entitydefs # replace entitydefs from sgmllib
        
            def __init__(self):
                HTMLParser.HTMLParser.__init__(self)
                self.result = ""
                self.endTagList = []
        
            def handle_data(self, data):
                if data:
                    self.result = self.result + data
        
            def handle_charref(self, name):
                self.result = "%s&#%s;" % (self.result, name)
        
            def handle_entityref(self, name):
                if self.entitydefs.has_key(name): 
                    x = ';'
                else:
                    # this breaks unstandard entities that end with ';'
                    x = ''
                self.result = "%s&%s%s" % (self.result, name, x)
        
            def handle_starttag(self, tag, attrs):
                """ Delete all tags except for legal ones """
                if tag in self.valid_tags:       
                    self.result = self.result + '<' + tag
                    for k, v in attrs:
                        if string.lower(k[0:2]) != 'on' and string.lower(v[0:10]) != 'javascript':
                            self.result = '%s %s="%s"' % (self.result, k, v)
                    endTag = '</%s>' % tag
                    self.endTagList.insert(0,endTag)    
                    self.result = self.result + '>'
        
            def handle_endtag(self, tag):
                if tag in self.valid_tags:
                    self.result = "%s</%s>" % (self.result, tag)
                    remTag = '</%s>' % tag
                    self.endTagList.remove(remTag)
        
            def cleanup(self):
                """ Append missing closing tags """
                for j in range(len(self.endTagList)):
                        self.result = self.result + self.endTagList[j]    
        
        
        def strip(s):
            """ Strip illegal HTML tags from string s """
            parser = StrippingParser()
            parser.feed(s)
            parser.close()
            parser.cleanup()
            return parser.result
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2011-10-10
          • 2010-10-28
          • 2012-12-03
          • 1970-01-01
          • 2011-01-25
          • 1970-01-01
          • 1970-01-01
          • 2013-05-13
          相关资源
          最近更新 更多