【问题标题】:OpenTbs convert html tags to MS Word tagsOpenTbs 将 html 标签转换为 MS Word 标签
【发布时间】:2012-03-08 02:26:56
【问题描述】:

我正在使用 OpenTbs,http://www.tinybutstrong.com/plugins/opentbs/tbs_plugin_opentbs.html。

我有一个 template.docx,可以用内容替换字段,但如果内容有 html 代码,它会显示在模板创建的文档中。

First list <br /> Second Line

我尝试过使用:

$TBS->LoadTemplate('document.docx', OPENTBS_ALREADY_XML); 

认为这样可以让我用 ms office 标签替换我的 html 标签,但它只是在文档中显示了 MS Office 标签:

First Line<w:br/> Second Line

如何将 HTML 标记转换为 MS Office XML 等效项。

【问题讨论】:

  • 我在这里拉头发,这应该很常见,我用 [b.thetext] 替换标签的文本将具有 HTML 样式,我想将其转换为 MS Word 样式,如粗体和斜体,我找到了 MS Word 等价物,但无法让它们作为代码发挥作用,它们只是作为文本输出到模板中......任何帮助将不胜感激。

标签: php templates ms-word tinybutstrong opentbs


【解决方案1】:

既然你有一个HTML到DOCX的转换函数,那么你可以在OpenTBS中使用自定义的PHP函数和参数“onformat”来实现它。

以下函数只转换换行符:

function f_html2docx($FieldName, &$CurrVal) {
  $CurrVal= str_replace('<br />', '<w:br/>', $CurrVal);
} 

在 DOCX 模板中使用:

[b.thetext;onformat=f_html2docx]

关于将 HTML 转换为 DOCX:

将一个格式化文本转换成另一个格式化文本通常是一场噩梦。这就是为什么存储纯数据而不是格式化数据是明智的。

将 HTML 转换为 DOCX 是一场真正的噩梦,因为格式的结构不同。

例如,在 HTML 标签中我可以嵌套,像这样:

<i> hello <b> this is important </b> to know </i>

在 DOCX 中,它将显示为交叉,如下所示:

  <w:r>
    <w:rPr><w:b/></w:rPr>
    <w:t>hello</w:t>
  </w:r>

  <w:r>
    <w:rPr><w:b/><w:i/></w:rPr>
    <w:t>this is important</w:t>
  </w:r>

  <w:r>
    <w:rPr><w:i/></w:rPr>
    <w:t>to know</w:t>
  </w:r>

目前除了换行符之外,我没有其他转换标签的解决方案。对此感到抱歉。 而且我认为编写一个代码会非常困难。

【讨论】:

  • 我想知道你最终找到解决方案了吗?我正在寻找类似的东西。我只会在 HTML 中进行基本格式设置,例如粗体、下划线和斜体等。
  • 我没有研究过这样的解决方案,我只是报告别人的解决方案。
  • 谢谢,我想我将使用 PDF 来代替,因为我知道这会起作用。谢谢你
【解决方案2】:

感谢 Skrol 就我所有的 openTBS 问题提供意见,只是注意到你是它的创造者,它是一门很棒的课程,经过一天的学习 MS Word 格式后,你上面所说的是真实的,我有一个大脑Wave,我现在可以生成您在上面指定的格式,并且可以使用粗斜体和下划线,这是我所需要的,我希望这可以为您提供改进的基础。

我基本上注意到,在您放置的示例中,您只需要一个样式数组,当您找到一个结束标记时,您会从样式数组中删除。每次找到标签时都需要关闭&lt;w:r&gt; 并创建一个新标签,我已经对其进行了测试,效果非常好。

class printClass {
    private static $currentStyles = array();    

    public function __construct() {}

    public function format($string) {
            if($string !=""){
            return preg_replace_callback("#<b>|<u>|<i>|</b>|</u>|</i>#",
                                        'printClass::replaceTags',
                                        $string);
        }else{
            return false;
        }
    }


    private static function applyStyles() {

        if(count(self::$currentStyles) > 0 ) {

            foreach(self::$currentStyles as $value) {

                if($value == "b") {
                    $styles .= "<w:b/>";
                }   

                if($value == "u") {
                    $styles .= "<w:u w:val=\"single\"/>";
                }   

                if($value == "i") {
                    $styles .= "<w:i/>";
                }
            }

            return "<w:rPr>" . $styles . "</w:rPr>";
        }else{
            return false;
        }
    }



    private static function replaceTags($matches) {

        if($matches[0] == "<b>") {
            array_push(self::$currentStyles, "b");
        }   

        if($matches[0] == "<u>") {
            array_push(self::$currentStyles, "u");
        }   

        if($matches[0] == "<i>") {
            array_push(self::$currentStyles, "i");
        }

        if($matches[0] == "</b>") {
            self::$currentStyles = array_diff(self::$currentStyles, array("b"));
        }   

        if($matches[0] == "</u>") {
            self::$currentStyles = array_diff(self::$currentStyles, array("u"));
        }   

        if($matches[0] == "</i>") {
            self::$currentStyles = array_diff(self::$currentStyles, array("i"));
        }

        return "</w:t></w:r><w:r>" . self::applyStyles() . "<w:t xml:space=\"preserve\">";
    }
}

【讨论】:

    【解决方案3】:
    public function f_html2docx($currVal) {
    
        // handling <i> tag
    
        $el = 'i';
        $tag_open  = '<' . $el . '>';
        $tag_close = '</' . $el . '>';
        $nb = substr_count($currVal, $tag_open);
    
        if ( ($nb > 0) && ($nb == substr_count($currVal, $tag_open)) ) {
            $currVal= str_replace($tag_open,  '</w:t></w:r><w:r><w:rPr><w:i/></w:rPr><w:t>', $currVal);
            $currVal= str_replace($tag_close, '</w:t></w:r><w:r><w:t>', $currVal);
        }
    
        // handling <b> tag
    
        $el = 'b';
        $tag_open  = '<' . $el . '>';
        $tag_close = '</' . $el . '>';
        $nb = substr_count($currVal, $tag_open);
    
        if ( ($nb > 0) && ($nb == substr_count($currVal, $tag_open)) ) {
            $currVal= str_replace($tag_open,  '</w:t></w:r><w:r><w:rPr><w:b/></w:rPr><w:t>', $currVal);
            $currVal= str_replace($tag_close, '</w:t></w:r><w:r><w:t>', $currVal);
        }
    
        // handling <u> tag
    
        $el = 'u';
        $tag_open  = '<' . $el . '>';
        $tag_close = '</' . $el . '>';
        $nb = substr_count($currVal, $tag_open);
    
        if ( ($nb > 0) && ($nb == substr_count($currVal, $tag_open)) ) {
            $currVal= str_replace($tag_open,  '</w:t></w:r><w:r><w:rPr><w:u w:val="single"/></w:rPr><w:t>', $currVal);
            $currVal= str_replace($tag_close, '</w:t></w:r><w:r><w:t>', $currVal);
        }
    
        // handling <br> tag
    
        $el = 'br';
        $currVal= str_replace('<br />', '<w:br/>', $currVal);
    
        return $currVal;
    }
    
    public function f_handleUnsupportedTags($fieldValue){
        $fieldValue = strip_tags($fieldValue, '<b><i><u><br>');
    
        $fieldValue = str_replace('&nbsp;',' ',$fieldValue);
        $fieldValue = str_replace('<br>','<br />',$fieldValue);
    
        return $fieldValue;
    }
    

    现在调用这个函数:

    $fieldVal = $this->f_html2docx($this->f_handleUnsupportedTags($fieldVal));
    

    【讨论】:

    • 虽然此代码可以解决问题,including an explanation 说明如何以及为什么解决问题将真正有助于提高您的帖子质量,并可能导致更多的赞成票。请记住,您正在为将来的读者回答问题,而不仅仅是现在提问的人。请edit您的回答添加解释并说明适用的限制和假设。
    猜你喜欢
    • 2023-04-04
    • 1970-01-01
    • 2017-05-07
    • 1970-01-01
    • 2023-03-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多