【问题标题】:Rendering text with HTML tags to Formatted text in a Word table using VBA使用 VBA 将带有 HTML 标记的文本呈现为 Word 表格中的格式化文本
【发布时间】:2019-02-20 23:00:31
【问题描述】:

我有一个带有 html 标签的 word 文档,我需要将其转换为格式化文本。例如,我希望 <strong>Hello</strong> 改为显示为 Hello。

我以前从未使用过 VBA,但我一直在尝试拼凑一些东西,让我可以从 Word 中的特定表格单元格复制 html 文本,使用 IE 显示该文本的格式化版本,复制来自 IE 的格式化文本,然后将其粘贴回同一个 Word 表格单元格。我想我已经能够弄清楚一些代码,但我认为我没有正确引用表格单元格。任何人都可以帮忙吗?这是我目前所拥有的:

Dim Ie As Object

Set Ie = CreateObject("InternetExplorer.Application")

With Ie
    .Visible = False

    .Navigate "about:blank"

    .Document.body.InnerHTML = ActiveDocument.Tables(1).Cell(2, 2)
    
    .Document.execCommand "SelectAll"
    
    .Document.execCommand "Copy"
    
    ActiveDocument.Paste Destination = ActiveDocument.Tables(1).Cell(2, 2)

    .Quit
End With
End Sub

【问题讨论】:

    标签: html vba internet-explorer ms-word formatted-text


    【解决方案1】:

    .cell(2,2) 的两种用途需要两种不同的方法。

    要从单元格中获取文本,您需要修改第一行以读取

    .Document.body.InnerHTML = ActiveDocument.Tables(1).Cell(2, 2).range.text  
    

    在第二种情况下,您的术语不正确。它应该是

    ActiveDocument.Tables(1).Cell(2, 2).range.paste
    

    您可以很容易地获得有关各个关键字/属性的帮助。在 VBA IDE 中,只需将光标放在关键字/属性上,然后按 F1。您将被带到关键字/属性的 MS 帮助页面。有时,当有多个备选方案时,您会有额外的选择步骤。

    您还应该知道属性 .cell(row,column) 很容易失败,因为它依赖于它们在表格中没有合并的单元格。更强大的方法是使用 .cells(index) 属性。

    您可能可以采用另一种方法并使用通配符搜索来查找标签,然后在应用合适的链接样式的同时替换您需要的部分(您将无法使用段落样式,因为您将尝试仅格式化段落的一部分并且字符样式似乎不适用于查找/替换)。

    下面是删除 HTML 标记并格式化剩余文本的此类代码示例

    Option Explicit
    
    Sub replaceHTML_WithFormattedText()
    
    ' a comma seperated list of HTML tags
    Const myTagsList                          As String = "strong,small,b,i,em"
    
    ' a list of linked styles chosen or designed for each tag
    ' Paragraph  styles cannot be used as we are replacing only part of a paragraph
    ' Character styles just don't seem to work
    ' The linked styles below were just chosen from the default Word styles as an example
    Const myStylesList                        As String = "Heading 1,Heading 9,Comment Subject,Intense Quote,Message Header"
    
    ' <, > and / are special characters therefore need escaping with '\' to get the actual character
    Const myFindTag                           As String = "(\<Tag\>)(*)(\<\/Tag\>)"
    Const myReplaceStr                        As String = "\2"
    
    Dim myTagsHTML()                        As String
    Dim myTagsStyles()                      As String
    Dim myIndex                             As Long
    
        myTagsHTML = Split(myTagsList, ",")
        myTagsStyles = Split(myStylesList, ",")
    
        If UBound(myTagsHTML) <> UBound(myTagsStyles) Then
            MsgBox "Different number of tags and Styles", vbOKOnly
            Exit Sub
    
        End If
    
        For myIndex = 0 To UBound(myTagsHTML)
    
            With ActiveDocument.StoryRanges(wdMainTextStory).Find
                .ClearFormatting
                .Format = True
                .Text = Replace(myFindTag, "Tag", Trim(myTagsHTML(myIndex)))
                .MatchWildcards = True
                .Replacement.Text = myReplaceStr
                .Replacement.Style = myTagsStyles(myIndex)
                .Execute Replace:=wdReplaceAll
    
            End With
    
        Next
    
    End Sub
    

    【讨论】:

    • 这非常有帮助!谢谢!我按照您的建议修改了这两行,目前效果很好,我正在考虑使用通配符选项。它似乎更可靠,因为它减少了对表格和 IE 的需求。
    • @steven-laycock «您还应该知道属性 .cell(row,column) 容易失败,因为它依赖于它们在表中没有合并单元格。»事实并非如此。当用作建议的 OP 时,它是非常可靠的。表格中的每个单元格都有一个 RowIndex 和一个 ColumnIndex;正确提供这些,不会有问题。
    • 我的经验是,如果一个表格不统一,因为它包含水平或垂直合并的单元格,那么当您访问合并的单元格时 .cell(x,y) 将失败,而 .cells(x) 永远不会失败。我试图提醒 OP 在使用 x,y 坐标扫描表时出现 VBA 错误的情况。
    • @modz 由于您是 Stack Overflow 的新手:在网站上单击您认为回答您的问题的贡献旁边的复选标记是合适的。这告诉其他有类似问题的人该贡献是最有帮助的;向寻找问题以回答的其他人表明该问题已“关闭”;奖励网站“点”给贡献者。在某些时候,您还可以对您认为有帮助的网站上的任何贡献进行投票,即使它不是“最佳”答案。您还可以对不属于您的投稿(其他人的问题和回复)进行投票(和否决)。
    【解决方案2】:

    尝试以下方式:

    Sub ReformatHTML()
    Application.ScreenUpdating = False
    With ActiveDocument.Range.Find
      .ClearFormatting
      .Format = True
      .Forward = True
      .MatchWildcards = True
      .Wrap = wdFindContinue
      .Replacement.Text = "\2"
      .Replacement.ClearFormatting
      .Text = "\<(u\>)(*)\</\1"
      .Replacement.Font.Underline = True
      .Execute Replace:=wdReplaceAll
      .Replacement.ClearFormatting
      .Text = "\<(b\>)(*)\</\1"
      .Replacement.Font.Bold = True
      .Execute Replace:=wdReplaceAll
      .Replacement.ClearFormatting
      .Text = "\<(i\>)(*)\</\1"
      .Replacement.Font.Italic = True
      .Execute Replace:=wdReplaceAll
      .Replacement.ClearFormatting
      .Text = "\<(h\>)(*)\</\1"
      .Replacement.Highlight = True
      .Execute Replace:=wdReplaceAll
    End With
    Application.ScreenUpdating = True
    End Sub
    

    上述宏使用“普通”HTML 代码进行粗体、斜体、下划线和突出显示。

    由于您的文档似乎使用了不同的约定(可能是样式名称?),例如,您可以将代码中的 (b>) 替换为 (strong>)。而且,如果它旨在与 Word 自己的“强”风格相关,您也可以更改:

    .Replacement.Font.Bold = True
    

    到:

    .Replacement.Style = "Strong"
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-03-06
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-08-29
      相关资源
      最近更新 更多