【问题标题】:Removing style from a string retrieved from WordDocument with Open XML Office SDK使用 Open XML Office SDK 从从 WordDocument 检索的字符串中删除样式
【发布时间】:2012-06-25 10:31:53
【问题描述】:

我正在使用 Open XML Office SDK 2.0 在 word 文档中搜索字符串并列出这些字符串。

    MatchCollection Matches;
    using (WordprocessingDocument wordDoc = WordprocessingDocument.Open(txtLocation.Text, true))
    {
        string docText = null;
        using (StreamReader sr = new StreamReader(wordDoc.MainDocumentPart.GetStream()))
        {
            docText = sr.ReadToEnd();
        }
        Regex regex = new   Regex(@"\(.*?\)");
        Matches = regex.Matches(docText);
    }
    int i = 0;
    while (i < Matches.Count)
    {    Label lb = new Label();
         lb.Text = Matches[i].ToString();
         lb.Location = new System.Drawing.Point(24, (28 + i * 24));
         this.panel1.Controls.Add(lb);
         i++;
     }

问题是有时它返回正确的字符串,例如:(HelloWorld),但有时它与标签完全不同,例如:

我该如何摆脱这些?

【问题讨论】:

    标签: c# regex string search ms-word


    【解决方案1】:

    找出我必须做的,将字符串运行到​​另一个 Regex.Replace。 这个替换了所有 标签(所以 XML/HTML)

    String str = Matches[i].ToString();
    str = Regex.Replace(str, @"<(.|\n)*?>", "");
    lb.Text  = str;
    

    【讨论】:

    • 是的,这是另一种方法,虽然它可能会使用更多的处理器时间。
    【解决方案2】:

    大概所有的格式化标签都是 XML 风格的(在尖括号之间)。在这种情况下,您可以使用 String.StartsWith 和 String.EndsWith 方法判断字符串是否为 XML 标记:

    // ...
    while (i < Matches.Count)
    {
         String str = Matches[i].ToString();
         if (!(str.StartsWith("<") && str.EndsWith(">"))) {
             // ...
         }
         i++;
    }
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2015-01-13
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-05-07
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多