【问题标题】:Extract word document comments and the text they comment on提取word文档评论和他们评论的文本
【发布时间】:2022-01-13 12:18:47
【问题描述】:

我需要提取 word 文档 cmets 和他们评论的文本。以下是我目前的解决方案,但它没有按预期工作

public class Main {

    public static void main(String[] args) throws Exception {
        var document = new Document("sample.docx");
        NodeCollection<Paragraph> paragraphs = document.getChildNodes(PARAGRAPH, true);
        List<MyComment> myComments = new ArrayList<>();

        for (Paragraph paragraph : paragraphs) {
            var comments = getComments(paragraph);
            int commentIndex = 0;

            if (comments.isEmpty()) continue;

            for (Run run : paragraph.getRuns()) {
                var runText = run.getText();

                for (int i = commentIndex; i < comments.size(); i++) {
                    Comment comment = comments.get(i);
                    String commentText = comment.getText();

                    if (paragraph.getText().contains(runText + commentText)) {
                        myComments.add(new MyComment(runText, commentText));
                        commentIndex++;
                        break;
                    }
                }
            }
        }

        myComments.forEach(System.out::println);
    }

    private static List<Comment> getComments(Paragraph paragraph) {
        @SuppressWarnings("unchecked")
        NodeCollection<Comment> comments = paragraph.getChildNodes(COMMENT, false);
        List<Comment> commentList = new ArrayList<>();

        comments.forEach(commentList::add);

        return commentList;
    }

    static class MyComment {
        String text;
        String commentText;

        public MyComment(String text, String commentText) {
            this.text = text;
            this.commentText = commentText;
        }

        @Override
        public String toString() {
            return text + "-->" + commentText;
        }
    }
}

sample.docx 内容为:

输出是(不正确的):

factors-->This is word comment
%–10% of cancers are caused by inherited genetic defects from a person's parents.-->Second paragraph comment

预期输出是:

factors-->This is word comment
These factors act, at least partly, by changing the genes of a cell. Typically, many genetic changes are required before cancer develops. Approximately 5%–10% of cancers are caused by inherited genetic defects from a person's parents.-->Second paragraph comment
These factors act, at least partly, by changing the genes of a cell. Typically, many genetic changes are required before cancer develops. Approximately 5%–10% of cancers are caused by inherited genetic defects from a person's parents.-->First paragraph comment

请帮助我用一种更好的方法来提取 word 文档 cmets 和他们评论的文本。如果您需要更多详细信息,请告诉我,我将提供所有必需的详细信息

【问题讨论】:

    标签: java aspose aspose.words


    【解决方案1】:

    注释文本由特殊节点CommentRangeStart 和CommentRangeEnd 标记。 CommentRangeStart 和 CommentRangeEnd 节点具有 Id,它对应于范围链接到的 Comment id。所以你需要在对应的开始节点和结束节点之间提取内容。 顺便说一句,Aspose.Words API 参考中的代码示例显示了如何使用文档访问者打印所有 cmets 的内容及其评论范围。看起来正是您正在寻找的东西。

    编辑:您可以使用如下代码来完成您的任务。我没有提供在节点之间提取内容的完整代码,在GitHub 上可用

    Document doc = new Document("C:\\Temp\\in.docx");
    
    // Get the comments in the document.
    Iterable<Comment> comments = doc.getChildNodes(NodeType.COMMENT, true);
    Iterable<CommentRangeStart> commentRangeStarts = doc.getChildNodes(NodeType.COMMENT_RANGE_START, true);
    Iterable<CommentRangeEnd> commentRangeEnds = doc.getChildNodes(NodeType.COMMENT_RANGE_END, true);
    
    for (Comment c : comments)
    {
        System.out.println(String.format("Comment %d : %s", c.getId(), c.toString(SaveFormat.TEXT)));
    
        CommentRangeStart start = null;
        CommentRangeEnd end = null;
    
        // Search for an appropriate start and end.
        for (CommentRangeStart s : commentRangeStarts)
        {
            if (c.getId() == s.getId())
            {
                start = s;
                break;
            }
        }
    
        for (CommentRangeEnd e : commentRangeEnds)
        {
            if (c.getId() == e.getId())
            {
                end = e;
                break;
            }
        }
    
        if (start != null && end != null)
        {
            // Extract content between the start and end nodes.
            // Code example how to extract content between nodes is here
            // https://github.com/aspose-words/Aspose.Words-for-Java/blob/master/Examples/src/main/java/com/aspose/words/examples/programming_documents/document/ExtractContentBetweenCommentRange.java
        }
        else
        {
            System.out.println(String.format("Comment %d Does not have comment range"));
        }
    
    }
    

    【讨论】:

    • 如果您尝试这些示例,您只会得到评论文本,而不是他们评论的文本。请找到这些示例的输出 [Comment range start] ID: 0 | [评论开始] 对于评论范围 ID 0,作者未知,于 27/10/21,下午 7:58 | | [运行]“这是文字评论” | [评论结束] [评论范围结束] ID:0
    • 我已经编辑了我的答案。
    • 代码正是这样做的,我只是没有在 CommentRangeStart 和 CommentRangeEnd 节点之间放置提取内容的实现。这段代码可以在这里找到github.com/aspose-words/Aspose.Words-for-Java/blob/master/…见代码中的cmets。
    • 好的,谢谢!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-12-11
    • 2018-09-22
    • 2019-09-21
    • 2011-07-05
    • 2014-09-10
    • 1970-01-01
    相关资源
    最近更新 更多