【问题标题】:atom feed: script to combine multiple <author> items into one? [closed]atom feed:将多个 <author> 项目合并为一个的脚本? [关闭]
【发布时间】:2021-03-05 12:42:38
【问题描述】:

我想编写一个命令行脚本,它将来自 atom 提要的多个 &lt;author&gt; 标签合并为一个。例如,这样的条目:

<entry>
    <id>someid</id>
    <published>somedate</published>
    <title>Title</title>
    <summary>Summary</summary>
    <author>
      <name>Author One</name>
    </author>
    <author>
      <name>Author Two</name>
    </author>
    <author>
      <name>Author Three</name>
    </author>
  </entry>

应该变成:

<entry>
    <id>someid</id>
    <published>somedate</published>
    <title>Title</title>
    <summary>Summary</summary>
    <author>
      <name>Author One, Author Two, Author Three</name>
    </author>
  </entry>

我想我可以自己使用 Perl 和正则表达式来完成,但是,由于使用正则表达式解析 XML 不是一个好主意,我会感谢使用适当的 xml 解析器的更优雅的解决方案。

【问题讨论】:

    标签: python php perl atom-feed


    【解决方案1】:

    Ted 的想法是正确的,但有些事情以比需要的更复杂的方式完成,而且他们不知道 Atom 格式的属性(例如,它使用 namespaces)。

    use XML::LibXML               qw( );
    use XML::LibXML::XPathContext qw( );
    
    my $xpc = XML::LibXML::XPathContext->new();
    $xpc->registerNs(a => 'http://www.w3.org/2005/Atom');
    
    # See XML::LibXML::Parser for more ways to create the document object.
    my $doc = XML::LibXML->load_xml( location => 'atom.xml' );
    
    for my $entry_node ($xpc->findnodes('/a:feed/a:entry', $doc)) {
       my @author_names;
       for my $author_node ($xpc->findnodes('a:author', $entry_node)) {
          push @author_names, $xpc->findvalue('a:name', $author_node);
          $author_node->unbindNode();
       }
    
       my $author_node = XML::LibXML::Element->new('author');
       my $name = $author_node->appendTextChild('name', join(", ", @author_names));
       $entry_node->appendChild($author_node);
    }
    
    $doc->toFile('atom.new.xml');
    

    【讨论】:

    • 酷,是的,我一个小时前才读到它 :-)
    • @Ted Lyngmo 注意使用unbindNode 删除节点。以及如何更简单地获取作者节点而不是名称节点
    • 我一定会看看的。我是 XML::LibXML 和 Xpath 的新手(我以前只用过一次,这里是为了回答)。
    【解决方案2】:

    在 Perl 中,我建议使用 XML::LibXML。

    在这里,我使用Xpath 查询来查找name 节点,然后将所有名称推送到一个数组中,同时删除author 节点。最后,我创建了一个新的 author 节点并被追加。

    #!/usr/bin/perl
    
    use strict;
    use warnings;
    
    use XML::LibXML;
    
    # example loading the xml from a file
    my $dom = XML::LibXML->load_xml(location => 'atom.xml', no_blanks => 1);
    my $root = $dom->documentElement();
    
    # the Xpath query
    my $query = q{
        /entry/author/name
    };
    
    my @authornames;
    
    foreach my $namenode ($dom->findnodes($query)) {
        # save the name
        push @authornames, $namenode->to_literal();
    
        # remove the author node
        $namenode->getParentNode->getParentNode->removeChild($namenode->getParentNode);
    
        #or:
        # $root->removeChild($namenode->getParentNode);
    }
    
    # build a new author node
    my $author = XML::LibXML::Element->new('author');
    $author->appendTextChild('name', join(", ",@authornames));
    
    # and add it
    $root->appendChild($author);
    
    # print the result
    print $dom->serialize(1);
    
    #or, if you don't want the <?xml...> header:
    # print $root->serialize(1) . "\n";
    

    输出:

    <?xml version="1.0"?>
    <entry>
      <id>someid</id>
      <published>somedate</published>
      <title>Title</title>
      <summary>Summary</summary>
      <author>
        <name>Author One, Author Two, Author Three</name>
      </author>
    </entry>
    

    【讨论】:

    • 我会选择池上的简短答案,但无论如何感谢您的解决方案!
    • @n_flanders 不客气 - 我也会接受这个答案。
    猜你喜欢
    • 2013-01-23
    • 1970-01-01
    • 2011-06-14
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多