【问题标题】:Symfony2 DomCrawler and FB2 book format parserSymfony2 DomCrawler 和 FB2 书籍格式解析器
【发布时间】:2013-11-14 13:53:03
【问题描述】:

全部!

如何使用 Symfony2 DomCrawler 组件正确解析描述的 XML 文件?

我需要拆分所有部分,并与当前部分一起收集一个内部标签(碑文、p、诗等),它只属于这个部分。

我有如下描述的标准 FB2 书籍 XML 格式:

<?xml version="1.0" encoding="utf-8"?>
<FictionBook xmlns="http://www.gribuser.ru/xml/fictionbook/2.0" xmlns:l="http://www.w3.org/1999/xlink">
<description></description>
<body>
<section>
    <title><p><strong>Level 1, section 1</strong></p></title>
    <section>
        <title><p><strong>Level 2, section 2</strong></p></title>
        <section>
            <title><p><strong>Level 3, section 3</strong></p></title>
            <p>Level 3, section 3, paragraph 1</p>
            <poem>
                <stanza>
                    <v>bla-bla-bla 1</v>
                    <v>bla-bla-bla 2</v>
                    <v>bla-bla-bla 3</v>
                </stanza>
            </poem>
            <p>Level3, section 3, paragraph 2</p>
            <subtitle><strong>x x x</strong></subtitle>
        </section>
        <section>
            <title><p><strong>Level 3, section 4</strong></p></title>
            <p>Level 3, section 4, paragraph 1</p>
            <p>Level 3, section 4, paragraph 2</p>
            <subtitle><strong>x x x</strong></subtitle>
        </section>
        <section>
            <title><p><strong>Level 3, section 5</strong></p></title>
            <p>Level 3, section 5, paragraph 1</p>
            <p>Level 3, section 5, paragraph 2</p>
            <p>Level 3, section 5, paragraph 3</p>
            <empty-line/>
            <subtitle>This file was created</subtitle>
            <subtitle>with BookDesigner program</subtitle>
            <subtitle>bookdesigner@the-ebook.org</subtitle>
            <subtitle>22.04.2004</subtitle>
        </section>
    </section>
</section>
</body>
</FictionBook>

下面的代码不起作用,有人可以帮我解决这个问题吗? 顺便说一句,标题解析正确......但部分的标签不......

private function loadBookSections(Crawler $crawler)
{
    $sections = $crawler->filter('section')->each(function(Crawler $node) {
        $c = $node->filter('section')->reduce(function(Crawler $node, $i) {
            return ($i == 0);
        });

        return array(
            'title' => $node->filter('title')->text(),
            'inner' => $c->html(),
        );
    });

    echo "*******************************************\n";

    foreach($sections as $section ) {
        echo ">>> ".$section['title']."\n";
        echo "!!! ".$section['inner']."\n";
    }
}

感谢您的帮助!

【问题讨论】:

  • 你可以在你的 xml 中使用内置的序列化器/反序列化器而不是 dom 爬虫吗? look here
  • 哟! (这个称呼让我笑了)

标签: php symfony xml-parsing fb2


【解决方案1】:

四天后...我通过 XPath 找到了解决方案...

private function loadBookSections(Crawler $crawler)
{

    $sections = $crawler->filter('section')->each(function(Crawler $node) {
        return array(
            'title' => $node->filter('title')->text(),
            'inner' => $node->filterXPath("//*[not(section)]")->html(),
        );
    });

    foreach($sections as $section) {
        echo "TITLE: ".$section['title']."\n";
        echo "INNER: ".$section['inner']."\n";
    }
}

【讨论】:

    【解决方案2】:

    如果你减少你的 XML 文件,你会得到这样的结果:

    <section>
        <section>
            <!-- ... -->
        </section>
        <section>
            <!-- ... -->
        </section>
        <section>
            <!-- ... -->
        </section>
    </section>
    

    你想捕捉子 section 元素,而不是父元素。

    目前您只遍历父 section 元素的列表,这意味着您只能获取父 section 元素的 HTML。

    要遍历孩子,您需要选择section section 而不是section。


    进一步改进代码的辅助信息:不要使用丑陋的reduce 调用,只需使用-&gt;first() 来获取节点列表的第一个元素。


    总的来说,您的代码将是:

    $sections = $crawler->filter('section section')->each(function(Crawler $node) {
        $c = $node->filter('section')->first();
    
        return array(
            'title' => $node->filter('title')->text(),
            'inner' => $c->html(),
        );
    });
    

    【讨论】:

    • 谢谢,Wouter J.,但是您的解决方案只是丢失了父数据的第一个条目...此外,如果我们将使用 3 个或更多级别的输入文档 - 子句 $c- >html() 返回属于这个根的所有子节...
    • 我认为在我的情况下,DomCrawler 应该包含 filter('section') 的相反方法(类似于 NOT filter() 或 discard(...)),它返回除属于部分标签之外的所有内容。 ..
    猜你喜欢
    • 2023-03-19
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多