【问题标题】:What is the proper Stormcrawler settings to capture a meta tag into an index?将元标记捕获到索引中的正确 Stormcrawler 设置是什么?
【发布时间】:2019-06-11 21:00:47
【问题描述】:

更新:我想通了。见底部...但如果我遗漏了什么,请随时纠正我...

对于来自以下元标记的信息,crawler-conf.yaml(以及其他地方,如果需要)中的正确设置是什么:

<meta name="college" content="artdesign"/>

要正确捕获到字段名称为“college”或“seed”的索引中吗?

我看到以下设置可能需要设置,但尝试了各种变体,似乎没有捕获数据。

在crawler-conf.yaml:

# lists the metadata to persist to storage
  # these are not transfered to the outlinks
  metadata.persist:
   - _redirTo
   - error.cause
   - error.source
   - isSitemap
   - isFeed
   - college
   - seed

不确定“持久存储”是否意味着索引?

crawler-conf.yaml 中的另一个选项是:

# configuration for the classes extending AbstractIndexerBolt
  indexer.md.mapping:
  - parse.title=title
  - parse.keywords=keywords
  - parse.description=description
  - domain=domain
  - college=college
  - college=seed

我之前曾问过这样一个事实,即“种子”的某些值似乎正在传播到获取的没有元标记的文档。该设置是:

  # metadata to transfer to the outlinks
  # used by Fetcher for redirections, sitemapparser, etc...
  # these are also persisted for the parent document (see below)
  # metadata.transfer:
  # - seed

因此,正如标题中所问的,我的问题是如何在crawler-conf.yaml(或任何其他配置)中配置这些选项,以可靠地从该问题顶部列出的元标记中捕获数据,而无需将其传播到获取的没有该元标记的文档?

【问题讨论】:

    标签: elasticsearch stormcrawler


    【解决方案1】:

    这是我整理出来的。上面引用代码中的“parse.title”中引用的“解析”是对@987654321 中顶级类下的自定义条目的引用(编辑:元标记的键,然后由其检索) @ 文件。我进去了,加了一个

    "parse.college": "//META[@name=\"college\"]/@content"

    在那些存在但仍在顶级类中的那些下方。

    然后,我将indexer.md.mapping 下对大学的引用更改为- parse.college=college,并重建了爬虫并运行了它。然后它开始正确抓取&lt;meta name="college" content="artdesign"/&gt; 标签并将其发送到索引中的college 字段。

    【讨论】:

    • 准确来说 parse.title 是对元数据对象中键的引用。您配置的 parsefilter 提取数据并将其放入该键下的元数据中
    猜你喜欢
    • 2011-06-23
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-04-20
    • 2010-10-30
    • 2017-06-08
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多