【问题标题】:html2pdf : Prevent the crawl of the google robothtml2pdf : 防止谷歌机器人抓取
【发布时间】:2021-03-08 11:33:55
【问题描述】:

我使用以下脚本:

https://www.html2pdf.fr/en/home

此脚本将我的 php 文件转换为 pdf 文件。

url 示例:mywebsite.com/pdf/url.php?id=8 将生成一个 PDF 文件。 另一个例子:https://github.com/spipu/html2pdf/blob/master/examples/example01.php

我不希望谷歌机器人将这些页面编入索引。

我在我的 htaccess 文件中添加了以下代码,但它不会阻止 google 抓取页面,因为它是在 PHP 中的: #Word和PDF文件的块索引 Header Set X-Robots-Tag "noindex, nofollow

我屏蔽不了怎么办?

【问题讨论】:

  • 在 PDF 的链接中,放一个 rel no follow:stackoverflow.com/a/2509022/231316
  • 谢谢,你是对的。然而,这些页面已经被 Google 机器人编入索引。如何取消索引?
  • 对于现有内容,请登录 Google Search Console 和 request to have the URLs removed。此外,更新您的 htaccess 以包含动态 PDF 生成器的 URL 规则。 htaccess 不在乎它是静态文件还是动态文件,它只是 URL 模式。

标签: php html-to-pdf


【解决方案1】:

您可以在生成 PDF 文件的脚本中添加 X-Robots-Tag HTTP 响应标头。

示例:header("X-Robots-Tag: noindex, nofollow", true);

Reference.

【讨论】:

    猜你喜欢
    • 2015-04-20
    • 2016-09-09
    • 1970-01-01
    • 2020-03-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多