【问题标题】:Is there a regular expession or similar simple approach for checking whether a MediaWiki PageTitle is valid?是否有常规的表达或类似的简单方法来检查 MediaWiki PageTitle 是否有效?
【发布时间】:2020-07-27 17:15:56
【问题描述】:

https://www.mediawiki.org/wiki/Manual:Page_title 为 MediaWiki pageTitle 可能不包含的内容陈述了很多条件。使用这种方法检查字符串是否是有效的 MediaWiki PageTitle 似乎并不容易。

什么是正则表达式或类似的简单方法来检查页面标题是否有效?

到目前为止,我能找到的最好的是一些 Java 代码(来自https://github.com/MER-C/wiki-java/blob/master/src/org/wikipedia/Wiki.java)。不过,我的目标语言是 python。

    /**
     *  Convenience method for normalizing MediaWiki titles. (Converts all
     *  underscores to spaces).
     *  @param s the string to normalize
     *  @return the normalized string
     *  @throws IllegalArgumentException if the title is invalid
     *  @throws IOException if a network error occurs (rare)
     *  @since 0.27
     */
    public String normalize(String s) throws IOException
    {
        // remove leading colon
        if (s.startsWith(":"))
            s = s.substring(1);
        if (s.isEmpty())
            return s;

        int ns = namespace(s);
        // localize namespace names
        if (ns != MAIN_NAMESPACE)
        {
            int colon = s.indexOf(":");
            s = namespaceIdentifier(ns) + s.substring(colon);
        }
        char[] temp = s.toCharArray();
        if (wgCapitalLinks)
        {
            // convert first character in the actual title to upper case
            if (ns == MAIN_NAMESPACE)
                temp[0] = Character.toUpperCase(temp[0]);
            else
            {
                int index = namespaceIdentifier(ns).length() + 1; // + 1 for colon
                temp[index] = Character.toUpperCase(temp[index]);
            }
        }

        for (int i = 0; i < temp.length; i++)
        {
            switch (temp[i])
            {
                // illegal characters
                case '{':
                case '}':
                case '<':
                case '>':
                case '[':
                case ']':
                case '|':
                    throw new IllegalArgumentException(s + " is an illegal title");
                case '_':
                    temp[i] = ' ';
                    break;
            }
        }
        // https://www.mediawiki.org/wiki/Unicode_normalization_considerations
        String temp2 = new String(temp).trim().replaceAll("\\s+", " ");
        return Normalizer.normalize(temp2, Normalizer.Form.NFC);
    }

【问题讨论】:

  • 这可能是相当绝望的。其中很多也取决于每个 wiki 的配置。例如,您必须知道 wiki 有哪些命名空间。
  • @Tgr - 感谢您对此进行调查。这适用于可以完全访问后端 wiki 的前端,检查的主要原因是避免可能由伪装成页面标题的代码引起的漏洞。排除非法字符是否完全减轻了这种风险?
  • 如果您的意思是 XSS 攻击,禁止 &lt;&gt; 可能是个好主意。在某些情况下,&amp;、引号或空格可以用作攻击向量,但这些都是有效的标题字符;您必须确保标题在使用时正确转义。如果你正在编写运行在 wiki 上的 JS 代码,mw.Html 有一堆转义函数。此外,mw.Title 有一些相当复杂的验证(与后端逻辑不是 100% 等效,但很接近)。
  • 由于我对后端 wiki 有一些控制权,我可以修改页面标题的“合法性”,例如禁止引号和 & 我也有点害怕转义字符,比如 % - 所以它也取决于 url 解码过程发生的状态。

标签: java python mediawiki


【解决方案1】:

如果您可以调用目标 wiki API 来进行标准化,那么这是一个标准化页面标题的 API 调用示例:

规范化的标题将在/query/normalized/0/to 中。您可以一次发送多个标题进行标准化,并用| 分隔它们。

示例取自https://www.mediawiki.org/wiki/API:Query#Example_2:_Title_normalization。

【讨论】:

    猜你喜欢
    • 2011-12-19
    • 1970-01-01
    • 2019-05-13
    • 2011-01-16
    • 1970-01-01
    • 2021-06-27
    • 1970-01-01
    • 2018-02-09
    • 2010-10-25
    相关资源
    最近更新 更多