【发布时间】:2020-07-27 17:15:56
【问题描述】:
https://www.mediawiki.org/wiki/Manual:Page_title 为 MediaWiki pageTitle 可能不包含的内容陈述了很多条件。使用这种方法检查字符串是否是有效的 MediaWiki PageTitle 似乎并不容易。
什么是正则表达式或类似的简单方法来检查页面标题是否有效?
到目前为止,我能找到的最好的是一些 Java 代码(来自https://github.com/MER-C/wiki-java/blob/master/src/org/wikipedia/Wiki.java)。不过,我的目标语言是 python。
/**
* Convenience method for normalizing MediaWiki titles. (Converts all
* underscores to spaces).
* @param s the string to normalize
* @return the normalized string
* @throws IllegalArgumentException if the title is invalid
* @throws IOException if a network error occurs (rare)
* @since 0.27
*/
public String normalize(String s) throws IOException
{
// remove leading colon
if (s.startsWith(":"))
s = s.substring(1);
if (s.isEmpty())
return s;
int ns = namespace(s);
// localize namespace names
if (ns != MAIN_NAMESPACE)
{
int colon = s.indexOf(":");
s = namespaceIdentifier(ns) + s.substring(colon);
}
char[] temp = s.toCharArray();
if (wgCapitalLinks)
{
// convert first character in the actual title to upper case
if (ns == MAIN_NAMESPACE)
temp[0] = Character.toUpperCase(temp[0]);
else
{
int index = namespaceIdentifier(ns).length() + 1; // + 1 for colon
temp[index] = Character.toUpperCase(temp[index]);
}
}
for (int i = 0; i < temp.length; i++)
{
switch (temp[i])
{
// illegal characters
case '{':
case '}':
case '<':
case '>':
case '[':
case ']':
case '|':
throw new IllegalArgumentException(s + " is an illegal title");
case '_':
temp[i] = ' ';
break;
}
}
// https://www.mediawiki.org/wiki/Unicode_normalization_considerations
String temp2 = new String(temp).trim().replaceAll("\\s+", " ");
return Normalizer.normalize(temp2, Normalizer.Form.NFC);
}
【问题讨论】:
-
这可能是相当绝望的。其中很多也取决于每个 wiki 的配置。例如,您必须知道 wiki 有哪些命名空间。
-
@Tgr - 感谢您对此进行调查。这适用于可以完全访问后端 wiki 的前端,检查的主要原因是避免可能由伪装成页面标题的代码引起的漏洞。排除非法字符是否完全减轻了这种风险?
-
如果您的意思是 XSS 攻击,禁止
<>可能是个好主意。在某些情况下,&、引号或空格可以用作攻击向量,但这些都是有效的标题字符;您必须确保标题在使用时正确转义。如果你正在编写运行在 wiki 上的 JS 代码,mw.Html有一堆转义函数。此外,mw.Title有一些相当复杂的验证(与后端逻辑不是 100% 等效,但很接近)。 -
由于我对后端 wiki 有一些控制权,我可以修改页面标题的“合法性”,例如禁止引号和 & 我也有点害怕转义字符,比如 % - 所以它也取决于 url 解码过程发生的状态。