【问题标题】:Perl- How do I insert a space before each capital letter except for the first occurrence or existing?Perl-我如何在每个大写字母之前插入一个空格,除了第一次出现或现有的?
【发布时间】:2015-04-22 21:36:54
【问题描述】:

我有一个类似的字符串:

 SomeCamel WasEnteringText

我已经找到了使用 php str_replace 拆分字符串和插入空格的各种方法,但是我在 perl 中需要它。

有时字符串前可能有一个空格,有时没有。有时字符串中会有空格,但有时没有。

我试过了:

    my $camel = "SomeCamel WasEnteringText";
    #or
    my $camel = " SomeCamel WasEntering Text";
    $camel =~ s/^[A-Z]/\s[A-Z]/g;
    #and
    $camel =~ s/([\w']+)/\u$1/g;

还有更多 =~s//g 的组合;经过大量阅读。

我需要一位大师来引导这头骆驼走向答案的绿洲。

好的,根据下面的输入,我现在有了:

$camel =~ s/([A-Z])/ $1/g;
$camel =~ s/^ //; # Strip out starting whitespace
$camel =~ s/([^[:space:]]+)/\u$1/g;

它完成了它,但似乎过分了。虽然有效。

【问题讨论】:

  • P.S.我希望结果是“Some Camel Was Entering Text”

标签: perl space


【解决方案1】:
s/(?<!^)[A-Z][a-z]*+(?!\s+)\K/ /g;

还有少一些“搞砸了”的版本:

s/
 (?<!^)          #Something not following the start of line,
    [A-Z][a-z]*+ #That starts with a capital letter and is followed by
                 #Zero or more lowercased letters, not giving anything back,
 (?!\s+)          #Not followed by one or more spaces,
\K               #Better explained here [1]
/ /gx;            #"Replace" it with a space.

编辑:我注意到,当您在混音中添加标点符号时,这也会增加额外的空格,这可能不是 OP 想要的;幸运的是,修复只是将负面展望从 \s+ 更改为 \W+。虽然现在我开始想知道为什么我实际上添加了这些优点。臭屁,我!

EDIT2:Erm,抱歉,最初忘记了 /g 标志。

EDIT3:好的,有人反对我。我变得迟钝了。不需要对 ^ 进行负面的回顾 - 我真的把球放在了这个上。希望修复:

s/[A-Z][a-z]*+(?!\W)\K/ /gx;

1:http://perldoc.perl.org/perlre.html

【讨论】:

  • 我试过 $camel =~ s/(?
  • 奇数。你运行的是哪个版本的 Perl?我在gskinner.com/RegExr、Windows 中的 5.12.1 和 Linux 中的 5.13.7 中尝试过,但它可能在某些版本中出现故障?如果我记得,\K 是 5.10 的补充,是吗?
  • 这应该.. 可能.. 可以在 5.10 之前的版本上运行。虽然无法测试。 s/([A-Z][a-z]*+(?!\W))/$1 /gx;
  • 正则表达式中的嵌套量词;由 标记
  • 它从来没有到达\K,那么。 +(和 ?+ 和 ++)是所有格限定词——它们吃,但从不回溯。 5.10 也增加了,我想。编辑:对于它的价值, s/([A-Z][a-z])(?=[A-Z]+)/$1 /g;应该在 5.8 上工作。该死的,老Perl! :)
【解决方案2】:

试试:

$camel =~ s/(?<! )([A-Z])/ $1/g; # Search for "(?<!pattern)" in perldoc perlre 
$camel =~ s/^ (?=[A-Z])//; # Strip out extra starting whitespace followed by A-Z

请注意$camel =~ s/([^ ])([A-Z])/$1 $2/g;的明显尝试有一个错误:如果有大写字母一个接一个,它就不起作用(例如“ABCD”将转换为“ABCD”而不是“A B C D”)

【讨论】:

  • @BipedalShark - 误读了原始 Q... 使用零宽度负后视断言进行了调整。
  • 应该改为第二个匹配项?
  • 这些天我本能地厌恶看到[A-Z] 的模式。除非是 RFC,否则\p{Lu} 更正确,并且只需输入一个字符。
  • @tchrist - 给我上色 - 害怕改变的老屁,但我发现 A-Z 比 \p{Lu} 更可读
【解决方案3】:

尝试: s/(?

这会在小写字符之后(即不是空格或字符串开头)和大写字符之前插入空格。

【讨论】:

  • 当我结合两个时: $camel =~ s/(?
  • \p{Lu} 是大写字母; \p{Ll} 是小写字母。
【解决方案4】:

改进...

...在Hughmeir 上,这也适用于以小写字母开头的数字和单词。

s/[a-z0-9]+(?=[A-Z])\K/ /gx

测试

 myBrainIsBleeding     => my_Brain_Is_Bleeding
 MyBrainIsBleeding     => My_Brain_Is_Bleeding
 myBRAInIsBLLEding     => my_BRAIn_Is_BLLEding
 MYBrainIsB0leeding    => MYBrain_Is_B0leeding
 0My0BrainIs0Bleeding0 => 0_My0_Brain_Is0_Bleeding0

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多