【问题标题】:sed: replace spaces within quotes with underscoressed:用下划线替换引号内的空格
【发布时间】:2013-02-16 23:08:31
【问题描述】:

我的输入(例如,来自 OpenBSD 上的 ifconfig run0 scan)有一些由空格分隔的字段,但一些字段本身包含空格(幸运的是,这些包含空格的字段总是用引号括起来)。

我需要区分引号内的空格和分隔符空格。这个想法是用下划线替换引号内的空格。

样本数据:

%cat /tmp/ifconfig_scan | fgrep nwid | cut -f3
nwid Websense chan 6 bssid 00:22:7f:xx:xx:xx 59dB 54M short_preamble,short_slottime
nwid ZyXEL chan 8 bssid cc:5d:4e:xx:xx:xx 5dB 54M privacy,short_slottime
nwid "myTouch 4G Hotspot" chan 11 bssid d8:b3:77:xx:xx:xx 49dB 54M privacy,short_slottime

最终没有按照我想要的方式处理,因为我还没有用下划线替换引号内的空格:

%cat /tmp/ifconfig_scan | fgrep nwid | cut -f3 |\
    cut -s -d ' ' -f 2,4,6,7,8 | sort -n -k4
"myTouch Hotspot" 11 bssid d8:b3:77:xx:xx:xx
ZyXEL 8 cc:5d:4e:xx:xx:xx 5dB 54M
Websense 6 00:22:7f:xx:xx:xx 59dB 54M

【问题讨论】:

标签: sed awk tcsh openbsd


【解决方案1】:

对于sed-only 解决方案(我不一定提倡),请尝试:

echo 'a b "c d e" f g "h i"' |\
sed ':a;s/^\(\([^"]*"[^"]*"[^"]*\)*[^"]*"[^"]*\) /\1_/;ta'
a b "c_d_e" f g "h_i"

翻译:

  • 从行首开始。
  • 查找模式junk"junk",重复零次或多次,其中junk 没有引号,后跟junk"junk space。
  • 将最后的空格替换为_。
  • 如果成功,请跳回到开头。

【讨论】:

  • 它确实有效! :-) 即使在 OpenBSD 4.6 上有一个旧的 sed 还没有 -E 选项!但是为什么要转义括号呢? (虽然我尝试用( 替换\(,但它停止工作了。)另外,为什么你不必在第二个[] 中包含一个空格,例如不是"[^" ]*" 而不是"[^"]*"?它怎么知道不贪?除此之外,正则表达式本身非常有意义! :) 那么,:a 是标签a,而ta 是跳转到a?跳跃意味着倒回应用搜索/替换的行?漂亮!我必须把它放进我的武器库。 :-)
  • @cnst 替换向后工作。要查看各个步骤 (GNU sed),请将命令 l0 放在替换命令之后。即:a;s/.../.../;l0;ta
  • @potong,很好的其他选择,l0 在我的sed 中不起作用,但只是一个l,就像在;l;ta 中一样,似乎很好用,确实表明它是处理贪婪和倒退。在这种情况下,最好还是避免空间以使其不贪婪?
【解决方案2】:

试试这个:

awk -F'"' '{for(i=2;i<=NF;i++)if(i%2==0)gsub(" ","_",$i);}1' OFS="\"" file

它适用于一行中的多引号部分:

echo '"first part" foo "2nd part" bar "the 3rd part comes" baz'| awk -F'"' '{for(i=2;i<=NF;i++)if(i%2==0)gsub(" ","_",$i);}1' OFS="\"" 
"first_part" foo "2nd_part" bar "the_3rd_part_comes" baz

编辑替代形式:

awk 'BEGIN{FS=OFS="\""} {for(i=2;i<NF;i+=2)gsub(" ","_",$i)} 1' file

【讨论】:

  • 嗯,在我的 tcsh 中不起作用:cat /tmp/ifconfig_scan | fgrep nwid | cut -f3 | awk -F'"' '{for(i=2;i&lt;=NF;i++)if(i%2==0)gsub(" ","_",$i);}1' OFS="\"" | cut -s -d ' ' -f 2,4,6,7,8 | sort -n -k4 返回 Unmatched ".
  • 好的,这在 tcsh 中效果很好(只是将一些“更改为”):cat /tmp/ifconfig_scan | fgrep nwid | cut -f3 | awk -F'"' '{for(i=2;i&lt;=NF;i++)if(i%2==0)gsub(" ","_",$i);}1' OFS='"' | cut -s -d ' ' -f 2,4,6,7,8 | sort -n -k4
  • 我不同意awk/sed 适合这项任务,但这并不意味着它不能完成。如果您打算使用awk,您可以取消if 语句。只需使用i+=2 和i&lt;NF。
  • +1 用于该方法,但将 i++ 更改为 i+=2 并删除 if(i%2==0) 和 gsub() 之后的虚假尾随 ;。此外,如果您希望 FS 和 OFS 具有相同的值,最清楚的是在 BEGIN 部​​分中为它们分配相同的值作为BEGIN{FS=OFS="""}。
【解决方案3】:

另一个 awk 尝试:

awk '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=\"

去掉引号:

awk '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=

在@steve 完成的早期测试的基础上,使用三倍大小的测试文件进行了一些额外的测试。我必须稍微转换一下sed 语句,以便非GNU 的seds 也可以处理它。我包括awk (bwk) gawk3, gawk4 和 mawk:

$ for i in {1..1500000}; do echo 'a b "c d e" f g "h i" j k l "m n o "p q r" s t" u v "w x" y z' ; done > test
$ time perl -pe 's:"[^"]*":($x=$&)=~s/ /_/g;$x:ge' test >/dev/null

real    0m27.802s
user    0m27.588s
sys 0m0.177s
$ time awk 'BEGIN{FS=OFS="\""} {for(i=2;i<NF;i+=2)gsub(" ","_",$i)} 1' test >/dev/null

real    0m6.565s
user    0m6.500s
sys 0m0.059s
$ time gawk3 'BEGIN{FS=OFS="\""} {for(i=2;i<NF;i+=2)gsub(" ","_",$i)} 1' test >/dev/null

real    0m21.486s
user    0m18.326s
sys 0m2.658s
$ time gawk4 'BEGIN{FS=OFS="\""} {for(i=2;i<NF;i+=2)gsub(" ","_",$i)} 1' test >/dev/null

real    0m14.270s
user    0m14.173s
sys 0m0.083s
$ time mawk 'BEGIN{FS=OFS="\""} {for(i=2;i<NF;i+=2)gsub(" ","_",$i)} 1' test >/dev/null

real    0m4.251s
user    0m4.193s
sys 0m0.053s
$ time awk '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=\" test >/dev/null

real    0m13.229s
user    0m13.141s
sys 0m0.075s
$ time gawk3 '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=\" test >/dev/null

real    0m33.965s
user    0m26.822s
sys 0m7.108s
$ time gawk4 '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=\" test >/dev/null

real    0m15.437s
user    0m15.328s
sys 0m0.087s
$ time mawk '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=\" test >/dev/null

real    0m4.002s
user    0m3.948s
sys 0m0.051s
$ time sed -e :a -e 's/^\(\([^"]*"[^"]*"[^"]*\)*[^"]*"[^"]*\) /\1_/;ta' test > /dev/null

real    5m14.008s
user    5m13.082s
sys 0m0.580s
$ time gsed -e :a -e 's/^\(\([^"]*"[^"]*"[^"]*\)*[^"]*"[^"]*\) /\1_/;ta' test > /dev/null

real    4m11.026s
user    4m10.318s
sys 0m0.463s

mawk 呈现最快的结果...

【讨论】:

  • 不错的一个!两者都工作得很好,并且似乎是问题的最短解决方案,甚至比@Steve 的最短perl sn-p 更短(尽管在那方面不太可读)。我需要放弃sed,学习awk!
  • 在@steve 的测试中加入了一些额外的测试。
【解决方案4】:

您最好使用perl。代码更具可读性和可维护性:

perl -pe 's:"[^"]*":($x=$&)=~s/ /_/g;$x:ge'

根据您的输入,结果是:

a b "c_d_e" f g "h_i"

解释:

-p            # enable printing
-e            # the following expression...

s             # begin a substitution

:             # the first substitution delimiter

"[^"]*"      # match a double quote followed by anything not a double quote any
              # number of times followed by a double quote

:             # the second substitution delimiter

($x=$&)=~s/ /_/g;      # copy the pattern match ($&) into a variable ($x), then 
                       # substitute a space for an underscore globally on $x. The
                       # variable $x is needed because capture groups and
                       # patterns are read only variables.

$x            # return $x as the replacement.

:             # the last delimiter

g             # perform the nested substitution globally
e             # make sure that the replacement is handled as an expression

一些测试:

for i in {1..500000}; do echo 'a b "c d e" f g "h i" j k l "m n o "p q r" s t" u v "w x" y z' >> test; done

time perl -pe 's:"[^"]*":($x=$&)=~s/ /_/g;$x:ge' test >/dev/null

real    0m8.301s
user    0m8.273s
sys     0m0.020s

time awk 'BEGIN{FS=OFS="\""} {for(i=2;i<NF;i+=2)gsub(" ","_",$i)} 1' test >/dev/null

real    0m4.967s
user    0m4.924s
sys     0m0.036s

time awk '!(NR%2){gsub(FS,"_")}1' RS=\" ORS=\" test >/dev/null

real    0m4.336s
user    0m4.244s
sys     0m0.056s

time sed ':a;s/^\(\([^"]*"[^"]*"[^"]*\)*[^"]*"[^"]*\) /\1_/;ta' test >/dev/null

real    2m26.101s
user    2m25.925s
sys     0m0.100s

【讨论】:

  • 我很抱歉,但我不同意该代码比任何东西都更具可读性。其他人显然会不同意,但至少我完全无法理解,我诚实地试图弄清楚。你介意添加一个关于它在做什么的解释吗?
  • @EdMorton:没问题。很高兴我能帮上忙。 1) =~ 仅表示“针对此正则表达式运行此变量”。 2) Perl 的e 标志,就像seds 的e 标志。在父代替换中,替换值是第二个(子代)替换。默认情况下,Perl 不期望这样。所以需要e 标志。
  • @EdMorton: 3) 我的意思是,让父代成为 $x。
  • @EdMorton:如果您有兴趣,请发布一些有趣的时间安排。不过,我想我还是更喜欢perl。我认为它更好地描述了实际发生的事情。但在时间紧迫的管道中,在看到这些结果之后,我会权衡可读性并选择awk。
  • 好的,我编译了 mawk、gawk3、gawk4 和 GNU sed 并将它们添加到我的系统中,然后运行了一些进一步的测试。我将结果添加到帖子的末尾。
【解决方案5】:

不是答案,只是发布 @steve 的 perl 代码的 awk 等效代码以防万一有人感兴趣(并帮助我将来记住这一点):

@steve 发布:

perl -pe 's:"[^\"]*":($x=$&)=~s/ /_/g;$x:ge'

从阅读@steve 的解释来看,与该 perl 代码等效的最简短的 awk(不是首选的 awk 解决方案 - 请参阅@Kent 的答案)将是 GNU awk:

gawk '{
   head = ""
   while ( match($0,"\"[^\"]*\"") ) {
      head = head substr($0,1,RSTART-1) gensub(/ /,"_","g",substr($0,RSTART,RLENGTH))
      $0 = substr($0,RSTART+RLENGTH)
   }
   print head $0
}'

我们从具有更多变量的 POSIX awk 解决方案开始:

awk '{
   head = ""
   tail = $0
   while ( match(tail,"\"[^\"]*\"") ) {
      x = substr(tail,RSTART,RLENGTH)
      gsub(/ /,"_",x)
      head = head substr(tail,1,RSTART-1) x
      tail = substr(tail,RSTART+RLENGTH)
   }
   print head tail
}'

并使用 GNU awk 的 gensub() 保存一行:

gawk '{
   head = ""
   tail = $0
   while ( match(tail,"\"[^\"]*\"") ) {
      x = gensub(/ /,"_","g",substr(tail,RSTART,RLENGTH))
      head = head substr(tail,1,RSTART-1) x
      tail = substr(tail,RSTART+RLENGTH)
   }
   print head tail
}'

然后去掉变量x:

gawk '{
   head = ""
   tail = $0
   while ( match(tail,"\"[^\"]*\"") ) {
      head = head substr(tail,1,RSTART-1) gensub(/ /,"_","g",substr(tail,RSTART,RLENGTH))
      tail = substr(tail,RSTART+RLENGTH)
   }
   print head tail
}'

如果你不需要 $0、NF 等,然后在循环之后摆脱变量“tail”:

gawk '{
   head = ""
   while ( match($0,"\"[^\"]*\"") ) {
      head = head substr($0,1,RSTART-1) gensub(/ /,"_","g",substr($0,RSTART,RLENGTH))
      $0 = substr($0,RSTART+RLENGTH)
   }
   print head $0
}'

【讨论】:

    猜你喜欢
    • 2018-06-11
    • 1970-01-01
    • 2011-07-12
    • 1970-01-01
    • 2020-02-03
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多