【发布时间】:2011-03-28 09:16:13
【问题描述】:
我正在使用
awk '{ printf "%s", $3 }'
从空格分隔的行中提取一些字段。当然,当该字段被引用时,我会得到部分结果。请问有人可以提出解决方案吗?
【问题讨论】:
-
显示您的输入文件格式..以及您想要的输出!
我正在使用
awk '{ printf "%s", $3 }'
从空格分隔的行中提取一些字段。当然,当该字段被引用时,我会得到部分结果。请问有人可以提出解决方案吗?
【问题讨论】:
下次显示您的输入文件和所需的输出。要获取引用字段,
$ cat file
field1 field2 "field 3" field4 "field5"
$ awk -F'"' '{for(i=2;i<=NF;i+=2) print $i}' file
field 3
field5
【讨论】:
这实际上是相当困难的。我想出了以下awk 脚本,它手动拆分行并将所有字段存储在一个数组中。
{
s = $0
i = 0
split("", a)
while ((m = match(s, /"[^"]*"/)) > 0) {
# Add all unquoted fields before this field
n = split(substr(s, 1, m - 1), t)
for (j = 1; j <= n; j++)
a[++i] = t[j]
# Add this quoted field
a[++i] = substr(s, RSTART + 1, RLENGTH - 2)
s = substr(s, RSTART + RLENGTH)
if (i >= 3) # We can stop once we have field 3
break
}
# Process the remaining unquoted fields after the last quoted field
n = split(s, t)
for (j = 1; j <= n; j++)
a[++i] = t[j]
print a[3]
}
【讨论】:
这里有一个可能的替代解决方案来解决这个问题。它的工作原理是查找以引号开头或结尾的字段,然后将它们连接在一起。最后它会更新字段和 NF,因此如果您在进行合并的模式之后放置更多模式,您可以使用所有正常的 awk 功能处理(新)字段。
我认为这仅使用 POSIX awk 的功能,不依赖 gawk 扩展,但我不完全确定。
# This function joins the fields $start to $stop together with FS, shifting
# subsequent fields down and updating NF.
#
function merge_fields(start, stop) {
#printf "Merge fields $%d to $%d\n", start, stop;
if (start >= stop)
return;
merged = "";
for (i = start; i <= stop; i++) {
if (merged)
merged = merged OFS $i;
else
merged = $i;
}
$start = merged;
offs = stop - start;
for (i = start + 1; i <= NF; i++) {
#printf "$%d = $%d\n", i, i+offs;
$i = $(i + offs);
}
NF -= offs;
}
# Merge quoted fields together.
{
start = stop = 0;
for (i = 1; i <= NF; i++) {
if (match($i, /^"/))
start = i;
if (match($i, /"$/))
stop = i;
if (start && stop && stop > start) {
merge_fields(start, stop);
# Start again from the beginning.
i = 0;
start = stop = 0;
}
}
}
# This rule executes after the one above. It sees the fields after merging.
{
for (i = 1; i <= NF; i++) {
printf "Field %d: >>>%s<<<\n", i, $i;
}
}
在输入文件上,例如:
thing "more things" "thing" "more things and stuff"
它产生:
Field 1: >>>thing<<<
Field 2: >>>"more things"<<<
Field 3: >>>"thing"<<<
Field 4: >>>"more things and stuff"<<<
【讨论】:
如果您只是在寻找特定领域,那么
$ cat file
field1 field2 "field 3" field4 "field5"
awk -F"\"" '{print $2}' file
有效。它用 " 分割文件,所以上例中的第二个字段就是你想要的。
【讨论】: