【发布时间】:2019-04-12 15:18:50
【问题描述】:
我需要重新编目所有运行时间超过 5 小时的电影。
样本数据:
239835<TAB> 92075<TAB>Moonlighting, seasons one and two<TAB>NVIDEO<TAB>DVD<TAB>6 videodiscs (approximately 1200 min.) :
628328 180001 7th heaven. NVIDEO DVD 5 videodiscs (15 hr., 57 min.) :
773429 291072 Veronica Mars. NVIDEO DVD 6 videodiscs (842 min.) :
789908 379843 Castle in the Sky NVIDEO JDVD 2 videodiscs (approximately 125 min.) :
856287 208624 The Munsters. NVIDEO DVD 12 videodiscs (approximately 33 hr.) :
1076125 254085 From up on Poppy Hill (Rated PG) NVIDEO JDVD 2 videodiscs (91 min.) :
1154016 264851 Columbo. NVIDEO DVD 5 videodiscs (725 min.) :
1217001 113980 CSI, crime scene investigation. NVIDEO DVD 5 videodiscs (approximately 732 min.) :
1227803 280535 Seattle Seahawks NVIDEO DVD 3 videodiscs (500 min.) :
1227804 280535 Seattle Seahawks NVIDEO DVD 3 videodiscs (500 min.) :
1287497 293511 Seattle Seahawks : NVIDEO DVD 3 videodiscs (400 min.) :
1287499 293511 Seattle Seahawks : NVIDEO DVD 3 videodiscs (400 min.) :
1367994 228775 Spongebob Squarepants. NVIDEO JDVD 4 videodiscs (469 min.) :
1368002 257248 SpongeBob SquarePants. NVIDEO JDVD 4 videodiscs (589 min.) :
是否有一个快速的 perl 或 awk sn-p 或 one-liner 可以: * 如果打印整行 * # of "min" 大于 300 或 * # of "hr(s)" 大于 5
类似:
perl -F\\t -ane 'print if $F[6] <substring or capture group representing minutes> > 300' file.csv
与awk更亲密:
awk -F'\t' '$6 ~ /^.*\(.*[3-9][[:digit:]]{2}[[:space:]]+min.*\)/ {print}' minutes.csv
正则表达式模式:
超过 300 分钟:
/^.*\(.*[[:space:]][3-9][[:digit:]]{2}[[:space:]]+min.*\)/
分钟数大于 1000:
/^.*\(.*[[:digit:]]{4,}[[:space:]]+min.*\)/
小时数大于 5:
/^.*\(.*[[:space:]][5-9]{1}[[:space:]]+hr.*\)/
大于 10 小时:
/^.*\(.*[[:space:]][[:digit:]]{4}[[:space:]]+hr.*\)/
有没有更简单更简洁的方法?
【问题讨论】:
-
制表符分隔的数据不能很好地与 SO 配合使用。该示例中的列在哪里分开?
-
将它们添加到第一行,以便您查看。