【问题标题】:Extract specific lines from multiple text files从多个文本文件中提取特定行
【发布时间】:2016-11-21 19:59:22
【问题描述】:
我想打印文件夹中多个文本文件中的某些行,具体取决于文件名。考虑以下由下划线分隔的 3 个单词命名的文本文件:
Small_Apple_Red.txt
Small_Orange_Yellow.txt
Large_Apple_Green.txt
Large_Orange_Green.txt
如何实现以下目标?
if (first word of file name is "Small") {
// Print row 3 column 2 of the file (space delimited);
}
if (second word of file name is "Orange") {
// print row 1 column 4 of the file;
}
这可以通过 awk 实现吗?
【问题讨论】:
标签:
java
python
perl
awk
sed
【解决方案1】:
尝试如下。
使用glob 处理文件夹中的文件。
然后使用正则表达式检查文件名。这里
grep 用于从文件中提取特定内容。
my $path = "folderpath";
while (my $file = glob("$path/*"))
{
if($file =~/\/Small_Apple/)
{
open my $fh, "<", "$file";
print grep{/content what you want/ } <$fh>;
}
}
【解决方案2】:
use strict;
use warnings;
my @file_names = ("Small_Apple_Red.txt",
"Small_Orange_Yellow.txt",
"Large_Apple_Green.txt",
"Large_Orange_Green.txt");
foreach my $file ( @file_names) {
if ( $file =~ /^Small/){ // "^" marks the begining of the string
print "\n $file has the first word small";
}
elsif ( $file =~ /.*?_Orange/){ // .*? is non-greedy, this means that it matches anything<br>
// until the first "_" is found
print "\n $file has the second word orange";
}
}
还有一种特殊情况,即您的文件有“Small_Orange”,您必须决定哪个更重要。如果第二个词更重要,则将if部分的内容切换为elsif部分的内容
【解决方案3】:
在 awk 中:
awk 'FILENAME ~ /^Large/ {print $1,$4}
FILENAME ~ /^Small/ {print $3,$2}' *
在 Perl 中:
perl -naE 'say "$F[0] $F[3]" if $ARGV =~ /^Large/;
say "$F[2] $F[1]" if $ARGV =~ /^Small/ ' *
【解决方案4】:
试试这个:
use strict;
use warnings;
use Cwd;
use File::Basename;
my $dir = getcwd(); #or shift the input values from the user
my @txtfiles = glob("$dir/*.txt");
foreach my $each_txt_file (@txtfiles)
{
open(DATA, $each_txt_file) || die "reason: $!";
my @allLines = <DATA>;
(my $removeExt = $each_txt_file)=~s/\.txt$//g;
my($word1, $word2, $word3) = split/\_/, basename $removeExt; #Select the file name with matching case
if($word1=~m/small/i) #Select your match case
{
my @split_space = "";
my @allrows = split /\n/, $allLines[1]; #Mentioned the row number
my @allcolns = split /\s/, $allrows[0];
print "\n", $allcolns[1]; #Mentioned the column number
}
}