【发布时间】:2017-12-22 12:40:48
【问题描述】:
使用以下代码:
for root, dirs, files in os.walk(corpus_name):
for file in files:
if file.endswith(".v4_gold_conll"):
f= open(file)
lines = f.readlines()
tokens = [line.split()[3] for line in lines if line.strip()
and not line.startswith("#")]
print(tokens)
我收到以下错误:
Traceback(最近一次调用最后一次):文件“text_statistics.py”,行 28,在 corpus_reading_pos(corpus_name, option) 文件“text_statistics.py”,第 13 行,在 corpus_reading_pos f= open(file) FileNotFoundError: [Errno 2] No such file or directory: 'abc_0001.v4_gold_conll'
如您所见,该文件实际上已找到,但是当我尝试打开该文件时,它...找不到它?
编辑: 使用这个更新的代码,它在读取 7 个文件后停止,但有 172 个文件。
def corpus_reading_token_count(corpus_name, option="token"):
for root, dirs, files in os.walk(corpus_name):
tokens = []
file_count = 0
for file in files:
if file.endswith(".v4_gold_conll"):
with open((os.path.join(root, file))) as f:
tokens += [line.split()[3] for line in f if line.strip() and not line.startswith("#")]
file_count += 1
print(tokens)
print("File count:", file_count)
【问题讨论】:
-
您在
corpus_name中找到该文件,但您在当前工作目录中打开它。 -
所以我的语料库包含数百个文件,我只需要访问以“.v4_gold_conll”结尾的文件并提取信息。我不知道我会怎么做...
标签: python python-3.x