【问题标题】:Count repeated element in a csv file计算 csv 文件中的重复元素
【发布时间】:2018-01-08 14:52:37
【问题描述】:

这是我的 csv 文件:

2017-07-14  03:05:23    B2KPRT320   - Error1
2017-07-14  03:05:23    B2KPRT320   - Error1
2017-07-15  03:05:23    B2KPRT320   - Error2
2017-07-15  03:05:23    B2KPRT320   - Error3

我需要计算每天的错误数

这是我目前为止的脚本:

import collections
Data = []
string = ""
array = []
with open('out.csv') as f:
    for line in f:
        Data.append([word for word in line.strip().split("\t")])

for item in Data:
    try:
        date,error = item[0],item[3]
        string = date + "\t" + error + "\n"
        array.append([word for word in string.strip().split("\t")])
    except IndexError:
        print "A line in the file doesn't have enough entries."

最后,我需要将结果保存在另一个 csv 文件中 这是输出:

2017-07-14   - Error1   2
2017-07-15   - Error2   1
2017-07-15   - Error3   1

【问题讨论】:

  • 您需要什么帮助?
  • 我会看看collections.counter。将您的(日期,错误)元组列表提供给它,让它为您计算。

标签: python csv collections count


【解决方案1】:

您可以将文件读入列表并使用collections.Counter() 计算重复错误,然后split() 每行获取第一项和最后一项。例如:

import collections
Data = []
string = ""
array = []
with open('test.txt') as f:
    Data = collections.Counter(f.read().splitlines())

for item, c in Data.items():
    item = item.split()
    date, error = item[0], item[-1]
    string = "{}\t{}\t{}".format(date, error, c)
    array.append(string)


for elem in array:
    print elem

这将输出:

2017-07-15  Error3  1
2017-07-15  Error2  1
2017-07-14  Error1  2

编辑:

您不再需要try/except,因为使用item[-1] 会为您提供列表的最后一项。相反,您可以使用:

if len(item) < x:
    # print error
else:
    # the above code

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2023-03-11
    • 2021-12-22
    • 2017-11-06
    • 2011-05-07
    • 1970-01-01
    • 2016-03-07
    • 1970-01-01
    相关资源
    最近更新 更多