【问题标题】:Cumulative sum based on certain column values基于某些列值的累积和
【发布时间】:2018-04-09 03:07:58
【问题描述】:

[01/Aug/1995:00:54:59 -0400] "GET /images/opf-logo.gif HTTP/1.0" 200 32511 [01/Aug/1995:00:55:04 -0400] “获取 /images/ksclogosmall.gif HTTP/1.0" 200 4635 [01/Aug/1995:00:55:06 -0400] "GET /images/ksclogosmall.gif HTTP/1.0" 403 78787

我有一个来自 HTTP 服务器的文件,我需要根据最后一列的大小(以字节为单位)的累积总和列出前 10 个图像。

li = [i.strip().split() for i in open("input.txt").readlines()]

sorted_li = sorted(li, key = lambda cols : int(cols[6]), reverse = True)

sorted_out = {}

for l in sorted_li:

    if l[3] in sorted_out:
        sorted_out[l[3]] += int(l[6])
    else:
        sorted_out[l[3]] = int(l[6])

如何限制字典中的前 10 个值?有没有办法不使用 pandas 和 group by?

【问题讨论】:

    标签: python cumulative-sum


    【解决方案1】:

    您可以使用标准库中的Counter。

    from collections import Counter
    
    d = dict()
    with open('input.txt') as f:
        split_line_gen = (line.strip().split() for line in f)
        get_name_size_gen = ((line[3], int(line[-1])) for line in split_line_gen)
        for name, size in get_name_size_gen:
            d[name] = d.get(name, 0) + size
        c = Counter(d)
    

    要获得前 10 名,请使用 c.most_common(10)

    使用计数器可能有点开销。相反,您可以使用类似的东西

    sorted(d, key=d.get, reverse=True)[:10] 仅返回名称

    sorted(d.items(), key=lambda x: x[-1], reverse=True)[:10] 返回名称和大小

    但我会推荐使用 Counter -- 更易读,imo。

    【讨论】:

    • 但是,它不会做累积和,因为只获取前 10 条记录。
    • @paddu sorted 返回一个列表并且不会更改源,因此如果您要使用排序结果,则需要将其分配给某个变量,即result = sorted ...
    猜你喜欢
    • 1970-01-01
    • 2014-05-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-01-04
    • 2017-12-15
    • 1970-01-01
    相关资源
    最近更新 更多