【问题标题】:Python data aggregation ("Countifs" function in Excel)Python 数据聚合(Excel 中的“Countifs”函数)
【发布时间】:2021-11-07 07:22:48
【问题描述】:

长期使用 Excel 的用户在这里变成了新的 Python 用户。我有以下产品 ID 数据框:

productID              sales
6976849                194,518,557             
11197085               277,387,647
70689391               197,511,925
70827164               242,995,691
70942756               1,529,319,200

(它在界面中看起来并不漂亮,但在 Python 中,我设法将它放入一个数据框中,其中有一列用于 ID 和一列用于销售。)

每个产品 ID 都有总销售额。

我需要统计有多少产品的销售额超过 200,000,000,以及有多少产品的销售额低于 200,000,000。

存储桶总数 超过 200,000,000 x 200,000,000 岁以下

在 Excel 中,我会使用快速 Countif 函数来执行此操作,但我不确定它在 Python 中是如何工作的。

我很难找到如何做到这一点 - 谁能指出我正确的方向?即使只是函数的名称,以便我可以阅读它们,也会很有用!

谢谢!!

【问题讨论】:

  • 使用Pandas
  • 作为初学者,在您对语言本身相当熟悉之前,我个人不会推荐使用 pandas 数据框。它们非常强大,但是您的应用程序非常简单,并不严格需要它们的灵活性。将文件保存为 csv,然后读取文件并解析文本将是学习处理文件、字符串转换和处理简单数据结构(如列表等)的一个很好的练习
  • 你是对的@Aaron。请问您可以使用csv 模块回答问题吗?

标签: python excel data-analysis aggregation countif


【解决方案1】:

使用Pandas 和value_counts:

import pandas as pd

df = pd.read_excel('data.xlsx')
over, under = df['sales'].gt(200000000).value_counts().tolist()

输出:

>>> over
3

>>> under
2

一步一步:

# Display your data after load file
>>> df
   productID       sales
0    6976849   194518557
1   11197085   277387647
2   70689391   197511925
3   70827164   242995691
4   70942756  1529319200

# Select the column 'sales'
>>> df['sales']
0     194518557
1     277387647
2     197511925
3     242995691
4    1529319200
Name: sales, dtype: int64

# Sales are greater than 200000000? (IF part of COUNTIF)
>>> df['sales'].gt(200000000)
0    False
1     True
2    False
3     True
4     True
Name: sales, dtype: bool

# Count True (over) and False (under) (COUNT part of COUNTIF)
>>> df['sales'].gt(200000000).value_counts()
True     3
False    2
Name: sales, dtype: int64

# Convert to list
>>> df['sales'].gt(200000000).value_counts().tolist()
[3, 2]

# Set variables over / under
>>> over, under = df['sales'].gt(200000000).value_counts().tolist()

更新

我还应该补充一点,数据集中有 1 亿行,我需要更多的桶,比如 2 亿到 5 亿之间 5 亿到 1 亿到 2 亿之间 1 亿以下你能告诉我我是怎么做到的吗?会去设置水桶吗?

使用pd.cut 和value_counts:

df['buckets'] = pd.cut(df['sales'], right=False, ordered=True,
                       bins=[0, 100e6, 200e6, 500e6, np.inf],
                       labels=['under 100M', '100-200M',
                               '200-500M', 'over 500M'])
>>> df
   productID       sales    buckets
0    6976849   194518557   100-200M
1   11197085   277387647   200-500M
2   70689391   197511925   100-200M
3   70827164   242995691   200-500M
4   70942756  1529319200  over 500M

>>> df.value_counts('buckets', sort=False)
buckets
under 100M    0
100-200M      2
200-500M      2
over 500M     1
dtype: int64

【讨论】:

  • 谢谢!我还应该补充一点,数据集中有 1 亿行,我需要更多的存储桶,比如 2 亿到 5 亿之间 5 亿到 1 亿到 2 亿之间 1 亿以下 你能告诉我我会怎么做设置桶?
  • 我根据您的评论更新了我的答案。如果适合您的需要,请不要忘记accept an answer。这对我们所有人都很重要。
【解决方案2】:

有很多非常强大的库可以帮助你快速完成工作,但你提到你才刚刚开始,所以我建议坚持只使用 python 本身作为学习体验。

Excel 可以将文件保存为“CSV”,这是一种非常简单的文本文件格式。 CSV 文件通常只是电子表格中每一行的逐行表示,列用逗号分隔。您的问题的一个非常简单的解决方案可能是逐行读取此类文件,确定销售数量是大于还是小于给定数量,然后将一个添加到适当的组中。

从读取文件开始非常简单。在代码中读取文件与您可能习惯的有点不同,但它非常简单,并且在不同的编程语言之间几乎总是相同的。首先,您“打开”该文件,该文件基本上要求操作系统进行访问。操作系统将确定您是否应该有权访问,如果有,将允许您访问相当于可以扫描文件的游标。这个“光标”通常被称为“文件句柄”,或者在 python 中简称为“文件对象”。

file_obj = open("data.csv", "r")  #"r" for reading mode (instead of write)

一旦我们有了句柄,我们就可以调用它的read 方法来读取部分或全部文件。传递一个数字会告诉你要读取多少个字符(一个新行和一些其他不可见的代码算作一个字符),或者在 python 中你可以不传递任何参数,它将读取整个文件。因为我们以文本模式(默认)打开文件,所以我们会得到一个常规字符串。

file_contents = file_obj.read()

现在我们已经将文件的内容复制到一个字符串中,我们可以关闭文件并告诉操作系统我们已经完成了它。

file_obj.close()

您可能经常在 python 中看到这一点有点不同,使用 with 上下文以便稍微清理代码,并确保文件始终关闭:

with open("data.csv", "r") as file_obj:
    file_contents = file_obj.read()

现在您有了一个包含整个 CSV 文件的大长字符串。我们希望能够逐行进行,因此我们必须在找到换行符的任何地方将字符串分成几部分。 Python 字符串使用str.split 方法可以方便地执行此操作:

file_lines = file_contents.split("\n") #split the string into a list of strings by splitting on \n chars

现在,我们将使用此列表创建一个循环来读取每一行,并确定产品的销售额是多于还是少于 200,000,000。为此,我们必须找到每一行的适当部分,将其转换为数字,并使用if 语句来决定如何处理它。最简单的 python 循环形式是遍历一个简单的项目列表中的每个项目:for line in file_lines:。现在,在循环内的每次迭代中,line 变量将填充文件该行中的任何内容。只要 csv 文件没有空行,也没有标题,我们几乎只需要找到第二个值(用逗号分隔),并将其转换为数字。这里我们将再次使用str.split 方法,在逗号处分割字符串,并将第二项转换为整数(记住python使用基于0的索引,所以第二项将是索引1)。从这里我们可以添加一些逻辑来确定我们是否应该将一个计算为低于或超过 2 亿。

over = 0
under = 0
for line in file_lines:
    line_items = line.split(",")
    sales = int(line_items[1])
    if sales > 200000000:
        over = over + 1
    else:
        under = under + 1

在阅读了您关于需要更多“桶”的评论后,终于有了一个更高级的版本

buckets = [0] * 8 #list of 8 buckets each starting at 0
bucket_edges = [0, 1e6, 2e6, 5e6, 1e7, 2e7, 5e7, 1e8, 2e8] #an even more advanced version would find the max and min of the sales figures, and dynamically calculate the bucket edges
with open("data.csv") as f: #"r" mode is actually the default
    sales = [] #an empty list
    for line in f: #a default behavior of file handles is to read line-by-line in for loops
        line = line.strip() #remove leading and trailing whitespace
        if not line: #an empty string will act like "False"
            continue #advance to the next loop iteration
        sales.append(int(line.split(",")[1])) #combine a few operations in a single line, and build up a list of sales figures
for sale in sales:
    for i in range(len(buckets)):
        if bucket_edges[i] <= sale < bucket_edges[i+1]:
            buckets[i] += 1
for i in range(len(buckets)):
    print(bucket_edges[i], "to", bucket_edges[i+1], ":", buckets[i])

【讨论】:

  • 很好的解释,+1。但是你必须使用csv 模块(这是标准库的一个模块)。或许,您也可以使用 bisect (bisect_left) 和 collections (defaultdict) 模块。如果您不使用标准库中的任何模块,Python 将不会有用。
【解决方案3】:

如果使用 convtools 并将 xlsx 文件转换为 csv,那么简短的答案是:

# pip install convtools
from convtools import conversion as c

converter = c.aggregate({
    "below_200": c.ReduceFuncs.Count(where=c.item("sales") < 200000000),
    "above_200": c.ReduceFuncs.Count(where=c.item("sales") >= 200000000),
}).gen_converter()
results = converter(list_of_dicts)

更长的:

import csv

# pip install convtools
from convtools import conversion as c

buckets = [
    (None, 200000000),
    (200000000, 220000000),
    (220000000, 240000000),
    (220000000, None),
]


def bucket_to_condition(bucket, input_):
    conditions = []
    left, right = bucket
    if left is not None:
        conditions.append(input_ >= left)
    if right is not None:
        conditions.append(input_ < right)
    if not conditions:
        return {}
    return {
        "where": c.and_(*conditions) if len(conditions) > 1 else conditions[0]
    }


with open("input_data.csv", "w") as f:
    reader = csv.reader(f)

    # skip header (assuming it's known)
    next(header)

    converter = (
        c.iter(
            {
                "productID": c.item(0),
                "sales": c.item(1).as_type(int),
            }
        )
        .pipe(
            c.aggregate(
                {
                    f"{bucket}": c.ReduceFuncs.Count(
                        **bucket_to_condition(bucket, c.item("sales")),
                    )
                    for bucket in buckets
                }
            )
        )
        .gen_converter()
    )
    results = converter(reader)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-08-27
    • 2020-11-21
    • 2022-01-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多