【问题标题】:Python - Pandas: perform column value based data grouping across separate dataframe chunksPython - Pandas:跨单独的数据框块执行基于列值的数据分组
【发布时间】:2021-04-14 02:43:58
【问题描述】:

我在处理一个大的 csv 文件时遇到了这个问题。我正在读取 chunks 中的 csv 文件,并希望根据特定列的值提取子数据帧。

为了解释这个问题,这里是一个最小版本:

CSV(另存为 test1.csv,例如)

1,10
1,11
1,12
2,13
2,14
2,15
2,16
3,17
3,18
3,19
3,20
4,21
4,22
4,23
4,24

现在,如您所见,如果我以 5 行的块读取 csv,则第一列的值将分布在块中。我想要做的是仅将特定值的行加载到内存中。

我使用以下方法实现了它:

import pandas as pd

list_of_ids = dict()  # this will contain all "id"s and the start and end row index for each id

# read the csv in chunks of 5 rows
for df_chunk in pd.read_csv('test1.csv', chunksize=5, names=['id','val'], iterator=True):
    #print(df_chunk)

    # In each chunk, get the unique id values and add to the list
    for i in df_chunk['id'].unique().tolist():
        if i not in list_of_ids:
            list_of_ids[i] = []  # initially new values do not have the start and end row index

    for i in list_of_ids.keys():        # ---------MARKER 1-----------
        idx = df_chunk[df_chunk['id'] == i].index    # get row index for particular value of id
        
        if len(idx) != 0:     # if id is in this chunk
            if len(list_of_ids[i]) == 0:      # if the id is new in the final dictionary
                list_of_ids[i].append(idx.tolist()[0])     # start
                list_of_ids[i].append(idx.tolist()[-1])    # end
            else:                             # if the id was there in previous chunk
                list_of_ids[i] = [list_of_ids[i][0], idx.tolist()[-1]]    # keep old start, add new end
            
            #print(df_chunk.iloc[idx, :])
            #print(df_chunk.iloc[list_of_ids[i][0]:list_of_ids[i][-1], :])

print(list_of_ids)

skip = None
rows = None

# Now from the file, I will read only particular id group using following
#      I can again use chunksize argument to read the particular group in pieces
for id, se in list_of_ids.items():
    print('Data for id: {}'.format(id))
    skip, rows = se[0], (se[-1] - se[0]+1)
    for df_chunk in pd.read_csv('test1.csv', chunksize=2, nrows=rows, skiprows=skip, names=['id','val'], iterator=True):
        print(df_chunk)

我的代码的截断输出:

{1: [0, 2], 2: [3, 6], 3: [7, 10], 4: [11, 14]}
Data for id: 1
   id  val
0   1   10
1   1   11
   id  val
2   1   12
Data for id: 2
   id  val
0   2   13
1   2   14
   id  val
2   2   15
3   2   16
Data for id: 3
   id  val
0   3   17
1   3   18

我想问的是,我们有更好的方法吗?如果在代码中考虑MARKER 1,随着大小的增长,它的效率必然会降低。我确实节省了内存使用量,但是时间仍然是个问题。我们有一些现成的方法吗?

(我正在寻找完整的答案代码)

【问题讨论】:

  • 所以你要先读all one,all two等等?,还有什么是Marker 1?
  • 是的,在实际数据集中,所有1s(和其他人)可能有很多行。我想使用有限的块大小。 MARKER 1 在我分享的代码中:for i in list_of_ids.keys()
  • 所以您只希望前 5 行(1 行)或所有行(1 行)加载到内存中?
  • 为了确认,即使在读取所有1s 等时,我可能需要使用分块读取,但是,我想确保对于特定 id,我可以读取与它!

标签: python-3.x pandas csv


【解决方案1】:

我建议你为此使用itertools,如下:

import pandas as pd
import csv
import io

from itertools import groupby, islice
from operator import itemgetter


def chunker(n, iterable):
    """
    From answer: https://stackoverflow.com/a/31185097/4001592
    >>> list(chunker(3, 'ABCDEFG'))
    [['A', 'B', 'C'], ['D', 'E', 'F'], ['G']]
    """
    iterable = iter(iterable)
    return iter(lambda: list(islice(iterable, n)), [])


chunk_size = 5
with open('test1.csv') as csv_file:
    reader = csv.reader(csv_file)
    for _, group in groupby(reader, itemgetter(0)):
        for chunk in chunker(chunk_size, group):
            g = [','.join(e) for e in chunk]
            df = pd.read_csv(io.StringIO('\n'.join(g)), header=None)
            print(df)
            print('---')

输出 (部分)

   0   1
0  1  10
1  1  11
2  1  12
---
   0   1
0  2  13
1  2  14
2  2  15
3  2  16
---
   0   1
0  3  17
1  3  18
2  3  19
3  3  20
---
...

这种方法将首先按第 1 列分组读取:

for _, group in groupby(reader, itemgetter(0)):

每个组将被读取为 5 行的块(这可以使用 chunk_size 进行更改):

for chunk in chunker(chunk_size, group):

最后一部分:

g = [','.join(e) for e in chunk]
df = pd.read_csv(io.StringIO('\n'.join(g)), header=None)
print(df)
print('---')

创建一个合适的字符串传递给 pandas。

【讨论】:

  • 只是我刚刚在我的问题中更新的另一部分,是否可以分块读取组本身?看看我的代码中的最后一个循环,看看我的意思!还有itemgetter(0)这个是用来选择列的吧?
  • @anurag 是的,itemgetter(0) 是选择要分组的列,chunker 已经在分块读取组。在示例输出中看不到它,因为块大小为 5
  • 功能在答案第二部分解释here
猜你喜欢
  • 2021-09-29
  • 1970-01-01
  • 2015-01-15
  • 2019-06-05
  • 2014-03-28
  • 1970-01-01
  • 2018-10-17
  • 2016-11-13
  • 1970-01-01
相关资源
最近更新 更多