【发布时间】:2021-04-14 02:43:58
【问题描述】:
我在处理一个大的 csv 文件时遇到了这个问题。我正在读取 chunks 中的 csv 文件,并希望根据特定列的值提取子数据帧。
为了解释这个问题,这里是一个最小版本:
CSV(另存为 test1.csv,例如)
1,10
1,11
1,12
2,13
2,14
2,15
2,16
3,17
3,18
3,19
3,20
4,21
4,22
4,23
4,24
现在,如您所见,如果我以 5 行的块读取 csv,则第一列的值将分布在块中。我想要做的是仅将特定值的行加载到内存中。
我使用以下方法实现了它:
import pandas as pd
list_of_ids = dict() # this will contain all "id"s and the start and end row index for each id
# read the csv in chunks of 5 rows
for df_chunk in pd.read_csv('test1.csv', chunksize=5, names=['id','val'], iterator=True):
#print(df_chunk)
# In each chunk, get the unique id values and add to the list
for i in df_chunk['id'].unique().tolist():
if i not in list_of_ids:
list_of_ids[i] = [] # initially new values do not have the start and end row index
for i in list_of_ids.keys(): # ---------MARKER 1-----------
idx = df_chunk[df_chunk['id'] == i].index # get row index for particular value of id
if len(idx) != 0: # if id is in this chunk
if len(list_of_ids[i]) == 0: # if the id is new in the final dictionary
list_of_ids[i].append(idx.tolist()[0]) # start
list_of_ids[i].append(idx.tolist()[-1]) # end
else: # if the id was there in previous chunk
list_of_ids[i] = [list_of_ids[i][0], idx.tolist()[-1]] # keep old start, add new end
#print(df_chunk.iloc[idx, :])
#print(df_chunk.iloc[list_of_ids[i][0]:list_of_ids[i][-1], :])
print(list_of_ids)
skip = None
rows = None
# Now from the file, I will read only particular id group using following
# I can again use chunksize argument to read the particular group in pieces
for id, se in list_of_ids.items():
print('Data for id: {}'.format(id))
skip, rows = se[0], (se[-1] - se[0]+1)
for df_chunk in pd.read_csv('test1.csv', chunksize=2, nrows=rows, skiprows=skip, names=['id','val'], iterator=True):
print(df_chunk)
我的代码的截断输出:
{1: [0, 2], 2: [3, 6], 3: [7, 10], 4: [11, 14]}
Data for id: 1
id val
0 1 10
1 1 11
id val
2 1 12
Data for id: 2
id val
0 2 13
1 2 14
id val
2 2 15
3 2 16
Data for id: 3
id val
0 3 17
1 3 18
我想问的是,我们有更好的方法吗?如果在代码中考虑MARKER 1,随着大小的增长,它的效率必然会降低。我确实节省了内存使用量,但是时间仍然是个问题。我们有一些现成的方法吗?
(我正在寻找完整的答案代码)
【问题讨论】:
-
所以你要先读all one,all two等等?,还有什么是Marker 1?
-
是的,在实际数据集中,所有
1s(和其他人)可能有很多行。我想使用有限的块大小。 MARKER 1 在我分享的代码中:for i in list_of_ids.keys() -
所以您只希望前 5 行(1 行)或所有行(1 行)加载到内存中?
-
为了确认,即使在读取所有
1s 等时,我可能需要使用分块读取,但是,我想确保对于特定 id,我可以读取与它!
标签: python-3.x pandas csv