【问题标题】:Get the pandas dataframe in chunks without repetition?以块的形式获取熊猫数据框而不重复?
【发布时间】:2020-12-07 02:21:09
【问题描述】:

我和其他几个人一起研究了以下StackOverflow answer,我可能已经厌倦了我犯了这个错误并且无法弄清楚确切的位置。我基本上想将 pandas 数据帧拆分成块,并通过 JSON 将其逐个发送到 API 端点。我不希望多次发送同一行。我的问题在以下流程的第 4 步中。

可重现的例子

第 1 步:创建数据框

# Dataframe Creation

import numpy as np
import pandas as pd

filenames = ["file_"+str(x) for x in np.arange(1, 11)]
languages = ['en', 'en', 'fr', 'en', 'en', 'en', 'es', 'en', 'fr', 'en']

test_df = pd.DataFrame({'file': filenames, 'lang': languages})

第 1 步输出

file    lang
0   file_1  en
1   file_2  en
2   file_3  fr
3   file_4  en
4   file_5  en
5   file_6  en
6   file_7  es
7   file_8  en
8   file_9  fr
9   file_10 en

第 2 步 - 两个函数

def get_chunk_df(large_df, splits):
    """splits df into chunks"""
    for chunk_df in np.array_split(large_df, splits):
        yield chunk_df


def get_json_chunks(df, splits):
    """converts each chunk to a dict which is basically going to be a JSON load"""
    documents = {"documents": []}
    df_chunks = get_chunk_df(df, splits)
    for chunk_df in df_chunks:
        for idx, row in chunk_df.iterrows():
            documents["documents"].append({
                "id": str(idx + 1),
                "text": row["lang"]
            })
        yield documents

第 3 步 - 测试 get_chunk_df 函数的输出 - 没问题

chunk_gen = get_chunk_df(test_df, 3)
counter = 0
for chk in chunk_gen:
    counter = counter + 1
    print(f"***********PRINTING {counter} CHUNK...")
    print(chk)

第 3 步输出

***********PRINTING 1 CHUNK...
     file lang
0  file_1   en
1  file_2   en
2  file_3   fr
3  file_4   en
***********PRINTING 2 CHUNK...
     file lang
4  file_5   en
5  file_6   en
6  file_7   es
***********PRINTING 3 CHUNK...
      file lang
7   file_8   en
8   file_9   fr
9  file_10   en

第 4 步 - 我的问题在这里

json_chunks = get_json_chunks(test_df, 3)

for json_chk in json_chunks:
    print(f"First row: {json_chk['documents'][0]}")
    print(f"Last row: {json_chk['documents'][-1]}")

第 4 步输出

First row: {'id': '1', 'text': 'en'}
Last row: {'id': '4', 'text': 'en'}
First row: {'id': '1', 'text': 'en'}
Last row: {'id': '7', 'text': 'es'}
First row: {'id': '1', 'text': 'en'}
Last row: {'id': '10', 'text': 'en'}

但我希望 预期输出为:

First row: {'id': '1', 'text': 'en'}
Last row: {'id': '4', 'text': 'en'}
First row: {'id': '5', 'text': 'en'}
Last row: {'id': '7', 'text': 'es'}
First row: {'id': '8', 'text': 'en'}
Last row: {'id': '10', 'text': 'en'}

谢谢!

【问题讨论】:

  • 你在get_json_chunks 中使用append() 所以也许这就是为什么你总是在数据框中有旧行。
  • @furas 我确实注意到了这一点,并且正要回答我自己的问题:)。随意回答,我会接受的。

标签: python pandas dataframe generator


【解决方案1】:

您在for-loop 之前创建documents = {"documents": []},然后在append 中创建相同的documents,但您必须在for-loop 中创建新的documents

def get_json_chunks(df, splits):
    """converts each chunk to a dict which is basically going to be a JSON load"""
    
    #documents = {"documents": []}  # <-- wrong place
    df_chunks = get_chunk_df(df, splits)
    
    for chunk_df in df_chunks:

        documents = {"documents": []}  # <-- good place

        for idx, row in chunk_df.iterrows():
            documents["documents"].append({
                "id": str(idx + 1),
                "text": row["lang"]
            })

        yield documents
        

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2017-11-13
    • 2021-09-04
    • 2018-12-29
    • 2021-09-18
    • 1970-01-01
    • 2017-11-26
    • 2019-01-17
    相关资源
    最近更新 更多