【问题标题】:Python/Pandas Write File Line by Line :: Memory UsePython/Pandas 逐行写入文件 :: 内存使用
【发布时间】:2014-11-12 03:19:14
【问题描述】:

我有一个用 Pandas (~9GB) 加载到内存中的大型数据框。我正在尝试写出遵循给定格式(Vowpal Wabbit)的文本文件,并且对内存使用和性能感到困惑。虽然文件很大(4800 万行),但对 Pandas 的初始加载还不错。写出文件至少需要 6 个多小时,并且几乎压垮了我的笔记本电脑,几乎消耗了我的所有 RAM (32GB)。天真地,我假设这个操作一次只在一条线上操作,所以 RAM 的使用会非常小。有没有更有效的方法来处理这些数据?

with open("C:\\Users\\Desktop\\DATA\\train_mobile2.vw", "wb") as outfile:
    for index, row in train.iterrows():
        if row['click'] ==0:
            vwline=""
            vwline+="-1 "
        else:
            vwline=""
            vwline+="1 "
        vwline+="|a C1_"+ str(row['C1']) +\
        " |b banpos_"+ str(row['banner_pos']) +\
        " |c siteid_"+ str(row['site_id']) +\
        " sitedom_"+ str(row['site_domain']) +\
        " sitecat_"+ str(row['site_category']) +\
        " |d appid_"+ str(row['app_id']) +\
        " app_domain_"+ str(row['app_domain']) +\
        " app_cat_"+ str(row['app_category']) +\
        " |e d_id_"+ str(row['device_id']) +\
        " d_ip_"+ str(row['device_ip']) +\
        " d_os_"+ str(row['device_os']) +\
        " d_make_"+ str(row['device_make']) +\
        " d_mod_"+ str(row['device_model']) +\
        " d_type_"+ str(row['device_type']) +\
        " d_conn_"+ str(row['device_conn_type']) +\
        " d_geo_"+ str(row['device_geo_country']) +\
        " |f num_a:"+ str(row['C17']) +\
        " numb:"+ str(row['C18']) +\
        " numc:"+ str(row['C19']) +\
        " numd:"+ str(row['C20']) +\
        " nume:"+ str(row['C22']) +\
        " numf:"+ str(row['C24']) +\
        " |g c21_"+ str(row['C21']) +\
        " C23_"+ str(row['C23']) +\
        " |h hh_"+ str(row['hh']) +\
        " |i doe_"+ str(row['doe']) 
        outfile.write(vwline + "\n")

响应用户的建议,

我编写了以下代码,但是当它运行的最后一行显示“+: 'numpy.ndarray' 和 'str' 不支持的操作数类型”时出现错误

lines_T = np.where(train['click'] == 0, "-1 ", "1 ") +\
        "|a C1_" + train['C1'].astype('str') +\
        " |b banpos_"+ train['banner_pos'].astype('str') +\
....

        "|h hh_"+ train['hh'].astype('str')+\
        " |i doe_"+ train['doe'].astype('str')    #ERROR HERE

line_T.to_csv("C:\Users\Desktop\DATA\KAGGLE\mobile\train_mobile.vw",mode='a', header=False,index=False)

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    不确定内存使用情况,但这肯定会更快:

    lines = np.where(train['click'] == 0, "-1 ", "1 ") +
            "|a C1_" + train['C1'].astype('str') +
            " |b banpos_"+ train['banner_pos'].astype('str') +
            ...
    

    然后保存行

    lines.to_csv(outfile, index=False)
    

    如果内存成为问题,您也可以分批进行(例如几百万条记录)

    【讨论】:

    • 使用 iterrows() 效率低吗?
    • 一般来说,向量化操作速度更快,因为它们经过了高度优化。 Iterrows 较慢,但如果这是您的内存爆炸的原因,我会感到惊讶
    • 我重新编码,但在最后一行出现错误 - 我将添加到上面的问题。
    • 需要将 np.where() 与 pd.Series() 包装起来,就像在 pd.Series(np.where(...)) 中一样
    猜你喜欢
    • 2016-10-13
    • 2019-12-18
    • 2017-05-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多