【问题标题】:Speed up loop over each row for big dataset in Python加快Python中大数据集每一行的循环
【发布时间】:2019-02-19 09:45:32
【问题描述】:

我想通过根据其他列值(两列或三列以上)将值分配给新列来处理大型数据集。我有下面的 Python 代码。

我的数据集包含 1700 万条数据记录。运行脚本需要 40 多个小时。我是 Python 新手,对大数据的经验很少。

有人可以帮我加快脚本运行时间吗?

这是数据集的样本:

 PId    hZ  tId tPurp   ps  oZ  dZ  oT  dT
0   1   50  1040    32  762 748 10.5    12.5
0   1   50  1040    16  748 81  12.5    12.5
0   1   50  1040    2048    81  1   12.5    12.5
0   1   50  1040    1040    1   762 9.5 9.5
1   1   10  320 320 1   35  17.5    17.5
1   1   10  320 2048    35  1   19.5    19.5
2   1   50  1152    1152    297 102 11.5    12
2   1   50  1152    2048    102 1   12  12
2   1   50  1152    32  1   297 11.5    11.5
3   1   1   2   64  737 184 14  18
3   1   1   2   128 184 713 14  14
3   1   1   2   2048    184 1   18  18
3   1   1   2   2   1   737 9   9
4   1   1   2   2   1   856 9   9
4   1   1   2   2048    296 1   18  18
4   1   1   2   16  856 296 17  18
8   1   50  1056    16  97  7   15  15.5
8   1   50  1056    32  7   816 15.5    1
8   1   50  1056    2048    816 1   1   1
8   1   50  1056    1056    1   97  12  12

以下是 Python 代码

import pandas as pd 
import numpy as np
df_test = pd.read_csv("C:/users/test.csv")
df_test.sort_values(by=['PId','tId','oT','dT'],inplace=True)


ls2t = df_test.groupby(['PId','tId']).nth(-2)

ls2t.reset_index(level=(0,1),inplace=True)


ls2tps=ls2t[['PId','tId','ps']]

ls2tps=ls2tps.rename(columns = {'ps':'ls2ps'})

df_lst = pd.merge(df_test,
                 ls2tps,
                 on=['PId','tId'],
                 how='left')

for index,row in df_lst.iterrows():
    if df_lst.loc[index,'oZ']==df_lst.loc[index,'hZ'] and df_lst.loc[index,'ps']==2: 
       df_lst.loc[index,'d'] = 'A'
    elif df_lst.loc[index,'oZ']==df_lst.loc[index,'hZ'] and df_lst.loc[index,'ps']!=2:
         df_lst.loc[index,'d']='B'
    elif df_lst.loc[index,'ps']==2048 and (df_lst.loc[index,'ls2ps']==2 or df_lst.loc[index,'ls2ps']==514):
        df_lst.loc[index,'d']='A'
    elif df_lst.loc[index,'ps']==2048 and (df_lst.loc[index,'ls2ps']!=2 and df_lst.loc[index,'ls2ps']!=514):
        df_lst.loc[index,'d']='B'
    else:
        df_lst.loc[index,'d']='C'

od_aggpurp = df_lst.groupby(['oZ','dZ','d']).size().reset_index(name='counts')

od_aggpurp.to_csv('C:/users/test_result.csv')

【问题讨论】:

  • 好吧..还有一个类似于这个的问题..我认为可以有两种方法来解决这个问题..首先将您的数据帧分成多个块,然后在它们上异步执行您的逻辑(asyncio 会有所帮助).. 或尝试类似 hadoop 集群.. 只是一个建议.. 参数是受欢迎的

标签: python pandas performance loops bigdata


【解决方案1】:

你应该试试这个,而不是那个循环:

df_lst.loc[(df_lst['oZ'] == df_lst['hZ']) & (df_lst['ps'] == 2), 'd'] = 'A'  
df_lst.loc[(df_lst['oZ'] == df_lst['hZ']) & (df_lst['ps'] != 2), 'd'] = 'B'
df_lst.loc[(df_lst['ps'] == 2048) & ((df_lst['ls2ps'] == 2) | (df_lst['ls2ps'] == 514)), 'd'] = 'A'
df_lst.loc[(df_lst['ps'] == 2048) & ((df_lst['ls2ps'] != 2) & (df_lst['ls2ps'] != 514)), 'd'] = 'B'
df_lst.loc[(df_lst['d'] != 'A') & (df_lst['d'] != 'B'), 'd'] = 'C'

在这里,您从 df_lst(使用 .loc)中仅选择具有请求参数的行,但您仅修改 d 列。

请注意,在数据框 和 之间的 pandas 中是 &,or 是 |而 not 是 ~.

如果你喜欢这个应该会更好:

oZ_hZ = df_lst['oZ'] == df_lst['hZ']
ps_2 = df_lst['ps'] == 2

df_lst.loc[(oZ_hZ) & (ps_2), 'd'] = 'A'  
df_lst.loc[(oZ_hZ) & (~ps_2), 'd'] = 'B'

ps_2048 = df_lst['ps'] == 2048
ls2ps_2 = df_lst['ls2ps'] == 2
ls2ps_514 = df_lst['ls2ps'] == 514

df_lst.loc[(ps_2048) & ((ls2ps_2) | (ls2ps_514)), 'd'] = 'A'
df_lst.loc[(ps_2048) & ((~ls2ps_2) & (~ls2ps_514)), 'd'] = 'B'

df_lst.loc[(df_lst['d'] != 'A') & (df_lst['d'] != 'B'), 'd'] = 'C'

【讨论】:

  • 感谢您的帮助。我尝试了建议的代码。但它给了我一个错误:Series 的真值是模棱两可的。对代码使用 a.empty、a.bool()、a.item()、a.any() 或 a.all()。。我想这可能是因为 df_lst.loc[df_lst['oZ'] == df_lst['hZ'] & df_lst['ps']==2]['d'] 生成了一个系列,所以它不能被分配带有字符串值。所以我使用了 df_lst.loc[df_lst['oZ'] == df_lst['hZ'] & df_lst['ps']==2,'d'] = 'A'。还是同样的错误。无法弄清楚出了什么问题。
  • 对不起,我忘记了更多的括号。现在它应该可以工作了。是的,df_lst['d] == 2048 生成一个布尔系列,但在这种情况下,它仅用于从数据框中选择元素。
  • 非常感谢您的帮助。真的很有帮助!!!运行只需不到2分钟!!!但我对 df_lst.loc[(oZ_hZ) & (ps_2)]['d'] = 'A' 做了一个小修改。因为它没有为 d 列分配任何值。所以我将代码更改为 df_lst.loc[((oZ_hZ) & (ps_2)),'d'] = 'A'。我不太确定这两个代码的区别。但后一种有效。
  • 我真的很高兴它有帮助!你说的对!正如他们在这里所说的pandas.pydata.org/pandas-docs/stable/… 第一个版本有时有效,有时无效,因为它可能会在其他地方创建副本。
猜你喜欢
  • 1970-01-01
  • 2021-06-29
  • 2019-02-21
  • 2016-10-12
  • 2013-04-01
  • 2013-08-28
  • 2014-07-27
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多