【问题标题】:How to resolve Pandas performance warning "highly fragmented" after using many custom np.where statements?使用许多自定义 np.where 语句后如何解决 Pandas 性能警告“高度碎片化”?
【发布时间】:2022-06-11 00:24:46
【问题描述】:

我有一个项目,我将代码从 SQL 转换为 Pandas。我的数据集/数据框中有 80 个自定义元素 - 每个都需要自定义逻辑。在 SQL 中,我在单个 Select 中使用多个 case 语句,如下所示:

Select x, y, z,
(case when statement1 then 0
when statement2 then 0
else 1 end) as custom_element1,
next case statement...as custom_element2,
next case statement...as custom_element3,
etc...

现在在 Pandas 中,我希望就实现相同目标的最有效方法提供一些建议。为了更容易重现,这里有一个例子,它做了我想做的同样的事情。我需要创建 80 个自定义输出变量。在此示例中,我只是使用不同的 np.where 语句一次添加一个自定义元素。

df = pd.DataFrame({'num_legs': [2, 4, 8, 0],
                   'num_wings': [2, 0, 0, 0]},
                   index=['falcon', 'dog', 'spider', 'fish'])
df['custom1'] = np.where(df['num_legs'].values > 2, 1, 0)
df['custom2'] = np.where(df['num_wings'] == df['num_legs'], 1, 0)
df['custom3'] = np.where((df['num_wings'].values == 0) | (df['num_legs'].values == 0), 1, 0)

我可以从连续的 np.where 语句中获取输出,以与原始 SQL 的输出完全匹配,所以那里没有问题。

但是我看到了这个警告:

DataFrame is highly fragmented...poor performance...Consider using pd.concat 
instead...or use copy().

所以我的问题是,以我为例,如何提高性能?我将如何在这里使用 pd.concat ?有什么比我上面显示的更好的方式来构建代码?我试图在这个论坛中寻找答案,但没有找到任何东西。感谢您的回复。

【问题讨论】:

  • 我们应该如何猜测你在做什么?请提供minimal reproducible example
  • 几乎可以肯定,您应该使用完全不同的方法。无论如何,您确实必须提供有关您正在做什么的具体细节。

标签: python pandas copy concatenation where-clause


猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2022-01-27
  • 1970-01-01
  • 1970-01-01
  • 2010-09-08
  • 2019-12-31
  • 2015-06-18
  • 2013-09-30
相关资源
最近更新 更多