【问题标题】:Difference between every row and column in two DataFrames (Python / Pandas)两个 DataFrames (Python / Pandas) 中每一行和每一列的区别
【发布时间】:2014-10-25 03:09:43
【问题描述】:

有没有更有效的方法来比较一个 DF 中每一行中的每一列与另一个 DF 中每一行中的每一列?这对我来说感觉很草率,但是我的循环/应用尝试要慢得多。

df1 = pd.DataFrame({'a': np.random.randn(1000),
                   'b': [1, 2] * 500,
                   'c': np.random.randn(1000)},
                   index=pd.date_range('1/1/2000', periods=1000))
df2 = pd.DataFrame({'a': np.random.randn(100),
                'b': [2, 1] * 50,
                'c': np.random.randn(100)},
               index=pd.date_range('1/1/2000', periods=100))
df1 = df1.reset_index()
df1['embarrassingHackInd'] = 0
df1.set_index('embarrassingHackInd', inplace=True)
df1.rename(columns={'index':'origIndex'}, inplace=True)
df1['df1Date'] = df1.origIndex.astype(np.int64) // 10**9
df1['df2Date'] = 0
df2 = df2.reset_index()
df2['embarrassingHackInd'] = 0
df2.set_index('embarrassingHackInd', inplace=True)
df2.rename(columns={'index':'origIndex'}, inplace=True)
df2['df2Date'] = df2.origIndex.astype(np.int64) // 10**9
df2['df1Date'] = 0
timeit df3 = abs(df1-df2)

10 个循环,3 个循环中的最佳值:每个循环 60.6 毫秒

我需要知道进行了哪个比较,因此将每个相反的索引添加到比较 DF 中,这样它最终会出现在最终的 DF 中。

提前感谢您的帮助。

【问题讨论】:

  • 我忘了说,我的实际 DF 有数百万行和数十列要比较。有了这个大小,应用尝试需要几个小时。
  • @EdChum 是的,我看到了那个,它决定了两个 DF 之间的变化,而不是值的差异。

标签: python pandas


【解决方案1】:

您发布的代码显示了一种生成减法表的巧妙方法。但是,它并不能发挥 Pandas 的优势。 Pandas DataFrame 将基础数据存储在基于列的块中。因此,按列而不是按行完成数据的检索是最快的。由于所有行都具有相同的索引,因此减法是按行执行的(将每一行与每隔一行配对),这意味着df1-df2 中正在进行大量基于行的数据检索。这对 Pandas 来说并不理想,尤其是当并非所有列都具有相同的 dtype 时。

减法表是 NumPy 擅长的:

In [5]: x = np.arange(10)

In [6]: y = np.arange(5)

In [7]: x[:, np.newaxis] - y
Out[7]: 
array([[ 0, -1, -2, -3, -4],
       [ 1,  0, -1, -2, -3],
       [ 2,  1,  0, -1, -2],
       [ 3,  2,  1,  0, -1],
       [ 4,  3,  2,  1,  0],
       [ 5,  4,  3,  2,  1],
       [ 6,  5,  4,  3,  2],
       [ 7,  6,  5,  4,  3],
       [ 8,  7,  6,  5,  4],
       [ 9,  8,  7,  6,  5]])

您可以将x 视为df1 的一列,将y 视为df2 的一列。您将在下面看到 NumPy 可以使用基本相同的语法以基本相同的方式处理df1 的所有列和df2 的所有列。


下面的代码定义了orig 和using_numpy。 orig 是您发布的代码,using_numpy 是另一种使用 NumPy 数组执行减法的方法:

In [2]: %timeit orig(df1.copy(), df2.copy())
10 loops, best of 3: 96.1 ms per loop

In [3]: %timeit using_numpy(df1.copy(), df2.copy())
10 loops, best of 3: 19.9 ms per loop

import numpy as np
import pandas as pd
N = 100
df1 = pd.DataFrame({'a': np.random.randn(10*N),
                   'b': [1, 2] * 5*N,
                   'c': np.random.randn(10*N)},
                   index=pd.date_range('1/1/2000', periods=10*N))
df2 = pd.DataFrame({'a': np.random.randn(N),
                'b': [2, 1] * (N//2),
                'c': np.random.randn(N)},
               index=pd.date_range('1/1/2000', periods=N))

def orig(df1, df2):
    df1 = df1.reset_index() # 312 µs per loop
    df1['embarrassingHackInd'] = 0 # 75.2 µs per loop
    df1.set_index('embarrassingHackInd', inplace=True) # 526 µs per loop
    df1.rename(columns={'index':'origIndex'}, inplace=True) # 209 µs per loop
    df1['df1Date'] = df1.origIndex.astype(np.int64) // 10**9 # 23.1 µs per loop
    df1['df2Date'] = 0

    df2 = df2.reset_index()
    df2['embarrassingHackInd'] = 0
    df2.set_index('embarrassingHackInd', inplace=True)
    df2.rename(columns={'index':'origIndex'}, inplace=True)
    df2['df2Date'] = df2.origIndex.astype(np.int64) // 10**9
    df2['df1Date'] = 0
    df3 = abs(df1-df2) # 88.7 ms per loop  <-- this is the bottleneck
    return df3

def using_numpy(df1, df2):
    df1.index.name = 'origIndex'
    df2.index.name = 'origIndex'
    df1.reset_index(inplace=True) 
    df2.reset_index(inplace=True) 
    df1_date = df1['origIndex']
    df2_date = df2['origIndex']
    df1['origIndex'] = df1_date.astype(np.int64) 
    df2['origIndex'] = df2_date.astype(np.int64) 

    arr1 = df1.values
    arr2 = df2.values
    arr3 = np.abs(arr1[:,np.newaxis,:]-arr2) # 3.32 ms per loop vs 88.7 ms 
    arr3 = arr3.reshape(-1, 4)
    index = pd.MultiIndex.from_product(
        [df1_date, df2_date], names=['df1Date', 'df2Date'])
    result = pd.DataFrame(arr3, index=index, columns=df1.columns)
    # You could stop here, but the rest makes the result more similar to orig
    result.reset_index(inplace=True, drop=False)
    result['df1Date'] = result['df1Date'].astype(np.int64) // 10**9
    result['df2Date'] = result['df2Date'].astype(np.int64) // 10**9
    return result

def is_equal(expected, result):
    expected.reset_index(inplace=True, drop=True)
    result.reset_index(inplace=True, drop=True)

    # expected has dtypes 'O', while result has some float and int dtypes. 
    # Make all the dtypes float for a quick and dirty comparison check
    expected = expected.astype('float')
    result = result.astype('float')
    columns = ['a','b','c','origIndex','df1Date','df2Date']
    return expected[columns].equals(result[columns])

expected = orig(df1.copy(), df2.copy())
result = using_numpy(df1.copy(), df2.copy())
assert is_equal(expected, result)

x[:, np.newaxis] - y 的工作原理:

这个表达式利用了 NumPy 广播。 要理解广播——通常是 NumPy——了解数组的形状是值得的:

In [6]: x.shape
Out[6]: (10,)

In [7]: x[:, np.newaxis].shape
Out[7]: (10, 1)

In [8]: y.shape
Out[8]: (5,)

[:, np.newaxis] 在右侧 上为x 添加了一个新轴,因此形状为(10, 1)。所以x[:, np.newaxis] - y 是形状数组(10, 1) 与形状数组(5,) 的减法。

从表面上看,这没有任何意义,但 NumPy 数组广播它们的形状 according to certain rules 以尝试使它们的形状兼容。

第一条规则是可以在左边添加新轴。因此,形状(5,) 的数组可以将自己广播到形状(1, 5)。

下一条规则是长度为 1 的轴可以将自身广播到任意长度。数组中的值在额外维度上根据需要简单地重复。

所以当 (10, 1) 和 (1, 5) 形状的数组在 NumPy 算术运算中放在一起时,它们都会被广播到形状为 (10, 5) 的数组:

In [14]: broadcasted_x, broadcasted_y = np.broadcast_arrays(x[:, np.newaxis], y)

In [15]: broadcasted_x
Out[15]: 
array([[0, 0, 0, 0, 0],
       [1, 1, 1, 1, 1],
       [2, 2, 2, 2, 2],
       [3, 3, 3, 3, 3],
       [4, 4, 4, 4, 4],
       [5, 5, 5, 5, 5],
       [6, 6, 6, 6, 6],
       [7, 7, 7, 7, 7],
       [8, 8, 8, 8, 8],
       [9, 9, 9, 9, 9]])

In [16]: broadcasted_y
Out[16]: 
array([[0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4],
       [0, 1, 2, 3, 4]])

所以x[:, np.newaxis] - y 等价于broadcasted_x - broadcasted_y。

现在,有了这个更简单的例子,我们可以看看 arr1[:,np.newaxis,:]-arr2.

arr1 的形状为 (1000, 4),arr2 的形状为 (100, 4)。我们要减去长度为 4 的轴中的项目,对于沿 1000 长度轴的每一行,以及沿 100 长度轴的每一行。换句话说,我们希望减法形成一个形状为(1000, 100, 4) 的数组。

重要的是,我们不希望 1000-axis 与 100-axis 交互。 我们希望它们位于不同的轴上。

所以如果我们像这样给arr1添加一个轴:arr1[:,np.newaxis,:],那么它的形状就变成了

In [22]: arr1[:, np.newaxis, :].shape
Out[22]: (1000, 1, 4)

现在,NumPy 广播将两个阵列提升到(1000, 100, 4) 的共同形状。瞧,减法表。

要将值按摩到形状为(1000*100, 4) 的二维数据帧中,我们可以使用reshape:

arr3 = arr3.reshape(-1, 4)

-1 告诉 NumPy 用任何正整数替换 -1 以使重塑有意义。由于arr 有1000*100*4 个值,-1 被替换为1000*100。使用-1 比编写1000*100 更好,因为它允许代码工作,即使我们更改df1 和df2 中的行数。

【讨论】:

  • 您能解释一下x[:, np.newaxis] 的工作原理吗?我知道 x[:] 只是对整个表进行切片,但我不明白 np.newaxis 会发生什么(以及如何)。还有那个语法是什么?它是特定于 numpy 的吗?它可以以不同的方式使用吗?
  • 我添加了对x[:, np.newaxis] - y 工作原理的说明。
  • 非常感谢,这说明了很多
猜你喜欢
  • 2022-12-03
  • 1970-01-01
  • 2017-03-30
  • 2017-07-18
  • 2020-09-04
  • 2016-11-16
  • 1970-01-01
  • 2017-04-03
  • 2017-07-25
相关资源
最近更新 更多