【问题标题】:Interpret columns of zeros and ones as binary and store as an integer column将零和一的列解释为二进制并存储为整数列
【发布时间】:2016-10-05 14:42:16
【问题描述】:

我有一个零和一的数据框。我想将每一列视为其值是整数的二进制表示。进行这种转换的最简单方法是什么?

我想要这个:

df = pd.DataFrame([[1, 0, 1], [1, 1, 0], [0, 1, 1], [0, 0, 1]])

print df

   0  1  2
0  1  0  1
1  1  1  0
2  0  1  1
3  0  0  1

转换为:

0    12
1     6
2    11
dtype: int64

尽可能高效。

【问题讨论】:

  • 将每列乘以 2 的相关幂,然后将结果列相加?
  • @Evert 是的,我也刚刚想到。使用我指定的索引很方便。
  • 我发现如果您先转置数据集,这会更有意义:您现在正在“求和”每一列中的值,但最终结果显示的行只有一个数字;我曾期望有一个数字的列。转置时,很容易跨列求和,而不是跨行求和。

标签: python numpy pandas binary integer


【解决方案1】:

类似的解决方案,但速度更快:

print (df.T.dot(1 << np.arange(df.shape[0] - 1, -1, -1)))
0    12
1     6
2    11
dtype: int64

时间安排:

In [81]: %timeit df.apply(lambda col: int(''.join(str(v) for v in col), 2))
The slowest run took 5.66 times longer than the fastest. This could mean that an intermediate result is being cached.
1000 loops, best of 3: 264 µs per loop

In [82]: %timeit (df.T*(1 << np.arange(df.shape[0]-1, -1, -1))).sum(axis=1)
1000 loops, best of 3: 492 µs per loop

In [83]: %timeit (df.T.dot(1 << np.arange(df.shape[0] - 1, -1, -1)))
The slowest run took 6.14 times longer than the fastest. This could mean that an intermediate result is being cached.
1000 loops, best of 3: 204 µs per loop

【讨论】:

    【解决方案2】:

    您可以从列值创建一个字符串,然后使用int(binary_string, base=2) 转换为整数:

    df.apply(lambda col: int(''.join(str(v) for v in col), 2))
    Out[6]: 
    0    12
    1     6
    2    11
    dtype: int64
    

    不确定效率,乘以 2 的相关幂然后求和可能会更好地利用快速 numpy 操作,但这可能更方便。

    【讨论】:

      【解决方案3】:

      在概念上类似于使用 dot-product 的 @jezrael's solution,但有一些改进。我们可以通过从前面为dot-product 带来 2 次幂范围数组来避免转置。这对大型数组是有益的,因为转置它们会产生一些开销。此外,在 NumPy 数组上操作对于这些数字处理情况会更好,因此我们可以在 df.values 上操作。最后,我们需要转换为 pandas series/dataframe 以进行最终输出。

      因此,结合这两个改进,修改后的实现将是 -

      pd.Series((2**np.arange(df.shape[0]-1,-1,-1)).dot(df.values))
      

      运行时测试-

      In [159]: df = pd.DataFrame(np.random.randint(0,2,(4,10000)))
      
      In [160]: p1 = pd.Series((2**np.arange(df.shape[0]-1,-1,-1)).dot(df.values))
      
      # @jezrael's solution
      In [161]: p2 = (df.T.dot(1 << np.arange(df.shape[0] - 1, -1, -1)))
      
      In [162]: np.allclose(p1.values, p2.values)
      Out[162]: True
      
      In [163]: %timeit pd.Series((2**np.arange(df.shape[0]-1,-1,-1)).dot(df.values))
      1000 loops, best of 3: 268 µs per loop
      
      # @jezrael's solution
      In [164]: %timeit (df.T.dot(1 << np.arange(df.shape[0] - 1, -1, -1)))
      1000 loops, best of 3: 554 µs per loop
      

      【讨论】:

        猜你喜欢
        • 2012-03-15
        • 1970-01-01
        • 1970-01-01
        • 2016-01-12
        • 2014-05-21
        • 2019-06-30
        • 2018-10-06
        • 2021-03-29
        • 2012-12-15
        相关资源
        最近更新 更多