【问题标题】:Find all combinations of features查找所有特征组合
【发布时间】:2020-07-26 04:26:20
【问题描述】:

我需要将我的二进制编码特征矩阵转换为包含所有可能的特征交互组合的矩阵。我的意思是字面上的所有组合(每组 2 个、每组 3 个、每组 4 个、每组全部等)。

有人知道 sklearn.preprocessing 是否有办法做到这一点?还是其他库?

将此数组输入到某个函数或方法中:

array([[0, 1, 1],
       [1, 0, 0],
       [1, 1, 1]])

并将其作为输出

array([[0, 0, 1, 0],
       [0, 0, 0, 0],
       [1, 1, 1, 1]])

新矩阵中的每一行代表[x1*x2, x1*x3, x2*x3, x1*x2*x3]

【问题讨论】:

  • 请张贴minimal reproducible example,显示您的目标。
  • 你能解释一下为什么会这样输出吗?我似乎不清楚
  • 我也一样。
  • 抱歉矩阵不正确。我更正了输出矩阵。它是每个特征的详尽组合,组合长度从 2 到特征计数。每个特征都必须以所有可能的长度乘以其他所有特征。所以特征 1 与特征 2,特征 1 与特征 3,特征 2 与 3,所有 3 个特征相乘。
  • 当然,在我的实际应用中,我还有很多很多的功能。这种类型的全功能交互在我的领域中有特定领域的应用。

标签: python machine-learning scikit-learn feature-extraction


【解决方案1】:

您想要的是powerset。所以你想找到你的特征的幂集,然后乘以相应的二进制值,这基本上是一个np.bitwise_and。所以你可以这样做:

  • 获取powerset,查找所有特征组合,最长可达len(features)
  • 减少np.logical_and.reduce
  • 附加到包含 powerset 中所有 sets 的列表

a = np.array([[0, 1, 1],
              [1, 0, 0],
              [1, 1, 1]])

from itertools import chain, combinations

features = a.T.tolist()
power_set = []
for comb in chain.from_iterable(combinations(features, r) 
                               for r in range(2,len(features)+1)):
    power_set.append(np.logical_and.reduce(comb).view('i1').tolist())

这会给你:

np.array(power_set).T

array([[0, 0, 1, 0],
       [0, 0, 0, 0],
       [1, 1, 1, 1]])

【讨论】:

  • 非常感谢。这是一个优雅的解决方案。但是,当我的初始矩阵已经是 1,000,000 行乘 5,000 个特征(在该范围内)时,您认为这是创建幂集的可行方法吗?
  • 不,不是。据我所知,没有矢量化的方法可以创建电源集。所以这在计算上是昂贵的。也许限制组合的范围。或者看看其他一些功能工程工具,如果这是你所追求的,比如featuretools@MichaelHavlin
【解决方案2】:

更新:好的,一个 for 循环消失了,还有一个。

recipes section of itertools 中有一个很好的 powerset 函数:

def powerset(iterable):
    "powerset([1,2,3]) --> () (1,) (2,) (3,) (1,2) (1,3) (2,3) (1,2,3)"
    s = list(iterable)
    return chain.from_iterable(combinations(s, r) for r in range(len(s)+1))

这应该可以。它尚未针对速度进行优化。 如果这提供了您正在寻找的结果,则可能可以删除嵌套的 for 循环。

import numpy as np
from itertools import combinations, chain

features = np.array([[0, 1, 1],
                     [1, 0, 0],
                     [1, 1, 1]])

n = features.shape[1]

def powerset(iterable):
    "powerset([1,2,3]) --> () (1,) (2,) (3,) (1,2) (1,3) (2,3) (1,2,3)"
    s = list(iterable)
    return chain.from_iterable(combinations(s, r) for r in range(len(s)+1))

# get all combinations, we will use this as indices for the columns later
indices = list(powerset(range(n)))

# remove the empty subset
indices.pop(0)

print(indices)

data = []

for i in indices:

    print()
    print(i)
    _ = features[:, i]
    print(_)

    x = np.prod(_, axis=1)
    print(x)
    data.append(x)

print(np.column_stack(data))

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多