【问题标题】:Create histogram from two arrays从两个数组创建直方图
【发布时间】:2022-01-09 13:05:21
【问题描述】:

我有两个具有相同维度的 numpy 数组:权重和百分比。百分比是“真实”数据,权重是直方图中每个“真实”数据的数量。

例如)

weights = [[0, 1, 1, 4, 2]
           [0, 1, 0, 3, 5]]
percents = [[1, 2, 3, 4, 5]
            [1, 2, 3, 4, 5]]

(每一行百分比都是一样的)

我想将这些“相乘”在一起,从而产生 weights[x] * [percents[x]]:

results = [[0 * [1] + 1 * [2] + 1 * [3] + 4 * [4] + 2 * [5]
           [0 * [1] + 1 * [2] + 0 * [3] + 3 * [4] + 5 * [5]]
        = [[2, 3, 4, 4, 4, 4, 5, 5]
           [2, 4, 4, 4, 5, 5, 5, 5, 5]]

请注意,每行的长度可能不同。理想情况下,这可以在 numpy 中完成,但正因为如此,它可能最终成为列表列表。

编辑: 我已经能够将这些嵌套的 for 循环拼凑在一起,但显然这并不理想:

list_of_hists = []
for index in df.index:
    hist = []
    # Create a list of lists, later to be flattened to 'results'
    for i, percent in enumerate(percents):
        hist.append(
        # For each percent, create a list of [percent] * weight
            [percent]
            * int(
                df.iloc[index].values[i]
            )
        )
    # flatten the list of lists in hist
    results = [val for list_ in hist for val in list_]
    list_of_hists.append(results)

【问题讨论】:

  • 你实际上不需要循环。但由于您输入的数组长度不平衡,因此np.split 之类的内容可能是一个不错的选择。

标签: python list numpy nested-lists


【解决方案1】:

有一个np.repeat 专为此类操作而设计,但它不适用于 2D 情况。所以你需要改用扁平化的数组视图。

weights = np.array([[0, 1, 1, 4, 2], [0, 1, 0, 3, 5]])
percents = np.array([[1, 2, 3, 4, 5], [1, 2, 3, 4, 5]])
>>> np.repeat(percents.ravel(), weights.ravel())
array([2, 3, 4, 4, 4, 4, 5, 5, 2, 4, 4, 4, 5, 5, 5, 5, 5])

然后,您需要选择将其拆分的索引位置:

>>> np.split(np.repeat(percents.ravel(), weights.ravel()), np.cumsum(np.sum(weights, axis=1)[:-1]))
[array([2, 3, 4, 4, 4, 4, 5, 5]), array([2, 4, 4, 4, 5, 5, 5, 5, 5])]

请注意,np.split 是非常低效的操作以及您希望从不等长的行中创建数组。

【讨论】:

  • 这也是一个聪明的解决方案,谢谢!
【解决方案2】:

你可以从functools使用list-comprehension和reduce:

import functools
res=[functools.reduce(lambda x,y: x+y,
                [x*[y] for x, y in zip(w, p)])
                for w, p in zip(weights, percents)]

输出:

[[2, 3, 4, 4, 4, 4, 5, 5],
 [2, 4, 4, 4, 5, 5, 5, 5, 5]]

或者,只是列表理解解决方案:

res= [[j for i in [x*[y]
              for x, y in zip(w, p)]
                for j in i]
    for w, p in zip(weights, percents)]

输出:

[[2, 3, 4, 4, 4, 4, 5, 5],
 [2, 4, 4, 4, 5, 5, 5, 5, 5]]

【讨论】:

  • 这太好了,非常感谢!
猜你喜欢
  • 1970-01-01
  • 2015-02-10
  • 1970-01-01
  • 2016-04-25
  • 2013-05-02
  • 2015-05-30
  • 1970-01-01
  • 1970-01-01
  • 2017-10-09
相关资源
最近更新 更多