【问题标题】:Split vector in MATLABMATLAB中的分割向量
【发布时间】:2015-07-03 19:07:29
【问题描述】:

我正在尝试优雅地拆分矢量。例如,

vec = [1 2 3 4 5 6 7 8 9 10]

根据另一个相同长度的 0 和 1 向量,其中 1 表示向量应该在哪里分割 - 或者更确切地说是切割:

cut = [0 0 0 1 0 0 0 0 1 0]

给我们一个类似于以下的单元格输出:

[1 2 3] [5 6 7 8] [10]

【问题讨论】:

  • 虽然这些答案是有效的,但也许如果你给我们你想要完成的原始任务,会有比这些更好的方法。
  • +1 公平点。你猜对了,我正在做一些更具体的事情:)。尽管如此,我还是尝试将我的问题提炼成一个足够笼统的问题,其他人可能会觉得有用。
  • 我看到了,这就是我为你的问题 +1 的原因。但是,我认为这对很多人来说并不是很有用,哈哈
  • 好问题。我们中的一些人回答它很开心。
  • 这是一个很好的问题。谢谢你问。尽管我的解决方案可能不如 Tony 或 Divakar 的有效,但我想写一些不依赖于cumsum 的东西。玩得开心!

标签: arrays matlab vector


【解决方案1】:

不幸的是,MATLAB 中没有“反向连接”。如果你想解决这样的问题,你可以试试下面的代码。在您有两个分割点以在最后产生三个向量的情况下,它将为您提供所需的内容。如果你想要更多的拆分,你需要在循环之后修改代码。

结果是 n 向量形式。要将它们变成单元格,请在结果上使用 num2cell。

pos_of_one = 0;

% The loop finds the split points and puts their positions into a vector.
for kk = 1 : length(cut)
    if cut(1,kk) == 1
        pos_of_one = pos_of_one + 1;
        A(1,one_pos) = kk;
    end
end

F = vec(1 : A(1,1) - 1);
G = vec(A(1,1) + 1 : A(1,2) - 1);
H = vec(A(1,2) + 1 : end);

【讨论】:

    【解决方案2】:

    这是你需要的:

    function spl  = Splitting(vec,cut)
    n=1;
    j=1;
    for i=1:1:length(b)
        if cut(i)==0 
            spl{n}(j)=vec(i);
            j=j+1;
        else 
            n=n+1;
            j=1;
        end
    end
    end
    

    尽管我的方法很简单,但它在性能方面排名第二:

    -------------------- With CUMSUM + ACCUMARRAY
    Elapsed time is 0.264428 seconds.
    -------------------- With FIND + ARRAYFUN
    Elapsed time is 0.407963 seconds.
    -------------------- With CUMSUM + ARRAYFUN
    Elapsed time is 18.337940 seconds.
    -------------------- SIMPLE
    Elapsed time is 0.271942 seconds.
    

    【讨论】:

      【解决方案3】:

      对于这个问题,一个方便的函数是cumsum,它可以创建切割数组的累积和。产生输出元胞数组的代码如下:

      vec = [1 2 3 4 5 6 7 8 9 10];
      cut = [0 0 0 1 0 0 0 0 1 0];
      
      cutsum = cumsum(cut);
      cutsum(cut == 1) = NaN;  %Don't include the cut indices themselves
      sumvals = unique(cutsum);      % Find the values to use in indexing vec for the output
      sumvals(isnan(sumvals)) = [];  %Remove NaN values from sumvals
      output = {};
      for i=1:numel(sumvals)
          output{i} = vec(cutsum == sumvals(i)); %#ok<SAGROW>
      end
      

      正如另一个答案所示,您可以使用arrayfun 创建一个包含结果的元胞数组。要在此处应用它,您需要将 for 循环(以及输出的初始化)替换为以下行:

      output = arrayfun(@(val) vec(cutsum == val), sumvals, 'UniformOutput', 0);
      

      这很好,因为它最终不会增加输出元胞数组。

      这个例程的关键特性是变量cutsum,它最终看起来像这样:

      cutsum =
           0     0     0   NaN     1     1     1     1   NaN     2
      

      然后我们需要做的就是使用它来创建索引以将数据从原始vec 数组中提取出来。我们从零循环到最大值并提取匹配值。请注意,此例程处理一些可能出现的情况。例如,它在cut 数组的开头和结尾处处理 1 值,并优雅地处理cut 数组中的重复值,而不会在输出中创建空数组。这是因为使用unique 来创建要在cutsum 中搜索的值集,并且我们在sumvals 数组中丢弃了NaN 值。

      您可以使用 -1 而不是 NaN 作为不使用剪切位置的信号标志,但我喜欢 NaN 以提高可读性。 -1 值可能会更有效,因为您所要做的就是截断 sumvals 数组中的第一个元素。使用 NaN 作为信号标志只是我的偏好。

      这个的输出是一个带有结果的元胞数组:

      output{1} =
           1     2     3
      output{2} =
           5     6     7     8
      output{3} =
          10
      

      我们需要处理一些奇怪的情况。考虑情况:

      vec = [1 2 3 4 5 6 7 8 9 10 11 12 13 14];
      cut = [1 0 0 1 1 0 0 0 0 1  0  0  0  1];
      

      其中有重复的 1,以及开头和结尾的 1。这个例程可以正确处理所有这些,没有任何空集:

      output{1} = 
           2     3
      output{2} =
           6     7     8     9
      output{3} = 
          11    12    13
      

      【讨论】:

      • 目前看来是最短的方法,+1
      【解决方案4】:

      您可以结合使用find 和arrayfun:

      vec = [1 2 3 4 5 6 7 8 9 10];
      N = numel(vec);
      cut = [0 0 0 1 0 0 0 0 1 0];
      ind = find(cut);
      ind_before = [ind-1 N]; ind_before(ind_before < 1) = 1;
      ind_after = [1 ind+1]; ind_after(ind_after > N) = N;
      out = arrayfun(@(x,y) vec(x:y), ind_after, ind_before, 'uni', 0);
      

      因此我们得到:

      >> celldisp(out)
      
      out{1} =
      
           1     2     3         
      
      out{2} =
      
           5     6     7     8    
      
      out{3} =
      
          10
      

      那么这是如何工作的呢?嗯,第一行定义你的输入向量,第二行找出这个向量中有多少元素,第三行表示你的cut 向量,它定义了我们需要在向量中剪切的位置。接下来,我们使用find 来确定@​​987654329@ 中与向量中的分割点相对应的非零位置。如果您注意到,分割点决定了我们需要在哪里停止收集元素并开始收集元素。

      但是,我们需要考虑向量的开头和结尾。 ind_after 告诉我们需要开始收集值的位置,ind_before 告诉我们需要停止收集值的位置。要计算这些开始和结束位置,您只需将find 的结果分别加和减1。

      ind_after 和ind_before 中的每个对应位置都告诉我们需要在哪里开始和停止一起收集值。为了适应向量的开头,ind_after 需要在开头插入索引 1,因为索引 1 是我们应该从开头开始收集值的位置。同样,N 需要插入到 ind_before 的末尾,因为这是我们需要在数组末尾停止收集值的地方。

      现在对于ind_after 和ind_before,存在一种退化情况,即切割点可能位于向量的末尾或开头。如果是这种情况,那么减或加 1 将生成超出范围的开始和停止位置。我们在第 4 行和第 5 行代码中检查这一点,然后根据我们是在数组的开头还是结尾,简单地将它们设置为 1 或 N。

      最后一行代码使用arrayfun 并遍历每对ind_after 和ind_before 以切入我们的向量。每个结果都放入一个元胞数组中,然后是我们的输出。


      我们可以通过在cut 的开头和结尾放置一个 1 以及介于两者之间的一些值来检查退化情况:

      vec = [1 2 3 4 5 6 7 8 9 10];
      cut = [1 0 0 1 0 0 0 1 0 1];
      

      使用这个例子和上面的代码,我们得到:

      >> celldisp(out)
      
      out{1} =
      
           1
      
      out{2} =
      
           2     3         
      
      out{3} =
      
           5     6     7
      
      out{4} =
      
           9         
      
      out{5} =
      
          10
      

      【讨论】:

      • 漂亮。如果 cut 的第一个元素是 1,这可以修改吗?我想在这种情况下可以只检查一个空数组,具体取决于所需的内容。
      • @Tony - 是的,我现在正在写。我只是想先输入一个答案,这样你就可以看到我在这里大声笑。
      • @Tony - 已修复。现在要写解释了。
      • 很好地应用arrayfun来创建输出!
      • @Tony - 谢谢 :) 你的cumsum 方法也不是很糟糕。 +1。
      【解决方案5】:

      解决方案代码

      您可以使用 cumsum 和 accumarray 获得有效的解决方案 -

      %// Create ID/labels for use with accumarray later on
      id = cumsum(cut)+1   
      
      %// Mask to get valid values from cut and vec corresponding to ones in cut
      mask = cut==0        
      
      %// Finally get the output with accumarray using masked IDs and vec values 
      out = accumarray(id(mask).',vec(mask).',[],@(x) {x})
      

      基准测试

      以下是在列出的三种最流行的方法中使用大量输入来解决此问题时的一些性能数据 -

      N = 100000;  %// Input Datasize
      
      vec = randi(100,1,N); %// Random inputs
      cut = randi(2,1,N)-1;
      
      disp('-------------------- With CUMSUM + ACCUMARRAY')
      tic
      id = cumsum(cut)+1;
      mask = cut==0;
      out = accumarray(id(mask).',vec(mask).',[],@(x) {x});
      toc
      
      disp('-------------------- With FIND + ARRAYFUN')
      tic
      N = numel(vec);
      ind = find(cut);
      ind_before = [ind-1 N]; ind_before(ind_before < 1) = 1;
      ind_after = [1 ind+1]; ind_after(ind_after > N) = N;
      out = arrayfun(@(x,y) vec(x:y), ind_after, ind_before, 'uni', 0);
      toc
      
      disp('-------------------- With CUMSUM + ARRAYFUN')
      tic
      cutsum = cumsum(cut);
      cutsum(cut == 1) = NaN;  %Don't include the cut indices themselves
      sumvals = unique(cutsum);      % Find the values to use in indexing vec for the output
      sumvals(isnan(sumvals)) = [];  %Remove NaN values from sumvals
      output = arrayfun(@(val) vec(cutsum == val), sumvals, 'UniformOutput', 0);
      toc
      

      运行时

      -------------------- With CUMSUM + ACCUMARRAY
      Elapsed time is 0.068102 seconds.
      -------------------- With FIND + ARRAYFUN
      Elapsed time is 0.117953 seconds.
      -------------------- With CUMSUM + ARRAYFUN
      Elapsed time is 12.560973 seconds.
      

      特殊情况:如果您可能运行了 1,您需要修改下面列出的一些内容 -

      %// Mask to get valid values from cut and vec corresponding to ones in cut
      mask = cut==0  
      
      %// Setup IDs differently this time. The idea is to have successive IDs.
      id = cumsum(cut)+1
      [~,~,id] = unique(id(mask))
            
      %// Finally get the output with accumarray using masked IDs and vec values 
      out = accumarray(id(:),vec(mask).',[],@(x) {x})
      

      在这种情况下运行示例 -

      >> vec
      vec =
           1     2     3     4     5     6     7     8     9    10
      >> cut
      cut =
           1     0     0     1     1     0     0     0     1     0
      >> celldisp(out)
      out{1} =
           2
           3
      out{2} =
           6
           7
           8
      out{3} =
          10
      

      【讨论】:

      • 手掌。为什么我没有想到使用accumarray?
      • 嘿。我也是。好的! +1
      • @rayryeng 我也对此感到惊讶,所以我不得不在这个页面上检查几次,如果有任何 accumarray 解决方案! :)
      • @SanthanSalai 感谢您检查并找到此类案例。稍后删除空单元格将不会有效。因此,处理 ID 必须是有效的方法。编辑代码。
      • 很好地使用accumarray!
      【解决方案6】:

      另一种方式,但这次没有任何循环或累积......

      lengths = diff(find([1 cut 1])) - 1;    % assuming a row vector
      lengths = lengths(lengths > 0);
      data = vec(~cut);
      result = mat2cell(data, 1, lengths);    % also assuming a row vector
      

      diff(find(...)) 构造为我们提供了从每个标记到下一个标记的距离 - 我们将边界标记附加到 [1 cut 1] 以捕获任何触及末端的零运行。但是,每个长度都包含其标记,因此我们减去 1 来说明这一点,并删除任何刚好覆盖连续标记的部分,这样我们就不会在输出中得到任何不需要的空单元格。

      对于数据,我们屏蔽掉与标记相对应的所有元素,因此我们只有想要划分的有效部分。最后,准备好拆分的数据和拆分的长度,这正是mat2cell 的用途。

      另外,使用@Divakar's benchmark code;

      -------------------- With CUMSUM + ACCUMARRAY
      Elapsed time is 0.272810 seconds.
      -------------------- With FIND + ARRAYFUN
      Elapsed time is 0.436276 seconds.
      -------------------- With CUMSUM + ARRAYFUN
      Elapsed time is 17.112259 seconds.
      -------------------- With mat2cell
      Elapsed time is 0.084207 seconds.
      

      ...只是说';)

      【讨论】:

      • 这也是我的方法!除了我没想到lengths = lengths(lengths &gt; 0);很好看!
      猜你喜欢
      • 1970-01-01
      • 2014-07-09
      • 1970-01-01
      • 2021-12-16
      • 1970-01-01
      • 2014-08-13
      • 2015-03-01
      • 1970-01-01
      相关资源
      最近更新 更多