【问题标题】:Add columns to a pivot table (pandas)将列添加到数据透视表(熊猫)
【发布时间】:2016-07-27 20:40:46
【问题描述】:

我知道在 R 中我可以将 tidyr 用于以下用途:

data_wide <- spread(data_protein, Fraction, Count)

data_wide 将继承 data_protein 中所有未传播的列。

Protein Peptide  Start  Fraction  Count
1             A    122       F1     1
1             A    122       F2     2     
1             B    230       F1     3     
1             B    230       F2     4

变成

Protein Peptide  Start  F1  F2
1             A    122   1  2
1             B    230   3  4     

但在熊猫(Python)中,

data_wide = data_prot2.reset_index(drop=True).pivot('Peptide','Fraction','Count').fillna(0)

不继承函数中未指定的任何内容(索引、键、值)。 因此,我决定通过 df.join() 加入它:

data_wide2 = data_wide.join(data_prot2.set_index('Peptide')['Start']).sort_values('Start')

但这会产生重复的肽段,因为有多个起始值。有没有更直接的方法来解决这个问题?或者省略重复的连接的特殊参数?提前谢谢你。

【问题讨论】:

    标签: python pandas dataframe pivot-table


    【解决方案1】:

    试试这个:

    In [144]: df
    Out[144]:
       Protein Peptide  Start Fraction  Count
    0        1       A    122       F1      1
    1        1       A    122       F2      2
    2        1       B    230       F1      3
    3        1       B    230       F2      4
    
    In [145]: df.pivot_table(index=['Protein','Peptide','Start'], columns='Fraction').reset_index()
    Out[145]:
             Protein Peptide Start Count
    Fraction                          F1 F2
    0              1       A   122     1  2
    1              1       B   230     3  4
    

    您也可以明确指定Count 列:

    In [146]: df.pivot_table(index=['Protein','Peptide','Start'], columns='Fraction', values='Count').reset_index()
    Out[146]:
    Fraction  Protein Peptide  Start  F1  F2
    0               1       A    122   1   2
    1               1       B    230   3   4
    

    【讨论】:

    • 我应该在哪一步执行此操作?我什么时候指定 Count 是 Fraction 的值?
    【解决方案2】:

    使用stack:

    df.set_index(df.columns[:4].tolist()) \
      .Count.unstack().reset_index() \
      .rename_axis(None, axis=1)
    

    【讨论】:

      【解决方案3】:

      spread 在tidyr 中被pivot_wider 取代。

      使用遵循tidyr 的API 设计的datar 怎么样:

      >>> from datar.all import f, tribble, pivot_wider
      >>> data_protein = tribble(
      ...     f.Protein, f.Peptide,  f.Start,  f.Fraction,  f.Count,
      ...     1,         "A",        122,      "F1",        1,
      ...     1,         "A",        122,      "F2",        2,     
      ...     1,         "B",        230,      "F1",        3,     
      ...     1,         "B",        230,      "F2",        4,
      ... )
      >>> data_wide = pivot_wider(data_protein, names_from=f.Fraction, values_from=f.Count)
      >>> data_wide
        Peptide  Protein  Start  F1  F2
      0       A        1    122   1   2
      1       B        1    230   3   4
      

      我是包的作者。如果您有任何问题,请随时提交问题。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2021-06-09
        • 2019-04-28
        • 2021-02-28
        • 2023-01-11
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多