【问题标题】:Complicated (for me) reshaping from wide to long in PandasPandas 从宽到长的复杂(对我而言)重塑
【发布时间】:2013-07-15 07:39:58
【问题描述】:

个人(索引从 0 到 5)在两个位置之间进行选择:A 和 B。 我的数据具有广泛的格式,其中包含因个体而异的特征 (ind_var) 和仅因位置而异的特征 (location_var)。

例如,我有:

In [281]:

df_reshape_test = pd.DataFrame( {'location' : ['A', 'A', 'A', 'B', 'B', 'B'], 'dist_to_A' : [0, 0, 0, 50, 50, 50], 'dist_to_B' : [50, 50, 50, 0, 0, 0], 'location_var': [10, 10, 10, 14, 14, 14], 'ind_var': [3, 8, 10, 1, 3, 4]})

df_reshape_test

Out[281]:
    dist_to_A   dist_to_B   ind_var location location_var
0    0            50             3   A       10
1    0            50             8   A       10
2    0            50            10   A       10
3    50           0              1   B       14
4    50           0              3   B       14
5    50           0              4   B       14

变量“位置”是个人选择的变量。 dist_to_A 是从个人选择的位置到位置 A 的距离(与 dist_to_B 相同)

我希望我的数据具有这种形式:

    choice  dist_S  ind_var location    location_var
0    1        0       3         A           10
0    0       50       3         B           14
1    1        0       8         A           10
1    0       50       8         B           14
2    1        0      10         A           10
2    0       50      10         B           14
3    0       50       1         A           10
3    1        0       1         B           14
4    0       50       3         A           10
4    1        0       3         B           14
5    0       50       4         A           10
5    1        0       4         B           14

其中choice == 1 表示个人已选择该位置,dist_S 是与所选位置的距离。

我读到了.stack 方法,但不知道如何将它应用于这种情况。 感谢您的宝贵时间!

注意:这只是一个简单的例子。我正在查找的数据集具有不同数量的位置和每个位置的个人数量,因此如果可能,我正在寻找一个灵活的解决方案

【问题讨论】:

    标签: python pandas reshape


    【解决方案1】:

    其实pandas有一个wide_to_long命令可以方便的做你想做的事。

    df = pd.DataFrame( {'location' : ['A', 'A', 'A', 'B', 'B', 'B'], 
                    'dist_to_A' : [0, 0, 0, 50, 50, 50], 
                    'dist_to_B' : [50, 50, 50, 0, 0, 0], 
                    'location_var': [10, 10, 10, 14, 14, 14], 
                    'ind_var': [3, 8, 10, 1, 3, 4]})
    
    df['ind'] = df.index
    
    #The `location` and `location_var` corresponds to the choices, 
    #record them as dictionaries and drop them 
    #(Just realized you had a cleaner way, copied from yous). 
    
    ind_to_loc = dict(df['location'])
    loc_dict = dict(df.groupby('location').agg(lambda x : int(np.mean(x)))['location_var'])
    df.drop(['location_var', 'location'], axis = 1, inplace = True)
    # now reshape
    df_long = pd.wide_to_long(df, ['dist_to_'], i = 'ind', j = 'location') 
    
    # use the dictionaries to get variables `choice` and `location_var` back.
    
    df_long['choice'] = df_long.index.map(lambda x: ind_to_loc[x[0]])
    df_long['location_var'] = df_long.index.map(lambda x : loc_dict[x[1]])
    print df_long.sort()
    

    这将为您提供您要求的表格:

                  ind_var  dist_to_ choice  location_var
    ind location                                        
    0   A               3         0      A            10
        B               3        50      A            14
    1   A               8         0      A            10
        B               8        50      A            14
    2   A              10         0      A            10
        B              10        50      A            14
    3   A               1        50      B            10
        B               1         0      B            14
    4   A               3        50      B            10
        B               3         0      B            14
    5   A               4        50      B            10
        B               4         0      B            14
    

    当然,如果你想要的话,你可以生成一个接受 0 和 1 的选择变量。

    【讨论】:

    • 感谢Zhen Sun,这看起来是一种更简洁的方法。但是,它与我的问题不符。 location 变量没有错误。 dist_to_variable 应该具有个体(由索引给出)与位置之间的距离。选择表明个人在这个机会中选择了什么(她选择了位置 A 还是 B)。我试图在我的问题中说明这一点,但如果仍然不清楚,请告诉我,我可以重写一下。
    • @cd98,我想我明白你的问题了。我是说在您的长表中,location 变量看起来不太正确。 ID = 0 怎么可能同时拥有 location A 和 B? ID = 0 位于 A 基于您的宽表。此外,我不太明白为什么该方法与您的问题不匹配。
    • ID 标识个人。每个人在选项 A 或 B 之间进行选择。原始数据只有个人选择的数据。我正在寻找的“长”格式既有每个人的选择,也有选择的特征,包括仅取决于个人 (ind_var)、仅取决于位置 (location var) 或同时取决于个人和选择的特征(例如作为位置)。每个工业。只选择一个位置。这是一种奇怪的格式,但这是我的统计程序所需要的。抱歉,如果不清楚,希望是现在!感谢您的帮助!
    • @cd98,现在我明白了!数据确实以一种非常奇怪的方式存储。我认为location var 是个人的位置。事实上,这是选择的位置,显然,各个位置不在数据中。尽管如此,请查看编辑后的代码。在重塑之前,我制作了一个字典来创建个人和location 变量之间的对应关系。
    • 干得好。我接受了你的回答,因为它比我的干净得多。
    【解决方案2】:

    我有点好奇你为什么喜欢这种格式。可能有更好的方法来存储您的数据。但是这里有。

    In [137]: import numpy as np
    
    In [138]: import pandas as pd
    
    In [139]: df_reshape_test = pd.DataFrame( {'location' : ['A', 'A', 'A', 'B', 'B
    ', 'B'], 'dist_to_A' : [0, 0, 0, 50, 50, 50], 'dist_to_B' : [50, 50, 50, 0, 0, 
    0], 'location_var': [10, 10, 10, 14, 14, 14], 'ind_var': [3, 8, 10, 1, 3, 4]})
    
    In [140]: print(df_reshape_test)
       dist_to_A  dist_to_B  ind_var location  location_var
    0          0         50        3        A            10
    1          0         50        8        A            10
    2          0         50       10        A            10
    3         50          0        1        B            14
    4         50          0        3        B            14
    5         50          0        4        B            14
    
    In [141]: # Get the new axis separately:
    
    In [142]: idx = pd.Index(df_reshape_test.index.tolist() * 2)
    
    In [143]: df2 = df_reshape_test[['ind_var', 'location', 'location_var']].reindex(idx)
    
    In [144]: print(df2)
       ind_var location  location_var
    0        3        A            10
    1        8        A            10
    2       10        A            10
    3        1        B            14
    4        3        B            14
    5        4        B            14
    0        3        A            10
    1        8        A            10
    2       10        A            10
    3        1        B            14
    4        3        B            14
    5        4        B            14
    
    In [145]: # Swap the location for the second half
    
    In [146]: # replace any 6 with len(df) / 2 + 1 if you have more rows.d 
    
    In [147]: df2['choice'] = [1] * 6 + [0] * 6  # may need to play with this.
    
    In [148]: df2.iloc[6:].location.replace({'A': 'B', 'B': 'A'}, inplace=True)
    
    In [149]: df2 = df2.sort()
    
    In [150]: df2['dist_S'] = np.abs((df2.choice - 1) * 50)
    
    In [151]: print(df2)
       ind_var location  location_var  choice  dist_S
    0        3        A            10       1       0
    0        3        B            10       0      50
    1        8        A            10       1       0
    1        8        B            10       0      50
    2       10        A            10       1       0
    2       10        B            10       0      50
    3        1        B            14       1       0
    3        1        A            14       0      50
    4        3        B            14       1       0
    4        3        A            14       0      50
    5        4        B            14       1       0
    5        4        A            14       0      50
    

    它不会很好地概括,但可能有替代(更好)的方法来绕过更丑陋的部分,比如生成选择 col。

    【讨论】:

    • 感谢您的回复!我同意这是一种奇怪的格式,我只需要它,因为 Stata 的 asclogit 条件 logit 命令(即替代变化特征)需要具有这种形状的数据集。我今天会尝试你的解决方案,也许其他一些海报会加入。
    • 很公平。你看过statsmodels吗?这是一个用于计量经济学/统计推断的 python 包。我不确定是否有人实现了条件 logit,但它涵盖了所有基础知识。
    • 我查看了statsmodels,但到目前为止,它们只有多项 Logit(个体变化特征)。但是,他们现在似乎正在研究条件 Logit(请参阅this blog)
    • 我尝试将 TomAugspurger 代码应用到我的真实数据集(它有大约 60 个不同的选择位置,不仅是 A 和 B,而且每个选择的位置有不同数量的人),我还没有弄清楚如何让它在我的情况下工作。我正在研究针对个人和选择的多重索引,看看是否是这样。
    • 嗯,很遗憾听到它不起作用。让我知道是否还有其他问题。顺便说一句,pd.get_dummies() 功能可能会为您提供这么多位置。
    【解决方案3】:

    好的,这比我预期的要长,但这里有一个更通用的答案,适用于每个人的任意数量的选择。我确信有更简单的方法,所以如果有人可以为以下一些代码提供更好的方法,那就太好了。

    df = pd.DataFrame( {'location' : ['A', 'A', 'A', 'B', 'B', 'B'], 'dist_to_A' : [0, 0, 0, 50, 50, 50], 'dist_to_B' : [50, 50, 50, 0, 0, 0], 'location_var': [10, 10, 10, 14, 14, 14], 'ind_var': [3, 8, 10, 1, 3, 4]})
    

    给了

        dist_to_A   dist_to_B   ind_var location   location_var
    0    0           50          3     A            10
    1    0           50          8     A            10
    2    0           50         10     A            10
    3    50          0           1     B            14
    4    50          0           3     B            14
    5    50          0           4     B            14
    

    然后我们这样做:

    df.index.names = ['ind']
    
    # Add choice var
    
    df['choice'] = 1
    
    # Create dictionaries we'll use later
    
    ind_to_loc = dict(df['location'])
    # gives ind_to_loc equal to {0 : 'A', 1 : 'A', 2 : 'A', 3 : 'B', 4 : 'B', 5: 'B'}
    
    ind_dict = dict(df['ind_var'])
    #gives  { 0: 3, 1 : 8, 2 : 10, 3: 1, 4 : 3, 5: 4}
    
    loc_dict = dict(  df.groupby('location').agg(lambda x : int(np.mean(x)) )['location_var']  )
    # gives  {'A' : 10, 'B' : 14}
    

    现在我创建一个多索引并重新索引以获得长形状

    df = df.set_index( [df.index, df['location']] )
    
    df.index.names = ['ind', 'location']
    
    # re-index to long shape
    
    loc_list = ['A', 'B']
    ind_list = [0, 1, 2, 3, 4, 5]
    new_shape = [  (ind, loc) for ind in ind_list for loc in loc_list]
    idx = pd.Index(new_shape)
    df_long = df.reindex(idx, method = None)
    df_long.index.names = ['ind', 'loc']
    

    长形是这样的:

             dist_to_A  dist_to_B  ind_var location  location_var  choice
    ind loc                                                              
    0   A            0         50        3        A            10       1
        B          NaN        NaN      NaN      NaN           NaN     NaN
    1   A            0         50        8        A            10       1
        B          NaN        NaN      NaN      NaN           NaN     NaN
    2   A            0         50       10        A            10       1
        B          NaN        NaN      NaN      NaN           NaN     NaN
    3   A          NaN        NaN      NaN      NaN           NaN     NaN
        B           50          0        1        B            14       1
    4   A          NaN        NaN      NaN      NaN           NaN     NaN
        B           50          0        3        B            14       1
    5   A          NaN        NaN      NaN      NaN           NaN     NaN
        B           50          0        4        B            14       1
    

    所以现在用字典填充 NaN 值:

    df_long['ind_var'] = df_long.index.map(lambda x : ind_dict[x[0]] )
    df_long['location']  = df_long.index.map(lambda x : ind_to_loc[x[0]] )
    df_long['location_var'] = df_long.index.map(lambda x : loc_dict[x[1]] )
    
    # Fill in choice
    df_long['choice'] = df_long['choice'].fillna(0)
    

    最后,剩下的就是创建 dist_S
    我会在这里作弊并假设我可以创建一个像这样的嵌套字典

    nested_loc = {'A' : {'A' : 0, 'B' : 50}, 'B' : {'A' : 50, 'B' : 0}}
    

    (内容为:如果您在位置 A,则位置 A 位于 0 公里处,位置 B 位于 50 公里处)

    def nested_f(x):    
        return nested_loc[x[0]][x[1]]
    
    df_long = df_long.reset_index()
    df_long['dist_S'] = df_long[['loc', 'location']].apply(nested_f, axis=1)
    
    df_long = df_long.drop(['dist_to_A', 'dist_to_B', 'location'], axis = 1 )
    
    df_long
    

    给出想要的结果

        ind loc ind_var location_var    choice  dist_S
    0    0   A   3         10            1      0
    1    0   B   3         14            0      50
    2    1   A   8         10            1      0
    3    1   B   8         14            0      50
    4    2   A   10        10            1      0
    5    2   B   10        14            0      50
    6    3   A   1         10            0      50
    7    3   B   1         14            1      0
    8    4   A   3         10            0      50
    9    4   B   3         14            1      0
    10   5   A   4         10            0      50
    11   5   B   4         14            1      0
    

    【讨论】:

      猜你喜欢
      • 2021-06-25
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-12-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多