【问题标题】:A proper way to create dictionary from df or the way for computing jaccard similarity从df创建字典的正确方法或计算jaccard相似度的方法
【发布时间】:2019-07-10 09:47:51
【问题描述】:

我有一个超过 8000 列的 df。每列(除了第一列)代表二进制值 0 或 1。

|Name| t1| t2| t3|...| t4|  
| ..aa.. | 0 | 0 | 1 |...| 0 |  
| ..bb.. | 0 | 0 | 0 |...| 0 |  
| ..cc.. | 1 | 0 | 0 |...| 0 |

我的目标是计算 aa,bb,cc 之间的 jaccard 索引,以便我需要存储在列表中的值,这就是我要使用字典的原因。

字典必须如下所示:

{'aa': [0,0,1,...,0], 'bb': [0,0,0,...,0],...}

当 dict key=df index and value 是表示为列表的行时,我怎样才能达到这样的结果?

【问题讨论】:

  • index 或Name 列(毕竟)?
  • @RomanPerekhrest 索引是名称列

标签: python pandas dictionary


【解决方案1】:

您可以通过压缩Name 列和数据框的其余部分并从结果元组中调用dict 构造函数来构建字典:

dict(zip(df.Name, df.loc[:,'t1':].values.tolist()))
# dict(zip(df.index, df.loc[:,'t1':].values.tolist())) # if name is the index
# {'aa': [0, 0, 1, 0], 'bb': [0, 0, 0, 0], 'cc': [1, 0, 0, 0]}

输入数据:

   Name    t1     t2     t3     t4
0   aa      0      0      1      0
1   bb      0      0      0      0
2   cc      1      0      0      0

【讨论】:

    【解决方案2】:

    另一种方法:

    {k: list(v.values()) for k, v in df.set_index('Name').to_dict('index').items()}
    

    【讨论】:

      【解决方案3】:

      将Name 设置为索引并转置然后执行.to_dict():

      df.set_index('Name').T.to_dict('list')
      

      如果 Name 是索引,就这样做:

      df.T.to_dict('list')
      

      {'aa': [0, 0, 1, 0], 'bb': [0, 0, 0, 0], 'cc': [1, 0, 0, 0]}
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2011-02-28
        • 2022-01-04
        • 1970-01-01
        • 2017-06-09
        • 2016-07-21
        • 2017-03-27
        • 2017-02-28
        • 2015-08-04
        相关资源
        最近更新 更多