【问题标题】:How to convert list of Series to a two columns dataframe?如何将系列列表转换为两列数据框?
【发布时间】:2019-05-16 04:21:40
【问题描述】:

我有以下系列列表。

[LVH = 0    63 (88.73 %)
 LVH = 1      6 (8.45 %)
 LVH = 2      1 (1.41 %)
 LVH = 3      1 (1.41 %)
 dtype: object, LV diastolic dysfunction (guideline) = 0    60 (84.51 %)
 LV diastolic dysfunction (guideline) = 1     8 (11.27 %)
 LV diastolic dysfunction (guideline) = 4      3 (4.23 %)
 dtype: object, LV diastolic dysfunction grade (formula) = 0.0    60 (84.51 %)
 LV diastolic dysfunction grade (formula) = 1.0      4 (5.63 %)
 LV diastolic dysfunction grade (formula) = 3.0      4 (5.63 %)
 LV diastolic dysfunction grade (formula) = 4.0      3 (4.23 %)
 dtype: object, LV filling pressure(formula) = 0    67 (94.37 %)
 LV filling pressure(formula) = 1      4 (5.63 %)
 dtype: object, cause of hospitalization = 8      2 (2.82 %)
 cause of hospitalization = 1    43 (60.56 %)
 cause of hospitalization = 2    21 (29.58 %)
 cause of hospitalization = 3      1 (1.41 %)
 cause of hospitalization = 6      4 (5.63 %)
 dtype: object, simplfied cause of hospitalization = 1    43 (60.56 %)
 simplfied cause of hospitalization = 2    22 (30.99 %)
 simplfied cause of hospitalization = 3      4 (5.63 %)
 simplfied cause of hospitalization = 5      2 (2.82 %)
 dtype: object, ACC/AHA = A    10 (14.08 %)
 ACC/AHA = 0    56 (78.87 %)
 ACC/AHA = C      2 (2.82 %)
 ACC/AHA = B      3 (4.23 %)
 dtype: object, ACC-AHA -binary = 0    69 (97.18 %)
 ACC-AHA -binary = 1      2 (2.82 %)
 dtype: object, NYHA = I      65 (91.55 %)
 NYHA = II       2 (2.82 %)
 NYHA = III      4 (5.63 %)
 dtype: object, NYHA-binary = 0    66 (92.96 %)
 NYHA-binary = 1      5 (7.04 %)
 dtype: object]

对于列表的每个元素,即系列,我需要将它们转换为具有两列的数据框。例如,它应该如下所示:

Column 1                                      Column 2
LVH = 0                                       63 (88.73 %)
LVH = 1                                        6 (8.45 %)
LVH = 2                                        1 (1.41 %)    
LVH = 3                                        1 (1.41 %)
LV diastolic dysfunction (guideline) = 0      60 (84.51 %)
LV diastolic dysfunction (guideline) = 1       8 (84.51 %)
LV diastolic dysfunction (guideline) = 4       3 (84.51 %)
... 

等等。然后将其转换为 CSV 格式供人们下载。我只使用了基本的pd.DataFrame 和pd.DataFrame.from_items。第一个将其转换为数据框,但不是我想要的方式。第二个给出错误,但我认为这无论如何都没有帮助。我该如何解决这个问题?

更新

categorical_vars_multi_class = ['LVH','LV diastolic dysfunction (guideline)','LV diastolic dysfunction grade (formula)','LV filling pressure(formula)','cause of hospitalization','simplfied cause of hospitalization','ACC/AHA','ACC-AHA -binary','NYHA','NYHA-binary']

def getMultiClassData(index,table, prop):
    tab = pd.Series()
    for i in range(len(table)): 
        tab_str = str(table[i]) + " (" + str(prop[i]) + " %)"
        tab = tab.set_value(i,tab_str)
    tab.index = index
    return(tab)


def getMultiClassTable(data,name):
    table = pd.value_counts(data[name].values, sort=False)
    table.index = [name + ' = ' + str(x) for x in table.index]
    prop = (table/table.sum() * 100).round(2)

    return(getMultiClassData(table.index,table.values, prop))



m_cluster_1 = [getMultiClassTable(data,x) for x in categorical_vars_multi_class]

data 是一个数据框,其中包含变量的列名和测量值。数据集庞大且敏感。

【问题讨论】:

  • 给出有问题的代码以生成系列列表
  • @meW,我添加了代码。
  • 也定义data
  • @meW,你能详细说明一下吗? data 只是一个包含 287 列和 128000 行的数据框。提供的数据也很敏感。我只对用于查看其分布的分类数据感兴趣。还有其他几个具有类似数据的数据集。我将为每个数据框运行代码并放入数据框和(最终是 excel 电子表格),以供用户查看它们对于每个数据集组的不同(或相似性)

标签: pandas list dataframe series


【解决方案1】:

由于未知的data,我无法复制您的示例,因此我形成了自己的示例示例。您可以从中获得帮助 -

s1 = pd.Series(['1kg', '2kg'], index=['first', 'second'])
s2 = pd.Series(['3kg', '4kg'], index=['third', 'fourth'])
lst = [s1, s2]
lst

# [first     1kg
#  second    2kg
#  dtype: object, third     3kg
#  fourth    4kg
#  dtype: object]

ndf = pd.concat(lst,  axis = 1, keys=[s.name for s in lst], sort=False).fillna('').apply(lambda x: ''.join(x), axis=1)
ndf = pd.DataFrame(ndf).reset_index()
ndf.columns = ['Column 1', 'Column 2']
ndf


+---+----------+----------+
|   | Column 1 | Column 2 |
+---+----------+----------+
| 0 | first    | 1kg      |
| 1 | second   | 2kg      |
| 2 | third    | 3kg      |
| 3 | fourth   | 4kg      |
+---+----------+----------+

【讨论】:

  • 效果很好!!!!太感谢了!!!我只需要删除sort=False,因为我收到一个错误说明pd.concat has no sort function。我删除了它,它工作。我会进行排序以找出它不起作用的原因。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2019-04-18
  • 2019-11-16
  • 1970-01-01
  • 1970-01-01
  • 2020-07-29
  • 2020-04-17
相关资源
最近更新 更多