【发布时间】:2019-05-16 04:21:40
【问题描述】:
我有以下系列列表。
[LVH = 0 63 (88.73 %)
LVH = 1 6 (8.45 %)
LVH = 2 1 (1.41 %)
LVH = 3 1 (1.41 %)
dtype: object, LV diastolic dysfunction (guideline) = 0 60 (84.51 %)
LV diastolic dysfunction (guideline) = 1 8 (11.27 %)
LV diastolic dysfunction (guideline) = 4 3 (4.23 %)
dtype: object, LV diastolic dysfunction grade (formula) = 0.0 60 (84.51 %)
LV diastolic dysfunction grade (formula) = 1.0 4 (5.63 %)
LV diastolic dysfunction grade (formula) = 3.0 4 (5.63 %)
LV diastolic dysfunction grade (formula) = 4.0 3 (4.23 %)
dtype: object, LV filling pressure(formula) = 0 67 (94.37 %)
LV filling pressure(formula) = 1 4 (5.63 %)
dtype: object, cause of hospitalization = 8 2 (2.82 %)
cause of hospitalization = 1 43 (60.56 %)
cause of hospitalization = 2 21 (29.58 %)
cause of hospitalization = 3 1 (1.41 %)
cause of hospitalization = 6 4 (5.63 %)
dtype: object, simplfied cause of hospitalization = 1 43 (60.56 %)
simplfied cause of hospitalization = 2 22 (30.99 %)
simplfied cause of hospitalization = 3 4 (5.63 %)
simplfied cause of hospitalization = 5 2 (2.82 %)
dtype: object, ACC/AHA = A 10 (14.08 %)
ACC/AHA = 0 56 (78.87 %)
ACC/AHA = C 2 (2.82 %)
ACC/AHA = B 3 (4.23 %)
dtype: object, ACC-AHA -binary = 0 69 (97.18 %)
ACC-AHA -binary = 1 2 (2.82 %)
dtype: object, NYHA = I 65 (91.55 %)
NYHA = II 2 (2.82 %)
NYHA = III 4 (5.63 %)
dtype: object, NYHA-binary = 0 66 (92.96 %)
NYHA-binary = 1 5 (7.04 %)
dtype: object]
对于列表的每个元素,即系列,我需要将它们转换为具有两列的数据框。例如,它应该如下所示:
Column 1 Column 2
LVH = 0 63 (88.73 %)
LVH = 1 6 (8.45 %)
LVH = 2 1 (1.41 %)
LVH = 3 1 (1.41 %)
LV diastolic dysfunction (guideline) = 0 60 (84.51 %)
LV diastolic dysfunction (guideline) = 1 8 (84.51 %)
LV diastolic dysfunction (guideline) = 4 3 (84.51 %)
...
等等。然后将其转换为 CSV 格式供人们下载。我只使用了基本的pd.DataFrame 和pd.DataFrame.from_items。第一个将其转换为数据框,但不是我想要的方式。第二个给出错误,但我认为这无论如何都没有帮助。我该如何解决这个问题?
更新
categorical_vars_multi_class = ['LVH','LV diastolic dysfunction (guideline)','LV diastolic dysfunction grade (formula)','LV filling pressure(formula)','cause of hospitalization','simplfied cause of hospitalization','ACC/AHA','ACC-AHA -binary','NYHA','NYHA-binary']
def getMultiClassData(index,table, prop):
tab = pd.Series()
for i in range(len(table)):
tab_str = str(table[i]) + " (" + str(prop[i]) + " %)"
tab = tab.set_value(i,tab_str)
tab.index = index
return(tab)
def getMultiClassTable(data,name):
table = pd.value_counts(data[name].values, sort=False)
table.index = [name + ' = ' + str(x) for x in table.index]
prop = (table/table.sum() * 100).round(2)
return(getMultiClassData(table.index,table.values, prop))
m_cluster_1 = [getMultiClassTable(data,x) for x in categorical_vars_multi_class]
data 是一个数据框,其中包含变量的列名和测量值。数据集庞大且敏感。
【问题讨论】:
-
给出有问题的代码以生成系列列表
-
@meW,我添加了代码。
-
也定义
data -
@meW,你能详细说明一下吗?
data只是一个包含 287 列和 128000 行的数据框。提供的数据也很敏感。我只对用于查看其分布的分类数据感兴趣。还有其他几个具有类似数据的数据集。我将为每个数据框运行代码并放入数据框和(最终是 excel 电子表格),以供用户查看它们对于每个数据集组的不同(或相似性)
标签: pandas list dataframe series