【问题标题】:Error: Can't cast <list_of<character>> to <character> after using new pivot_wider() function in tidyr错误:在 tidyr 中使用新的 pivot_wider() 函数后,无法将 <list_of<character>> 转换为 <character>
【发布时间】:2019-04-10 11:36:49
【问题描述】:

我有一个庞大的病理结果数据集。每个患者都有一个唯一的标识符(在这种情况下为row_id。对于每个患者,他们在特定日期采集了样本(sample_date)。他们进行的测试范围非常多样化,并且输出不一(一些带有字符串和一些数字)。此外,并不是每个患者都在每个sample_date 进行过所有测试,因此应该有不少 NA。

执行的测试名称在test_name 列中,结果在result 列中。 我想把它变成一个广泛的数据集,使用test_name 作为列标题传播result 列,但将标识符保持为row_id 和sample_date。

tidyr 中的新 pivot_wider() 函数似乎非常适合我的需求,当我运行它时,它为我提供了我需要的数据框类型(即,行仍然由 row_id 和 sample_date 标识,但现在每个 test_name 和其中的结果都有列。

这是我的数据集的一个小样本:

structure(list(row_id = 1:81, sample_date = structure(c(16444, 
16444, 16444, 16444, 16444, 16444, 16444, 16444, 16444, 16444, 
16444, 16444, 16444, 16444, 16444, 16444, 16444, 16444, 16444, 
16444, 16444, 16444, 16444, 16444, 16444, 16444, 16444, 16444, 
16444, 16447, 16447, 16447, 16447, 16447, 16447, 16447, 16447, 
16447, 16447, 16447, 16447, 16447, 16447, 16447, 16447, 16447, 
16447, 16447, 16447, 16447, 16447, 16447, 16447, 16447, 16448, 
16448, 16448, 16448, 16448, 16448, 16442, 16442, 16442, 16442, 
16442, 16442, 16442, 16442, 16442, 16442, 16442, 16442, 16442, 
16442, 16442, 16442, 16442, 16442, 16442, 16442, 16442), class = "Date"), 
    test_name = c("Epidemic Typhus Group IgG Abs", "Epidemic Typhus Group IgM Abs", 
    "Spotted Fever Group IgG Abs", "Spotted Fever Group IgM Abs", 
    "Albumin", "Alkaline phosphatase", "Alanine transaminase", 
    "Basophils", "Bilirubin (total)", "Creatinine", "C-reactive protein", 
    "Eosinophils", "Estimated GFR", "Haemoglobin (g/L)", "HCT", 
    "Potassium", "Lymphocytes", "MCHC (g/L)", "MCH", "MCV", "Monocytes", 
    "MPV", "Sodium", "Neutrophils", "Platelet count", "Red cell count", 
    "RDW", "Urea", "White cell count", "Albumin", "Alkaline phosphatase", 
    "Alanine transaminase", "Basophils", "Bilirubin (total)", 
    "Creatinine", "C-reactive protein", "Eosinophils", "Estimated GFR", 
    "Haemoglobin (g/L)", "HCT", "Potassium", "Lymphocytes", "MCHC (g/L)", 
    "MCH", "MCV", "Monocytes", "MPV", "Sodium", "Neutrophils", 
    "Platelet count", "Red cell count", "RDW", "Urea", "White cell count", 
    "Creatinine", "C-reactive protein", "Estimated GFR", "Potassium", 
    "Sodium", "Urea", "Albumin", "Alkaline phosphatase", "Alanine transaminase", 
    "APTT Ratio", "APTT", "Basophils", "Bilirubin (total)", "Creatinine", 
    "C-reactive protein", "Eosinophils", "Fibrinogen", "Estimated GFR", 
    "Haemoglobin (g/L)", "HCT", "INR", "Potassium", "Lymphocytes", 
    "MCHC (g/L)", "MCH", "MCV", "Monocytes"), result = c("Not detected", 
    "Not detected", "Not detected", "Not detected", "47", "84", 
    "29", "0.3%  0.03", "12", "98", "3.3", "1.7%  0.15", "77\r\nUnits: mL/min/1.73sqm\r\nMultiply eGFR by 1.21 for people of African\r\nCaribbean origin. Interpret with regard to UK CKD\r\nguidelines: www.renal.org/information-resources\r\nUse with caution for adjusting drug dosages -\r\ncontact clinical pharmacist for advice.", 
    "156", "0.435", "3.8", "25.7%  2.31", "359", "30.4", "84.6", 
    "7.1%  0.64", "10.1", "140", "65.2%  5.86", "240", "5.14", 
    "12.4", "3.9", "8.99", "45", "53", "41", "0.3%  0.03", "10", 
    "59", "2.0", "2.8%  0.32", ">90\r\nUnits: mL/min/1.73sqm\r\nMultiply eGFR by 1.21 for people of African\r\nCaribbean origin. Interpret with regard to UK CKD\r\nguidelines: www.renal.org/information-resources\r\nUse with caution for adjusting drug dosages -\r\ncontact clinical pharmacist for advice.", 
    "126", "0.398", "4.5", "28.7%  3.30", "317", "25.7", "81.2", 
    "5.7%  0.65", "10.8", "143", "62.5%  7.18", "411", "4.90", 
    "14.7", "3.5", "11.49", "59", "76.2", ">90\r\nUnits: mL/min/1.73sqm\r\nMultiply eGFR by 1.21 for people of African\r\nCaribbean origin. Interpret with regard to UK CKD\r\nguidelines: www.renal.org/information-resources\r\nUse with caution for adjusting drug dosages -\r\ncontact clinical pharmacist for advice.", 
    "4.2", "139", "3.4", "46", "47", "40", "1.3", "39", "0.4%  0.01", 
    "8", "74", "7.0", "0.4%  0.01", "2.50", ">90\r\nUnits: mL/min/1.73sqm\r\nMultiply eGFR by 1.21 for people of African\r\nCaribbean origin. Interpret with regard to UK CKD\r\nguidelines: www.renal.org/information-resources\r\nUse with caution for adjusting drug dosages -\r\ncontact clinical pharmacist for advice.", 
    "146", "0.441", "0.96", "4.3", "43.2%  1.14", "331", "29.1", 
    "87.8", "6.8%  0.18")), class = "data.frame", row.names = c(NA, 
-81L))

这是我使用的pivot_wider() 代码(上面调用了path_results 的数据集:

library(tidyr)

path_results_wide <- path_results %>%
  select(row_id, sample_date, test_name, result)%>%
  pivot_wider(
    id_cols = c(row_id,
                sample_date), 
    names_from = test_name, 
    values_from = result
  )

有些列应该是数字,有些应该是字符串,但pivot_wider() 已将它们全部解析为字符列表,当我尝试将它们更改为数字时,出现以下错误:

path_results_wide$Albumin <- as.numeric(path_results_wide$Albumin)

错误:无法将 > 转换为

非常欢迎任何关于我可以做些什么来解决这个问题的建议。 谢谢。

【问题讨论】:

    标签: r tidyr


    【解决方案1】:

    旧答案:

    不确定pivot_wider,但如果我没有正确理解,我认为这可能是你想要的,使用reshape2。由于每一行都是一个患者和一个日期,因此有多个 NA 值,在该日期上执行了特定的测试。

    library(reshape2)
    res <- dcast(path_results, row_id + sample_date ~ test_name)
    

    新答案:

    在另一个issue 中阅读了有关 dcast 函数的信息,我意识到我们需要添加另一个 id 列来唯一标识每个单独的行。然后我在阅读 dplyr 中的扩展函数时遇到了以下内容:

    path_results_wide <- path_results %>%
        rowid_to_column() %>%
        spread(test_name, result)
    

    【讨论】:

    • 谢谢。这绝对看起来像我需要的,当然,当我在上面发布的数据样本上尝试它时,它工作正常。当我在我的完整数据集上尝试它时(约 500 000 行导致来自 test_names 的大约 1000 个不同的列,这些列都被解析为整数(可能与结果存在于该特定单元格中的次数有关?) . 知道如何克服这个问题吗?
    • 你的意思是列名已经被解析为整数了吗?还是列中的值?也许检查您的完整数据集中的某些列是否是因子类型?
    • 谢谢。我不这么认为。完整数据集中的列都是字符,除了 sample_date。问题似乎是当演员不能变量不能识别每个输出单元格的单个观察时
    • 即。当在同一日期 (sample_date) 对患者 (row_id) 的给定测试 (test_name) 进行多次测量 (结果) 时。我认为当它无法识别每个输出单元格的单个观察值时,它会调用 fun.aggregate 并返回汇总统计信息而不是实际结果。理想情况下,我会让 dcast 为附加值创建一个附加行。不知道如何最好地克服这一点。我尝试了 fun.aggregate 的各种排列,但没有成功。
    猜你喜欢
    • 2023-04-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-12-26
    • 2018-07-28
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多