【问题标题】:how to clean weird symbols from csv file in r如何从r中的csv文件中清除奇怪的符号
【发布时间】:2020-08-19 07:17:40
【问题描述】:

我有一个 csv 文件,其中包含很多奇怪的符号。示例如下:

df = data.frame(comments = c('Korea¬Ãs Ministry of Food and Drug Safety is proposing an amendment seeking to amend the Standards and Specification','it is important to highlight:\n• Many maximum limits for drug',
                            'The European Parliament has published a decision, which aims to establish a special Committee to examine the EU¬Ãs authorization procedure'))
write.csv(df, './example.csv', row.names = FALSE)

有谁知道我如何在 R(或 python)中清理那些奇怪的符号。我不知道为什么会发生这种情况以及如何清理它们。非常感谢。

【问题讨论】:

  • 你能显示你的预期输出吗
  • 其中的“dupe”组件正在清理数据,它与写入 CSV 无关。 (此问题与 CSV 无关。)
  • @akrun 我不完全确定预期的输出。这是我得到的 csv 文件的一个例子。它包含那些符号。我怀疑与编码有关。但是想不出来....
  • @r2evans 我重新打开了这个问题,因为它不清楚预期的输出
  • @zesla 如果this 解决了问题,那么它可以被复制标记

标签: r csv read.csv


【解决方案1】:

假设“奇怪”是指任何不是“正常”字母、数字、点或逗号的东西:

gsub("[^A-z0-9\\. ,]", "", df$comment)
[1] "Koreas Ministry of Food and Drug Safety is proposing an amendment seeking to amend the Standards and Specification"                     
[2] "it is important to highlight Many maximum limits for drug"                                                                               
[3] "The European Parliament has published a decision, which aims to establish a special Committee to examine the EUs authorization procedure"

从这里开始,您可以添加更多允许的符号。

【讨论】:

    猜你喜欢
    • 2015-04-15
    • 2018-02-14
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-10-31
    • 2015-02-09
    • 1970-01-01
    相关资源
    最近更新 更多