【问题标题】:How to read non-english characters with read.delim in R?如何在 R 中使用 read.delim 读取非英文字符?
【发布时间】:2016-09-01 04:14:02
【问题描述】:

我有一个包含多种语言的文本文件,如何在 R 中使用read.delim 函数读取,

Encoding("file.tsv")
#[1] "unknown"

source_data = read.delim(file, header= F, fileEncoding= "windows-1252",
               sep = "\t", quote = "")
source_D[360]
#[1] "ð¿ð¾ð¸ñðº ð½ð° ññ‚ð¾ð¼ ñð°ð¹ñ‚ðµ"

但记事本中显示的source_D[360] 是'поиск на этом сайте'

【问题讨论】:

标签: r character-encoding non-english


【解决方案1】:

tidyverse 方法:

在 read_delim 中使用选项 locale。 (readr 函数有 _ 而不是 . 并且通常阅读起来更快更智能) 更多细节在这里:https://r4ds.had.co.nz/data-import.html#parsing-a-vector

source_data = read_delim(file, header= F, 
                         locale = locale(encoding = "windows-1252"),
                         sep = "\t", quote = "")

【讨论】:

    【解决方案2】:
    source_data = read.delim(file, header = F, sep = "\t", quote = "", stringsAsFactors = FALSE)
    Encoding(source_data)= "UTF-8"
    

    我试过了,如果你在 Windows 中运行 R,上面的代码对我有用。 如果你在 Unix 中运行 R,你可以使用下面的代码

    source_data = read.delim(file, header = F, fileEncoding="UTF-8", sep = "\t", quote = "", stringsAsFactors = FALSE)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2018-04-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多