【问题标题】:What does the flush function do on read.table in R?在 R 中的 read.table 上,flush 函数有什么作用?
【发布时间】:2019-12-01 14:36:55
【问题描述】:

我最近开始参加 R 讲座,目前正在研究文件扫描。在工作表上,我的一个问题是:

读取文件Table6.txt,先签出文件。请注意,信息是重复的,我们只想要第一个不重复的信息。确保这次只创建角色而不是因素。最后,我们不想要 cmets。

文件名为Table6.Txt

我设法编写了正确读取表格的代码,但答题卡在扫描函数中有一个额外的部分,上面写着flush=TRUE

我的代码是这样的:

df <- read.table("Table6.txt",skip = 1,header = TRUE,row.names = "Name",nrow
= 7,comment.char = "@",stringsAsFactors = FALSE)

答题卡显示

df <- read.table("Table6.txt",skip = 1,header = TRUE,row.names = "Name",nrow
= 7,flush = TRUE,comment.char = "@",stringsAsFactors = FALSE)

flush 函数在这里做什么?两个代码的输出给出相同的数据帧。

df <- read.table("Table6.txt",skip = 1,header = TRUE,row.names = "Name",nrow
                  = 7,flush = TRUE,comment.char = "@",stringsAsFactors = FALSE)
 df
         Age Height Weight Sex
Alex      25    177     57   F
Lilly     31    163     69   F
Mark      23    190     83   M
Oliver    52    179     75   M
Martha    76    163     70   F
Lucas     49    183     83   M
Caroline  26    164     53   F
 df <- read.table("Table6.txt",skip = 1,header = TRUE,row.names = "Name",nrow
                  = 7,comment.char = "@",stringsAsFactors = FALSE)
 df
         Age Height Weight Sex
Alex      25    177     57   F
Lilly     31    163     69   F
Mark      23    190     83   M
Oliver    52    179     75   M
Martha    76    163     70   F
Lucas     49    183     83   M
Caroline  26    164     53   F

【问题讨论】:

  • 这很有趣。 read.table 使用 scan 函数进行实际扫描。这些以太币的文档说 - will flush to the end of the line after reading the last of the fields requested. This allows putting comments after the last field. 但是这里的评论无论如何都应该被忽略。并且将 flush 设置为 False 对以太没有影响。用于扫描的 R 源代码也不是很有帮助,因为主要的扫描功能是在 C 中实现的
  • 所以,我在想的是,将comment.char 设置为“@”已经使程序将表中的cmets 识别为不必要的,因此我不需要在这里进行flush 争论。正如您所说,将其设置为 false 没有任何效果,这似乎是一个双重检查程序,以确保代码对我来说运行顺利。老实说,我对 read.table 的帮助页面也不太了解。我想我会找到解决方案的制定者,直接问他为什么写这个。感谢您的回答。
  • 如果可以的话,那就太好了。找到原因后,您可以在此处发布它作为您自己问题的答案。我会很有帮助的。
  • 如果我能做到,我一定会的。再次感谢。
  • 当您使用debug(read.table) 时,您会看到,它依赖于scan。看看?scan 和那里的例子。但我也不明白这一点。

标签: r scanning read.table


【解决方案1】:

我阅读了read.table 和scan 的文档,这是我用简单的话理解的。 flush 尝试通过忽略额外字符(如果有)来完成数据帧。

例如,让我们获取您共享的相同数据

read.table(text = 'Age Height Weight Sex
          Alex      25    177     57   F
          Lilly     31    163     69   F
          Mark      23    190     83   M
          Oliver    52    179     75   M
          Martha    76    163     70   F
          Lucas     49    183     83   M
          Caroline  26    164     53   F', header = TRUE)

这按预期工作并返回

#         Age Height Weight Sex
#Alex      25    177     57   F
#Lilly     31    163     69   F
#Mark      23    190     83   M
#Oliver    52    179     75   M
#Martha    76    163     70   F
#Lucas     49    183     83   M
#Caroline  26    164     53   F

现在让我们在末尾添加一个额外的字符。

read.table(text = "Age Height Weight Sex
          Alex      25    177     57   F
          Lilly     31    163     69   F
          Mark      23    190     83   M
          Oliver    52    179     75   M
          Martha    76    163     70   F
          Lucas     49    183     83   M
          Caroline  26    164     53   F A", header = TRUE)
                                         ^ #Notice this A

报错

扫描错误(文件 = 文件,什么 = 什么,sep = sep,quote = quote,dec = dec,: 第 7 行没有 5 个元素

这是有道理的,因为最后一行中有一个额外的字符。

我们可以加fill = TRUE

read.table(text = "Age Height Weight Sex
          Alex      25    177     57   F
          Lilly     31    163     69   F
          Mark      23    190     83   M
          Oliver    52    179     75   M
          Martha    76    163     70   F
          Lucas     49    183     83   M
          Caroline  26    164     53   F A", header = TRUE, fill = TRUE)

#         Age Height Weight Sex
#Alex      25    177     57   F
#Lilly     31    163     69   F
#Mark      23    190     83   M
#Oliver    52    179     75   M
#Martha    76    163     70   F
#Lucas     49    183     83   M
#Caroline  26    164     53   F
#A         NA     NA     NA    

这会根据列的类型填充NA 或空字符,从而在末尾添加一个额外的行。

现在如果我们添加flush = TRUE

read.table(text = "Age Height Weight Sex
          Alex      25    177     57   F
          Lilly     31    163     69   F
          Mark      23    190     83   M
          Oliver    52    179     75   M
          Martha    76    163     70   F
          Lucas     49    183     83   M
          Caroline  26    164     53   F A", header = TRUE, flush = TRUE)

#         Age Height Weight Sex
#Alex      25    177     57   F
#Lilly     31    163     69   F
#Mark      23    190     83   M
#Oliver    52    179     75   M
#Martha    76    163     70   F
#Lucas     49    183     83   M
#Caroline  26    164     53   F

它会忽略末尾的附加"A",将其视为注释并制作完整的数据框。


在您的情况下,这对最终输出没有任何影响,因为您的数据是完整的并且没有任何不完整的信息。如果您正在读取您不知道其结构的数据,您可以将此视为一种安全的编程实践。

希望这能澄清一点。


正如@Christoph 所评论的,这里有一个例子来说明comment.char 和flush 之间的区别

read.table(text = 'Age Height Weight Sex
          Alex      25    177     57   F
          Lilly     31    163     69   F
          Mark      23    190     83   M
          Oliver    52    179     75   M
          Martha    76    163     70   F
        @ Lucas     49    183     83   M 
          Caroline  26    164     53   F @', header = TRUE,flush = TRUE)

#           Age Height Weight Sex
#Alex        25    177     57   F
#Lilly       31    163     69   F
#Mark        23    190     83   M
#Oliver      52    179     75   M
#Martha      76    163     70   F
#@        Lucas     49    183  83
#Caroline    26    164     53   F


read.table(text = 'Age Height Weight Sex
          Alex      25    177     57   F
          Lilly     31    163     69   F
          Mark      23    190     83   M
          Oliver    52    179     75   M
          Martha    76    163     70   F
        @ Lucas     49    183     83   M 
          Caroline  26    164     53   F @', header = TRUE,comment.char = '@')

#         Age Height Weight Sex
#Alex      25    177     57   F
#Lilly     31    163     69   F
#Mark      23    190     83   M
#Oliver    52    179     75   M
#Martha    76    163     70   F
#Caroline  26    164     53   F

如果flush = TRUE @ 出现在倒数第二行的开头,则不会忽略最后一个字符 (M)。但是,使用comment.char,我们可以忽略文本任何部分的确切字符。

【讨论】:

  • 但是如果你把最后一行改成...Caroline 26 164 53 F @A", header = TRUE, comment.char = "@"),也可以不用flush。那么你为什么需要冲洗呢?
  • @Christoph 我在答案中添加了一个示例来解释flush 和comment.char 之间的区别。
猜你喜欢
  • 2022-01-10
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-10-29
  • 1970-01-01
  • 2012-10-26
相关资源
最近更新 更多