【问题标题】:String seeming to be one single space character, but isn't字符串似乎是一个空格字符,但不是
【发布时间】:2020-04-17 00:26:54
【问题描述】:

我正在使用rvest 进行一些网络抓取,但遇到了一些奇怪的事情。有一个字符串看起来像 " " 但不是。我已经在两台计算机上重现了这一点,一台运行 R 3.6.3 的 Mac OSX 系统和一台运行 R 3.6.3 的 Windows 10 系统。

library(rvest)
library(stringr)
# scrape website, no issue
webpage <- rvest::read_html("https://www.usms.org/longdist/ldnats00/1hrf4044.php")
html <- rvest::html_nodes(webpage, css = "td")
results <- rvest::html_text(html)
# cleaning results a bit, no issue
results <- stringr::str_replace(results, "\\\r\\\n", "")
results <- results[results != ""]
# the mystery string
results[605]
[1] " "

如果我将results[605] 与" " 进行比较,或者与打印results[605] 的复制粘贴结果进行比较

results[605] == " "
[1] FALSE

如果我将results[605] 存储在一个值中

string_605 <- results[605]
string_605
[1] " "
results[605] == string_605
[1] TRUE
string_605 == " "
[1] FALSE

作为健全性检查

" " == " "
[1] TRUE

这个神秘的字符串是什么,我该如何匹配它?我想像results &lt;- results[results != mystery string]一样摆脱它

【问题讨论】:

  • 问题可能出在反斜杠上。
  • 也许是标签? \t
  • 不是标签,因为 str_detect(results[605], "\\t") 返回 FALSE 虽然有趣的是 str_detect(results[605], "\\s") 返回 TRUE 所以它是某种空格

标签: r string rvest


【解决方案1】:

这里的字符串是&lt;U+00A0&gt;

我的解决方案总是尝试clipr::write_clip(results[605]) 并粘贴到任何地方。然后你可以看到这个字符串的代码也可以粘贴到谷歌搜索它:)

之后你可以这样做results &lt;- results[results != '\U00A0']

【讨论】:

  • clipr::write_clip(results[605]) 确实有效,但您是如何确定该字符名为“\U00A0”的?当我粘贴 write_clip 的结果时,我得到了一个看起来像空格的东西,但它确实匹配 results[605]
  • Frank 没有说,因为提供外部网站的链接是违反 SO 政策的。 babelstone.co.uk/Unicode/whatisit.html
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2013-05-13
  • 2016-03-26
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2015-03-13
相关资源
最近更新 更多