【发布时间】:2019-01-19 19:59:02
【问题描述】:
我需要做 LIWC(语言查询和字数统计),我正在使用 quanteda/quanteda.dictionaries。我需要“加载”自定义词典:我将我的单词列表保存为单独的 .txt 文件,并通过 readlines“加载”(例如只有一本词典):
autonomy = readLines("Dictionary/autonomy.txt", encoding = "UTF-8")
EODic<-quanteda::dictionary(list(autonomy=autonomy),encoding = "auto")
这是我正在尝试的文本
txt <- c("12th Battalion Productions is producing a fully holographic feature length production. Presenting a 3D audio-visual projection without a single cast member present, to give the illusion of live stage performance.")
然后我运行它:
liwcalike(txt, EODic, what = "word")
并得到这个错误:
Error in stri_replace_all_charclass(value, "\\p{Z}", concatenator) :
invalid UTF-8 byte sequence detected; perhaps you should try calling stri_enc_toutf8()
显然,问题出在我的 txt 文件上。我有很多字典,宁愿将它们作为文件加载。
如何解决此错误?在 readlines 中指定编码似乎没有帮助
这里是文件https://drive.google.com/file/d/12plgfJdMawmqTkcLWxD1BfWdaeHuPTXV/view?usp=sharing
更新:在 Mac 上解决此问题的最简单方法是在 Word 而不是 TextEdit 中打开 .txt 文件。 Word 提供了与默认 TextEdit 不同的编码选项!
【问题讨论】:
-
如果没有可重现的文件示例,就不可能知道错误的来源。但听起来你的输入单词列表没有编码为 UTF-8。
readlines()中的“encoding”参数不会重新编码文件,它只会告诉 R 将文本视为 UTF-8。我的建议是在文本编辑器中打开文件并将其明确保存为 UTF-8。或者,提供文件链接以使问题可重现。 -
感谢 Ken,在那里添加了一个链接,我在 Mac 上,当我在 TextEdit 中打开并保存它时,它没有给我编码选项