【问题标题】:custom dictionaries in quantedaquanteda 中的自定义词典
【发布时间】:2019-01-19 19:59:02
【问题描述】:

我需要做 LIWC(语言查询和字数统计),我正在使用 quanteda/quanteda.dictionaries。我需要“加载”自定义词典:我将我的单词列表保存为单独的 .txt 文件,并通过 readlines“加载”(例如只有一本词典):

autonomy = readLines("Dictionary/autonomy.txt", encoding = "UTF-8")

EODic<-quanteda::dictionary(list(autonomy=autonomy),encoding = "auto")

这是我正在尝试的文本

txt <- c("12th Battalion Productions is producing a fully holographic feature length production. Presenting a 3D audio-visual projection without a single cast member present, to give the illusion of live stage performance.")

然后我运行它:

liwcalike(txt, EODic, what = "word")

并得到这个错误:

Error in stri_replace_all_charclass(value, "\\p{Z}", concatenator) : 


invalid UTF-8 byte sequence detected; perhaps you should try calling stri_enc_toutf8()

显然,问题出在我的 txt 文件上。我有很多字典,宁愿将它们作为文件加载。

如何解决此错误?在 readlines 中指定编码似乎没有帮助

这里是文件https://drive.google.com/file/d/12plgfJdMawmqTkcLWxD1BfWdaeHuPTXV/view?usp=sharing

更新:在 Mac 上解决此问题的最简单方法是在 Word 而不是 TextEdit 中打开 .txt 文件。 Word 提供了与默认 TextEdit 不同的编码选项!

【问题讨论】:

  • 如果没有可重现的文件示例,就不可能知道错误的来源。但听起来你的输入单词列表没有编码为 UTF-8。 readlines() 中的“encoding”参数不会重新编码文件,它只会告诉 R 将文本视为 UTF-8。我的建议是在文本编辑器中打开文件并将其明确保存为 UTF-8。或者,提供文件链接以使问题可重现。
  • 感谢 Ken,在那里添加了一个链接,我在 Mac 上,当我在 TextEdit 中打开并保存它时,它没有给我编码选项

标签: text encoding quanteda


【解决方案1】:

好的,问题不在于编码问题,因为您链接的文件中的所有内容都可以完全用低 128 字符 ASCII 编码。问题是由空行引起的空白。还有一些需要删除的领先空间。使用一些子集和一些 stringi 清理操作很容易做到这一点。

library("quanteda")
## Package version: 1.3.14

autonomy <- readLines("~/Downloads/risktaking.txt", encoding = "UTF-8")
head(autonomy, 15)
##  [1] "adventuresome"  " adventurous"   " audacious"     " bet"          
##  [5] " bold"          " bold-spirited" " brash"         " brave"        
##  [9] " chance"        " chancy"        " courageous"    " danger"       
## [13] ""               "dangerous"      " dare"

# strip leading or trailing whitespace
autonomy <- stringi::stri_trim_both(autonomy)
# get rid of empties
autonomy <- autonomy[!autonomy == ""]

现在您可以创建字典并应用quanteda.dictionaries::liwcalike() 函数。

# now define the quanteda dictionary
EODic <- dictionary(list(autonomy = autonomy))

txt <- c("12th Battalion Productions is producing a fully holographic feature length production. Presenting a 3D audio-visual projection without a single cast member present, to give the illusion of live stage performance.")

library("quanteda.dictionaries")
liwcalike(txt, dictionary = EODic)
##   docname Segment WC  WPS Sixltr Dic autonomy AllPunc Period Comma Colon
## 1   text1       1 35 15.5  34.29   0        0   11.43   5.71  2.86     0
##   SemiC QMark Exclam Dash Quote Apostro Parenth OtherP
## 1     0     0      0 2.86     0       0       0   8.57

【讨论】:

  • 肯,谢谢。我确实工作了,有趣的是,保存单词也可以工作。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2023-03-10
  • 2013-02-07
  • 1970-01-01
相关资源
最近更新 更多