【问题标题】:Error in nchar(Terms(x), type = "chars") : invalid multibyte string, element 204, when inspecting document term matrixnchar(Terms(x), type = "chars") 中的错误:检查文档术语矩阵时,多字节字符串无效,元素 204
【发布时间】:2018-07-10 15:27:20
【问题描述】:

这是我使用的源代码:

 MyData <- Corpus(DirSource("F:/Data/CSV/Data"),readerControl = list(reader=readPlain,language="cn"))
    SegmentedData <- lapply(MyData, function(x) unlist(segmentCN(x)))
    temp <- Corpus(DataframeSource(SegmentedData), readerControl = list(reader=readPlain, language="cn"))

预处理数据

temp <- tm_map(temp, removePunctuation)
temp <- tm_map(temp,removeNumbers)
removeURL <- function(x)gsub("http[[:alnum:]]*"," ",x)
temp <- tm_map(temp, removeURL)
temp <- tm_map(temp,stripWhitespace)
dtmxi <- DocumentTermMatrix(temp)
dtmxi <- removeSparseTerms(dtmxi,0.83)

**inspect(t(dtmxi))** ---This is where I get the error

【问题讨论】:

    标签: r text-mining


    【解决方案1】:

    我相信您的文件中有一些中文字符。为了克服这个问题,也可以使用这行代码来阅读它们:

    Sys.setlocale('LC_ALL','C')
    

    【讨论】:

      【解决方案2】:

      我的RStudio 在设置Sys.setlocale( 'LC_ALL','C' ) 并运行TermDocumentMatrix( mycorpus ) 函数后重新启动会话。

      【讨论】:

        【解决方案3】:

        你可以使用这个代码: txt

        【讨论】:

        • 正如目前所写,您的答案尚不清楚。请edit 添加其他详细信息,以帮助其他人了解这如何解决所提出的问题。你可以找到更多关于如何写好答案的信息in the help center。
        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2014-06-19
        • 2020-05-07
        • 1970-01-01
        • 2022-11-03
        相关资源
        最近更新 更多