【问题标题】:*Still stumped* Sentiment analysis of tweets using R with the twitteR, plyr, stringr, and RMySQL packages*仍然难倒*使用 R 和 twitteR、plyr、stringr 和 RMySQL 包对推文进行情绪分析
【发布时间】:2013-06-22 13:44:48
【问题描述】:

我正在学习 R twitteR 软件包的初学者教程,但遇到了障碍。 (教程网址:https://sites.google.com/site/miningtwitter/questions/sentiment/analysis)

我没有通过第 4 节中详述的 searchTwitter 函数导入推文列表,而是从 MySQL 数据库中导入推文数据框。我可以很好地从 MySQL 导入推文,但是当我尝试执行时:

wine_txt = sapply(wine_tweets, function(x) x$getText())

我收到一个错误:

x$getText 中的错误:$ 运算符对原子向量无效

数据已经是 data.frame 形式,我随后再次将其强制转换为 data.frame 以确保我仍然得到相同的错误。我在下面粘贴了我的完整代码,任何帮助将不胜感激。

library(twitteR)
library(plyr)
library(stringr)
library(RMySQL)

tweets.con<-dbConnect(MySQL(),user="XXXXXXXX",password="XXXXXXXX",dbname="XXXXXXX",host="XXXXXXX")
wine_tweets<-dbGetQuery(tweets.con,"select `tweet_text` from `tweets` where `created_at` BETWEEN timestamp(DATE_SUB(NOW(), INTERVAL 11 MINUTE)) AND timestamp(NOW())")

# function score.sentiment
score.sentiment = function(sentences, pos.words, neg.words, .progress='none')
{
# Parameters
# sentences: vector of text to score
# pos.words: vector of words of postive sentiment
# neg.words: vector of words of negative sentiment
# .progress: passed to laply() to control of progress bar

# create simple array of scores with laply
scores = laply(sentences,
function(sentence, pos.words, neg.words)
{
  # remove punctuation
  sentence = gsub("[[:punct:]]", "", sentence)
  # remove control characters
  sentence = gsub("[[:cntrl:]]", "", sentence)
  # remove digits?
  sentence = gsub('\\d+', '', sentence)

  # define error handling function when trying tolower
  tryTolower = function(x)
  {
     # create missing value
     y = NA
     # tryCatch error
     try_error = tryCatch(tolower(x), error=function(e) e)
     # if not an error
     if (!inherits(try_error, "error"))
     y = tolower(x)
     # result
     return(y)
  }
  # use tryTolower with sapply 
  sentence = sapply(sentence, tryTolower)

  # split sentence into words with str_split (stringr package)
  word.list = str_split(sentence, "\\s+")
  words = unlist(word.list)

  # compare words to the dictionaries of positive & negative terms
  pos.matches = match(words, pos.words)
  neg.matches = match(words, neg.words)

  # get the position of the matched term or NA
  # we just want a TRUE/FALSE
  pos.matches = !is.na(pos.matches)
  neg.matches = !is.na(neg.matches)

  # final score
  score = sum(pos.matches) - sum(neg.matches)
  return(score)
  }, pos.words, neg.words, .progress=.progress )

# data frame with scores for each sentence
scores.df = data.frame(text=sentences, score=scores)
return(scores.df)
}

# import positive and negative words
pos = readLines("/home/jgraab/R/scripts/positive_words.txt")
neg = readLines("/home/jgraab/R/scripts/negative_words.txt")
wine_txt = sapply(wine_tweets, function(x) x$getText())

【问题讨论】:

    标签: r twitter


    【解决方案1】:

    $ 用于抓取数据框(或列表等)的列,您不能使用它来应用功能。你想要类似的东西

    getText(x)
    

    那里。

    【讨论】:

    猜你喜欢
    • 2017-11-06
    • 2022-01-10
    • 1970-01-01
    • 1970-01-01
    • 2012-05-01
    • 2018-12-29
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多