【发布时间】:2018-05-10 08:00:01
【问题描述】:
我正在尝试从网页中抓取评论以确定词频。但是,当评论较长时,只会给出部分评论。您必须点击“更多”才能让网页显示完整评论。这是我用来提取评论文本的代码。我如何“点击”更多以获得完整评论?
library(rvest)
tripAdvisorURL <- "https://www.tripadvisor.com/Hotel_Review-g33657-d85704-
Reviews-Hotel_Bristol-Steamboat_Springs_Colorado.html#REVIEWS"
webpage <-read_html(tripAdvisorURL)
reviewData <- xml_nodes(webpage,xpath = '//*[contains(concat( " ", @class, "
" ), concat( " ", "partial_entry", " " ))]')
head(reviewData)
xml_text(reviewData[[1]])
[1] "The rooms were clean and we slept so good we had room 10 and 12 we
didn’t use 12 but it joins 10 .kind of strange but loved the hotel ..me
personally I would take the hot tub out it was kinda old..the lady
that...More"
【问题讨论】:
-
你研究过 Rselenium 吗?
-
您也可以点击标题中的链接访问全文。使用 ShowUserReviews 在页面中查找链接
-
您是否查看过旅行顾问的服务条款/条款和条件以及
robots.txt,或者您只是喜欢伤害他人? -
我没有看到他们的服务条款,@hrbrmstr。我只是将它用于在 R 中练习编程。我将转向另一个网站进行抓取。感谢您指出这一点。
标签: r web-scraping rvest