【问题标题】:Scraping a webpage and related subsequent page using R or Python使用 R 或 Python 抓取网页和相关的后续页面
【发布时间】:2020-04-08 13:23:30
【问题描述】:

我想对歌词做一些 NLP 来按几十年对情绪进行分类。现在,给定一个特定艺术家的歌词页面,例如 The Smiths,我的首页会显示所有歌曲名称:

https://www.azlyrics.com/s/smiths.html

绕着喷泉转\n

你现在已经得到了一切\n

.....

每个标题都是指向实际歌词页面的链接

https://www.azlyrics.com/lyrics/smiths/reelaroundthefountain.html https://www.azlyrics.com/lyrics/smiths/youvegoteverythingnow.html

现在,如何从https://www.azlyrics.com/lyrics/smiths/XXX.html 中抓取所有歌词,其中XXX 是第一页https://www.azlyrics.com/s/smiths.html 上的标题。

感谢您的帮助!正如我所写,R 或 Python。真的没关系。最好,我希望将每个歌词保存在单独的 *.txt 文件中。

我试过了:

    from bs4 import BeautifulSoup
import requests
list =[title1, title2, .....]
for x in list:
    url= "https://www.azlyrics.com/lyrics/smiths?x".format(str)
    r=requests.get(url)
    soup= BeautifulSoup(r.text)

    for span in soup.findAll('span', attrs={'class': 'views-field views-field-created'}) :
        print r.get_text()

但是失败了。但是,如果随后的页面被编号,它就可以工作。

【问题讨论】:

  • 你不能在没有代码的情况下发布问题并要求人们为你编写它。该网站旨在帮助人们尝试解决问题。
  • 我建议你从 iframe 中提取所有带有Rcrawler 包的href,然后你应该使用Rvest 来获取文本或使用Rselenium。
  • @EricTruett 我试过了。编辑了问题。

标签: python r pandas web-scraping nlp


【解决方案1】:
import requests
from bs4 import BeautifulSoup

# GET request to scrape the page for lyric links
r = requests.get('https://www.azlyrics.com/s/smiths.html')
# create soup
soup = BeautifulSoup(r.text, 'lxml')
# base url
url = 'https://www.azlyrics.com/'
# list comprehension to get all the links to the song lyrics
album_list = [url+a['href'].strip('..') for a in soup.find(id='listAlbum').findAll('a', href=True)]

for song in album_list:
    # do stuff with song
    # resp = requests.get(song)
    # song_soup = BeautifulSoup(resp.text, 'lxml')
    # etc.

【讨论】:

  • 这真是太好了!谢谢你,@Chris!
【解决方案2】:

在 R 中,我们可以使用rvest。

首先,我们获取歌词的所有链接。

library(rvest)

url <- "https://www.azlyrics.com/s/smiths.html"
all_links <- url %>%
              read_html() %>%
              html_nodes('div.listalbum-item a') %>%
              html_attr('href') %>%
         {paste0('https://www.azlyrics.com/', sub('../', '', ., fixed = TRUE))}

然后从all_links中的每一页获取歌词。

all_lyrics <- purrr::map(all_links, ~.x %>%read_html() %>% html_nodes('div') %>% .[[20]] %>% html_text())

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2011-12-24
    相关资源
    最近更新 更多