【问题标题】:Getting texts from urls is returning empty dataframe从 url 获取文本返回空数据框
【发布时间】:2021-08-12 18:48:42
【问题描述】:

我正在尝试使用 for 循环从几个网站获取所有段落,但我得到一个空数据框。 程序的逻辑是

urls=[]
texts = []        

for r in my_list:
                try:
                    # Get text
                    url = urllib.urlopen(r)
                    content = url.read()
                    soup = BeautifulSoup(content, 'lxml')
                    # Find all of the text between paragraph tags and strip out the html
                    page = soup.find('p').getText()
                    texts.append(page)
                    urls.append(r)
                    
                except Exception as e:
                    print(e)
                    continue

df = pd.DataFrame({"Urls" : urls, "Texts:" : texts})
    
      

网址 (my_list) 的示例可能是:https://www.ford.com.au/performance/mustang/、https://soperth.com.au/perths-best-fish-and-chips-46154、https://www.tripadvisor.com.au/Restaurants-g255103-zfd10901-Perth_Greater_Perth_Western_Australia-Fish_and_Chips.html、https://www.bbc.co.uk/programmes/b07d2wy4

我怎样才能正确地存储该特定页面上的链接和文本(所以不是整个网站!)?

预期输出:

Urls                                                       Texts

https://www.ford.com.au/performance/mustang/         Nothing else offers the unique combination of classic style and exhilarating performance quite like the Ford Mustang. Whether it’s the Fastback or Convertible, 5.0L V8 or High Performance 2.3L, the Mustang has a heritage few other cars can match.
https://soperth.com.au/perths-best-fish-and-chips-46154 
https://www.tripadvisor.com.au/Restaurants-g255103-zfd10901-Perth_Greater_Perth_Western_Australia-Fish_and_Chips.html 
https://www.bbc.co.uk/programmes/b07d2wy4 

我应该在文本中为每个 url 包含该页面中包含的段落(即,所有

元素)。 即使是一个虚拟代码(所以不完全是我的)也有助于理解我的错误在哪里。我想我目前的错误可能在这一步:url = urllib.urlopen(r) 因为我没有文字。

【问题讨论】:

  • 任何输出?请参阅How to Ask 以及如何创建minimal reproducible example。
  • 本题没有问题。
  • 嗨,彼得。没有输出:只有 Urls Texts 作为标题,没有行。我提供了应该创建数据框的代码部分,其中包含有关链接和段落的抓取信息。我添加了这个问题。它是关于正确存储链接和文本。我将提供一个预期输出的示例
  • 如果你想找到所有段落,你不应该使用find_all方法
  • 对,@BhavyaParikh。但即使我将find 替换为find_all,我也会得到空数据框。

标签: python web-scraping beautifulsoup urllib


【解决方案1】:

我尝试了以下代码(python3:因此是 urllib.request),它可以工作。在 urlopen 挂起时添加了用户代理。

import pandas as pd
import urllib
from bs4 import BeautifulSoup

urls = []
texts = []
my_list = ["https://www.ford.com.au/performance/mustang/", "https://soperth.com.au/perths-best-fish-and-chips-46154",
           "https://www.tripadvisor.com.au/Restaurants-g255103-zfd10901-Perth_Greater_Perth_Western_Australia-Fish_and_Chips.html", "https://www.bbc.co.uk/programmes/b07d2wy4"]

for r in my_list:
    try:
        # Get text
        req = urllib.request.Request(
            r,
            data=None,
            headers={
                'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_9_3) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/35.0.1916.47 Safari/537.36'
            }
        )
        url = urllib.request.urlopen(req)
        content = url.read()
        soup = BeautifulSoup(content, 'lxml')

        # Find all of the text between paragraph tags and strip out the html
        page = ''
        for para in soup.find_all('p'):
            page += para.get_text()
        print(page)
        texts.append(page)
        urls.append(r)
    except Exception as e:
        print(e)
        continue

df = pd.DataFrame({"Urls": urls, "Texts:": texts})
print(df)

【讨论】:

  • 谢谢@SubhashR。那么你认为问题出在 urllib.request 和 userAgent 上吗?
  • 我相信是用户代理
猜你喜欢
  • 2023-04-05
  • 1970-01-01
  • 2020-05-14
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-09-02
相关资源
最近更新 更多