【发布时间】:2021-08-12 18:48:42
【问题描述】:
我正在尝试使用 for 循环从几个网站获取所有段落,但我得到一个空数据框。 程序的逻辑是
urls=[]
texts = []
for r in my_list:
try:
# Get text
url = urllib.urlopen(r)
content = url.read()
soup = BeautifulSoup(content, 'lxml')
# Find all of the text between paragraph tags and strip out the html
page = soup.find('p').getText()
texts.append(page)
urls.append(r)
except Exception as e:
print(e)
continue
df = pd.DataFrame({"Urls" : urls, "Texts:" : texts})
网址 (my_list) 的示例可能是:https://www.ford.com.au/performance/mustang/、https://soperth.com.au/perths-best-fish-and-chips-46154、https://www.tripadvisor.com.au/Restaurants-g255103-zfd10901-Perth_Greater_Perth_Western_Australia-Fish_and_Chips.html、https://www.bbc.co.uk/programmes/b07d2wy4
我怎样才能正确地存储该特定页面上的链接和文本(所以不是整个网站!)?
预期输出:
Urls Texts
https://www.ford.com.au/performance/mustang/ Nothing else offers the unique combination of classic style and exhilarating performance quite like the Ford Mustang. Whether it’s the Fastback or Convertible, 5.0L V8 or High Performance 2.3L, the Mustang has a heritage few other cars can match.
https://soperth.com.au/perths-best-fish-and-chips-46154
https://www.tripadvisor.com.au/Restaurants-g255103-zfd10901-Perth_Greater_Perth_Western_Australia-Fish_and_Chips.html
https://www.bbc.co.uk/programmes/b07d2wy4
我应该在文本中为每个 url 包含该页面中包含的段落(即,所有
元素)。
即使是一个虚拟代码(所以不完全是我的)也有助于理解我的错误在哪里。我想我目前的错误可能在这一步:url = urllib.urlopen(r) 因为我没有文字。
【问题讨论】:
-
任何输出?请参阅How to Ask 以及如何创建minimal reproducible example。
-
本题没有问题。
-
嗨,彼得。没有输出:只有 Urls Texts 作为标题,没有行。我提供了应该创建数据框的代码部分,其中包含有关链接和段落的抓取信息。我添加了这个问题。它是关于正确存储链接和文本。我将提供一个预期输出的示例
-
如果你想找到所有段落,你不应该使用
find_all方法 -
对,@BhavyaParikh。但即使我将
find替换为find_all,我也会得到空数据框。
标签: python web-scraping beautifulsoup urllib