【发布时间】:2020-11-08 02:09:26
【问题描述】:
我正在使用以下函数尝试从网络抓取的 url 列表中提取 titles。
我确实看过一些 SO 答案,但请注意许多人建议避免使用正则表达式解决方案。我想修复和构建我现有的解决方案,但很高兴提出其他优雅的解决方案。
试图从 url 中提取标题的代码(函数)。
def titleextract(url):
#return unquote(url[58:url.rindex("/",58)-8].replace('_',''))
cleanedtitle1=url[58:]
title= cleanedtitle1.strip("-_Google_Art_Project.jpg/220px-")
return title
以上对 URL 的影响如下:
Url 1:Rembrandt_-Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son-Google_Art_Project.jpg/220px-Rembrandt-Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son-_Google_Art_Project.jpg
网址 2:Rembrandt_van_Rijn_-Saskia_van_Uylenburgh%2C_the_Wife_of_the_Artist-Google_Art_Project.jpg/220px-Rembrandt_van_Rijn-Saskia_van_Uylenburgh%2C_the_Wife_of_the_Artist-_Google_Art_Project.jpg p>
但所需的输出是:
网址 1: Rembrandt_-_Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son
网址 2: Rembrandt_van_Rijn_-_Saskia_van_Uylenburgh2C_the_Wife_of_the_Artist
我正在努力解决的问题是在此之后删除所有内容:_-Google_Art_Project.jpg/220px-Rembrandt-Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son-_Google_Art_Project.jpg 用于每个独特的案例以及然后删除不需要的字符(如果它们存在),例如 url2 中的 %。
理想情况下,我还想去掉标题中的下划线。
任何使用我现有代码的建议以及适当的分步说明都将不胜感激。
我删除开头的尝试成功了:
cleanedtitle1=url[58:]
但我尝试了各种方法来剥离字符并删除结尾,但都没有奏效:
title= cleanedtitle1.strip("-_Google_Art_Project.jpg/220px-")
基于一个建议,我也尝试了:
return unquote(url[58:url.rindex("/",58)-8].replace('_',''))
..但这并不能正确删除不需要的文本,仅删除最后 8 个字符,但由于它是可变的,因此不起作用。
我也试过这个,再次删除下划线 - 没有运气。
cleanedtitle1=url[58:]
cleanedtitle2= cleanedtitle1.strip("-_Google_Art_Project.jpg/220px-")
title = cleanedtitle2.strip("_")
return title
到目前为止我的导入是:
from flask import Flask, render_template,url_for #importing flask class
from urllib.request import urlopen
from bs4 import BeautifulSoup
import re
from urllib.parse import unquote
出于学习目的,我很乐意接受与相关的答案,但理想情况下,我也希望能完成我已经开始的工作。
对于仅使用 BeautifulSoup 的答案,这里是完整的完整代码 (这也是非常有用的参考)
from flask import Flask, render_template,url_for #importing flask class
from urllib.request import urlopen
from bs4 import BeautifulSoup
import re
from urllib.parse import unquote
app = Flask(__name__) #setting app variable to instance of flask class
@app.route('/') #this is what we type into our browser to go to pages. we create these using routes
@app.route('/home')
def home():
images=imagescrape()
titles=(titleextract(src) for src in images)
images_titles=zip(images,titles)
return render_template('home.html',images=images,images_titles=images_titles)
def titleextract(url):
pos1 = url.rindex("/")
pos2 = url[:pos1].rindex("/")
cleanedtitle1 = url[pos2 + 1: pos1]
title = cleanedtitle1.replace("_-_Google_Art_Project.jpg", "")
title = title.replace("_", " ")
return title
def imagescrape():
result_images=[]
html = urlopen('https://en.wikipedia.org/wiki/Rembrandt')
bs = BeautifulSoup(html, 'html.parser')
images = bs.find_all('img', {'src':re.compile('.jpg')})
for image in images:
result_images.append("https:"+image['src']+'\n') #concatenation!
return result_images
【问题讨论】:
标签: python string replace slice strip