【问题标题】:Python extract title from URLPython从URL中提取标题
【发布时间】:2020-11-08 02:09:26
【问题描述】:

我正在使用以下函数尝试从网络抓取的 url 列表中提取 titles。

我确实看过一些 SO 答案,但请注意许多人建议避免使用正则表达式解决方案。我想修复和构建我现有的解决方案,但很高兴提出其他优雅的解决方案。

示例网址 1:https://upload.wikimedia.org/wikipedia/commons/thumb/b/bd/Rembrandt_van_Rijn_-_Self-Portrait_-_Google_Art_Project.jpg/220px-Rembrandt_van_Rijn_-_Self-Portrait_-_Google_Art_Project.jpg

示例网址 2: https://upload.wikimedia.org/wikipedia/commons/thumb/a/ae/Rembrandt_-_Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son_-_Google_Art_Project.jpg/220px-Rembrandt_-_Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son_-_Google_Art_Project.jpg

试图从 url 中提取标题的代码(函数)。

def titleextract(url):
    #return unquote(url[58:url.rindex("/",58)-8].replace('_',''))
    cleanedtitle1=url[58:]
    title= cleanedtitle1.strip("-_Google_Art_Project.jpg/220px-")
    return title

以上对 URL 的影响如下:

Url 1:Rembrandt_-Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son-Google_Art_Project.jpg/220px-Rembrandt-Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son-_Google_Art_Project.jpg

网址 2:Rembrandt_van_Rijn_-Saskia_van_Uylenburgh%2C_the_Wife_of_the_Artist-Google_Art_Project.jpg/220px-Rembrandt_van_Rijn-Saskia_van_Uylenburgh%2C_the_Wife_of_the_Artist-_Google_Art_Project.jpg p>

但所需的输出是:

网址 1: Rembrandt_-_Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son

网址 2: Rembrandt_van_Rijn_-_Saskia_van_Uylenburgh2C_the_Wife_of_the_Artist

我正在努力解决的问题是在此之后删除所有内容:_-Google_Art_Project.jpg/220px-Rembrandt-Rembrandt_and_Saskia_in_the_Scene_of_the_Prodigal_Son-_Google_Art_Project.jpg 用于每个独特的案例以及然后删除不需要的字符(如果它们存在),例如 url2 中的 %。

理想情况下,我还想去掉标题中的下划线。

任何使用我现有代码的建议以及适当的分步说明都将不胜感激。

我删除开头的尝试成功了:

cleanedtitle1=url[58:]

但我尝试了各种方法来剥离字符并删除结尾,但都没有奏效:

title= cleanedtitle1.strip("-_Google_Art_Project.jpg/220px-")

基于一个建议,我也尝试了:

return unquote(url[58:url.rindex("/",58)-8].replace('_',''))

..但这并不能正确删除不需要的文本,仅删除最后 8 个字符,但由于它是可变的,因此不起作用。

我也试过这个,再次删除下划线 - 没有运气。

cleanedtitle1=url[58:]
    cleanedtitle2= cleanedtitle1.strip("-_Google_Art_Project.jpg/220px-")
    title = cleanedtitle2.strip("_")
    return title

到目前为止我的导入是:

from flask import Flask, render_template,url_for #importing flask class
from urllib.request import urlopen
from bs4 import BeautifulSoup
import re
from urllib.parse import unquote

出于学习目的,我很乐意接受与相关的答案,但理想情况下,我也希望能完成我已经开始的工作。

对于仅使用 BeautifulSoup 的答案,这里是完整的完整代码 (这也是非常有用的参考)

from flask import Flask, render_template,url_for #importing flask class
from urllib.request import urlopen
from bs4 import BeautifulSoup
import re
from urllib.parse import unquote

app = Flask(__name__) #setting app variable to instance of flask class

@app.route('/') #this is what we type into our browser to go to pages. we create these using routes
@app.route('/home')
def home():
    images=imagescrape()
    titles=(titleextract(src) for src in images)
    images_titles=zip(images,titles)
    return render_template('home.html',images=images,images_titles=images_titles)   

def titleextract(url):
    pos1 = url.rindex("/")
    pos2 = url[:pos1].rindex("/")
    cleanedtitle1 = url[pos2 + 1: pos1]
    title = cleanedtitle1.replace("_-_Google_Art_Project.jpg", "")
    title = title.replace("_", " ")
    return title


def imagescrape():
    result_images=[]
    html = urlopen('https://en.wikipedia.org/wiki/Rembrandt')
    bs = BeautifulSoup(html, 'html.parser')
    images = bs.find_all('img', {'src':re.compile('.jpg')})
    for image in images:
        result_images.append("https:"+image['src']+'\n') #concatenation!
    return result_images

【问题讨论】:

    标签: python string replace slice strip


    【解决方案1】:

    从你的:

    cleanedtitle1=url[58:]
    

    这可行,但它对硬编码数字可能不是很健壮,所以让我们从倒数第二个“/”之后的字符开始。

    您可以使用正则表达式来做到这一点,但更简单地说,这可能看起来像:

    pos1 = url.rindex("/")  # index of last /
    pos2 = url[:pos1].rindex("/")  # index of second-to-last /
    cleanedtitle1 = url[pos2 + 1:]
    

    虽然实际上,您只对倒数第二个和最后一个 / 之间的位感兴趣,所以让我们改用我们发现的 pos1 作为中间位:

    pos1 = url.rindex("/")  # index of last /
    pos2 = url[:pos1].rindex("/")  # index of second-to-last /
    cleanedtitle1 = url[pos2 + 1: pos1]
    

    这里,cleanedtitle1 的值如下

    'Rembrandt_van_Rijn_-_Self-Portrait_-_Google_Art_Project.jpg'
    

    现在转到您的strip。这不会完全符合您的要求:它将遍历您给它的字符串,给出该字符串中的各个字符,然后删除 all 出现的 each那些字符。

    因此,让我们使用replace,并将字符串替换为空字符串。

    title = cleanedtitle1.replace("_-_Google_Art_Project.jpg", "")
    

    我们也可以这样做:

    title = title.replace("_", " ")
    

    然后我们最终得到:

    'Rembrandt van Rijn - Self-Portrait'
    

    把它放在一起:

    pos1 = url.rindex("/")
    pos2 = url[:pos1].rindex("/")
    cleanedtitle1 = url[pos2 + 1: pos1]
    title = cleanedtitle1.replace("_-_Google_Art_Project.jpg", "")
    title = title.replace("_", " ")
    return title
    

    更新

    我错过了这样一个事实,即 URL 可能包含我们希望替换的序列,例如 %2C。

    这些可以使用replace 以相同的方式完成,例如:

    url = url.replace("%2C", ",")
    

    但是你必须为所有可能发生的类似序列执行此操作,因此最好使用urllib 提供的unquote 函数。如果您在代码的顶部放置:

    from urllib.parse import unquote
    

    那么您可以使用

    进行这些替换
    url = unquote(url)
    

    在其余处理之前:

    from urllib.parse import unquote
    
    def titleextract(url):
        url = unquote(url)
        pos1 = url.rindex("/")
        pos2 = url[:pos1].rindex("/")
        cleanedtitle1 = url[pos2 + 1: pos1]
        title = cleanedtitle1.replace("_-_Google_Art_Project.jpg", "")
        title = title.replace("_", " ")
        return title
    

    【讨论】:

    • 这很好解释谢谢。它几乎可以工作,但并不完全:例如,在一个标题中:Rembrandt van Rijn - Saskia van Uylenburgh%2C 艺术家的妻子 % 仍然存在,但在其他人中,.jpg 仍然在末尾 Rembrandt van Rijn - A波兰贵族.jpg ..我正在研究你的答案,谢谢
    • 好的,这可以通过unquote 函数修复 - 我现在将对其进行编辑。
    • 你能不能留下你精彩的一步一步的解释并添加到它......我确实尝试使用 unquote 但真的不明白它是什么以及它是如何工作的(请参阅我的问题尝试)。所以那里的任何解释也会很有用。
    • @MissComputing 我只是在最后添加了一个更新,其他所有内容都保持原样。
    • 谢谢 - 你能把它们加到最后以便快速浏览一下吗?我可以检查一下
    【解决方案2】:

    这应该可行,有任何问题请告诉我

    def titleextract(url):
        title = url[58:]
        if "Google_Art_Project" in title:
            x = title.index("-_Google_Art_Project.jpg")
            title = title[:x] # Cut after where this is.
    
        disallowed_chars = "%" # Edit which chars should go.
        # Python will look at each character in turn. If it is not in the disallowed chars string, 
        # then it will be left. "".join() joins together all chars still allowed. 
        title = "".join(c for c in title if c not in disallowed_chars)
    
        title = title.replace("_"," ") # Change underscores to spaces.
        return title
    

    【讨论】:

    • 只是去检查它是否有效 - 您能否添加非常详细的 cmets 来解释每一行以用于学习目的?例如:join (c for c in title)
    • 好的,我已经编辑了那个。还有什么要解释的吗?
    • 谢谢。再一次,这几乎有效 - 所以非常适合 title1 title2 但再往下,Rembrandt van Rijn - Rembrandts zoon Titus in monniksdracht 28Rijksmuseum Amsterdam29.jpg/220px-Rembrandt van Rijn - Rembrandts zoon Titus in monniksdracht 28Rijksmuseum Amsterdam29.jpg ....所以我还在努力,
    【解决方案3】:

    有几种方法可以做到这一点:

    1. 如果您只想使用内置的 python 字符串函数,那么您可以通过首先在 / 的基础上拆分所有内容,然后在所有 URL 中剥离公共部分来实现。
    def titleextract(url):
        cleanedtitle1 = url.split("/")[-1]
        return cleanedtitle1[6:-4].replace('_',' ')
    
    1. 由于您已经在使用 bs4 导入,您可以通过以下方式进行:
    soup = BeautifulSoup(htmlString, 'html.parser')
    title = soup.title.text
    

    【讨论】:

    • 谢谢 - 我只是在检查答案。请让 cmets 把每一步都详细解释一下好吗?
    • 我认为事情是不言自明的。让我知道你不明白的地方。
    • 对于 bs4 的答案 - 是 htmlString = 到 URL,如果是,你如何传递整个列表(它被称为“图像”)
    • 对于您的第一个建议,它几乎适用,但并非适用于所有情况,并且 googleartproject 部分仍然存在 - Rembrandt van Rijn - Self-Portrait - Google Art Project。
    • 我对 BS4 的答案很感兴趣.....请您将其合并到我现有的代码中,以便我知道如何在 .... 中传递图像 url 列表。
    猜你喜欢
    • 1970-01-01
    • 2020-10-13
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-04-12
    • 1970-01-01
    • 1970-01-01
    • 2013-07-12
    相关资源
    最近更新 更多