【问题标题】:Pythonic URL ParsingPythonic URL 解析
【发布时间】:2009-05-19 04:36:31
【问题描述】:

有很多关于如何在 Python 中解析 URL 的问题,这个问题是关于最好或最 Pythonic 的方法。

在我的解析中,我需要 4 个部分:网络位置、URL 的第一部分、路径和文件名以及查询字符串部分。

http://www.somesite.com/base/first/second/third/fourth/foo.html?abc=123

应该解析成:

netloc = 'www.somesite.com'
baseURL = 'base'
path = '/first/second/third/fourth/'
file = 'foo.html?abc=123'

下面的代码产生了正确的结果,但是在 Python 中有没有更好的方法来做到这一点?

url = "http://www.somesite.com/base/first/second/third/fourth/foo.html?abc=123"

file=  url.rpartition('/')[2]
netloc = urlparse(url)[1]
pathParts = path.split('/')
baseURL = pathParts[1]

partCount = len(pathParts) - 1

path = "/"
for i in range(2, partCount):
    path += pathParts[i] + "/"


print 'baseURL= ' + baseURL
print 'path= ' + path
print 'file= ' + file
print 'netloc= ' + netloc

【问题讨论】:

标签: url python


【解决方案1】:

由于您对所需部分的要求与 urlparse 为您提供的不同,所以它会得到最好的。但是,您可以替换它:

partCount = len(pathParts) - 1

path = "/"
for i in range(2, partCount):
    path += pathParts[i] + "/"

有了这个:

path = '/'.join(pathParts[2:-1])

【讨论】:

    【解决方案2】:

    我倾向于从urlparse 开始。此外,您可以使用rsplit,以及split 和rsplit 的maxsplit 参数来简化一些事情:

    _, netloc, path, _, q, _ = urlparse(url)
    _, base, path = path.split('/', 2) # 1st component will always be empty
    path, file = path.rsplit('/', 1)
    if q: file += '?' + q
    

    【讨论】:

      猜你喜欢
      • 2020-10-15
      • 2011-12-12
      • 2014-01-23
      • 2014-04-02
      • 1970-01-01
      • 2018-08-30
      • 2023-04-06
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多