【发布时间】:2009-05-19 04:36:31
【问题描述】:
有很多关于如何在 Python 中解析 URL 的问题,这个问题是关于最好或最 Pythonic 的方法。
在我的解析中,我需要 4 个部分:网络位置、URL 的第一部分、路径和文件名以及查询字符串部分。
http://www.somesite.com/base/first/second/third/fourth/foo.html?abc=123
应该解析成:
netloc = 'www.somesite.com'
baseURL = 'base'
path = '/first/second/third/fourth/'
file = 'foo.html?abc=123'
下面的代码产生了正确的结果,但是在 Python 中有没有更好的方法来做到这一点?
url = "http://www.somesite.com/base/first/second/third/fourth/foo.html?abc=123"
file= url.rpartition('/')[2]
netloc = urlparse(url)[1]
pathParts = path.split('/')
baseURL = pathParts[1]
partCount = len(pathParts) - 1
path = "/"
for i in range(2, partCount):
path += pathParts[i] + "/"
print 'baseURL= ' + baseURL
print 'path= ' + path
print 'file= ' + file
print 'netloc= ' + netloc
【问题讨论】:
-
与 258746 不太一样,这个问题的目标略有不同,主要关注的是完成任务的最佳(Pythonic)方式。