【发布时间】:2016-05-17 10:27:25
【问题描述】:
我想从这个网站http://tweakers.net 获取智能手机的价格。这是一个荷兰网站。问题是价格不是从网站收集的。
文本文件“TweakersTelefoons.txt”包含 3 个条目:
三星-galaxy-s6-32gb-zwart
lg-nexus-5x-32gb-zwart
huawei-nexus-6p-32gb-zwart
我使用的是 python 2.7,这是我使用的代码:
import urllib
import re
symbolfile = open("TweakersTelefoons.txt")
symbolslist = symbolfile.read()
symbolslist = symbolslist.split("\n")
for symbol in symbolslist:
url = "http://tweakers.net/pricewatch/[^.]*/" +symbol+ ".html"
## http://tweakers.net/pricewatch/423541/samsung-galaxy-s6-32gb-zwart.html is the original html
htmlfile = urllib.urlopen(url)
htmltext = htmlfile.read()
regex = '<span itemprop="lowPrice">(.+?)</span>'
## <span itemprop="lowPrice">€ 471,95</span> is what the original code looks like
pattern = re.compile(regex)
price = re.findall(pattern, htmltext)
print "the price of", symbol, "is ", price
输出:
samsung-galaxy-s6-32gb-zwart的价格是[]
lg-nexus-5x-32gb-zwart的价格是[]
huawei-nexus-6p-32gb-zwart的价格是[]
价格未显示 我尝试使用 [^.] 去掉欧元符号,但没有奏效。
此外,在欧洲,我们可能使用“,”而不是“。”作为小数的分隔符。 请帮忙。
提前谢谢你。
【问题讨论】:
-
使用 html 解析器,您想从网站获取什么?
-
"tweakers.net/pricewatch/[^.]*/" 没有按照你的想法做。 URL 不能包含小丑作为“*”(或者它是一个非常特殊的服务器配置)。
-
tweakers.net/pricewatch/423541/… 是原始网址,使用 [^.]* 替换第 423541 号。我正在尝试获取这些智能手机的价格。
-
我建议使用
for symbol in symbolslist而不是您的while i < len(symbolslist)和symbolslist[i]代码。 -
是的,确实看起来更好。我会改的。
标签: python html http currency digit-separator