【问题标题】:Python: getting prices of smartphones from websitePython:从网站获取智能手机的价格
【发布时间】:2016-05-17 10:27:25
【问题描述】:

我想从这个网站http://tweakers.net 获取智能手机的价格。这是一个荷兰网站。问题是价格不是从网站收集的。

文本文件“TweakersTelefoons.txt”包含 3 个条目:

三星-galaxy-s6-32gb-zwart

lg-nexus-5x-32gb-zwart

huawei-nexus-6p-32gb-zwart

我使用的是 python 2.7,这是我使用的代码:

import urllib
import re

symbolfile = open("TweakersTelefoons.txt")
symbolslist = symbolfile.read()
symbolslist = symbolslist.split("\n")

for symbol in symbolslist:
    url = "http://tweakers.net/pricewatch/[^.]*/" +symbol+ ".html"
## http://tweakers.net/pricewatch/423541/samsung-galaxy-s6-32gb-zwart.html  is the original html

    htmlfile = urllib.urlopen(url)
    htmltext = htmlfile.read()

    regex = '<span itemprop="lowPrice">(.+?)</span>'
## <span itemprop="lowPrice">€ 471,95</span>  is what the original code looks like
    pattern = re.compile(regex)
    price = re.findall(pattern, htmltext)

    print "the price of", symbol, "is ", price

输出:

samsung-galaxy-s6-32gb-zwart的价格是[]

lg-nexus-5x-32gb-zwart的价格是[]

huawei-nexus-6p-32gb-zwart的价格是[]

价格未显示 我尝试使用 [^.] 去掉欧元符号,但没有奏效。

此外,在欧洲,我们可能使用“,”而不是“。”作为小数的分隔符。 请帮忙。

提前谢谢你。

【问题讨论】:

  • 使用 html 解析器,您想从网站获取什么?
  • "tweakers.net/pricewatch/[^.]*/" 没有按照你的想法做。 URL 不能包含小丑作为“*”(或者它是一个非常特殊的服务器配置)。
  • tweakers.net/pricewatch/423541/… 是原始网址,使用 [^.]* 替换第 423541 号。我正在尝试获取这些智能手机的价格。
  • 我建议使用for symbol in symbolslist 而不是您的while i &lt; len(symbolslist) 和symbolslist[i] 代码。
  • 是的,确实看起来更好。我会改的。

标签: python html http currency digit-separator


【解决方案1】:
import requests

from bs4 import BeautifulSoup

soup = BeautifulSoup(requests.get("http://tweakers.net/categorie/215/smartphones/producten/").content)

print [(p.a["href"], p.a.text) for p in soup.find_all("p",{"class":"price"})]

获取所有页面:

from bs4 import BeautifulSoup

# base url to pass page number to 1-69 in this case
base_url = "http://tweakers.net/categorie/215/smartphones/producten/?page={}"
soup = BeautifulSoup(requests.get("http://tweakers.net/categorie/215/smartphones/producten/").content, "lxml")

# get and store all prices and phone links
data = {1: (p.a["href"], p.a.text) for p in soup.find_all("p", {'class': "price"})}

pag = soup.find("span", attrs={"class":"pageDistribution"}).find_all("a")

# last page number
mx_pg = max(int(a.text) for a in pag if a.text.isdigit())

# get all the pages from the second to  mx_pg 
for i in range(2, mx_pg + 1):
    req = requests.get(base_url.format(i))
    print req
    soup = BeautifulSoup(req.content)
    data[i] = [(p.a["href"], p.a.text) for p in soup.find_all("p",{"class":"price"})]

您将需要requests、BeautifulSoup。如果您想抓取更多数据,该字典包含指向您可以访问的每个电话页面的链接。

【讨论】:

  • 回溯(最近一次调用最后一次):文件“C:\PyScripter\PyScripter\MyScripts\WebScrapers\test.py”,第 12 行,在 导入请求中 ImportError: No module named requests
  • @JohanVerburg,你需要安装,你可以用 pip 来做,pip install requests 和 bs4 一样
  • 我安装了 requests 和 BeautifulSoup。我在上面运行了代码并得到了以下内容:Traceback(最近一次调用最后):文件“”,第 18 行,在 文件“C:\Python27\lib\site-packages\bs4_init_.py",第 156 行,在 init % ",".join(features)) FeatureNotFound:找不到具有您请求的功能的树生成器:lxml。需要安装解析器库吗?
  • 像我现在一样删除lxml,我建议安装lxml,如果你打算做大量的抓取
  • 当我删除“lxml”时,我得到以下结果: Traceback(最近一次调用最后一次):文件“C:\PyScripter\PyScripter\MyScripts\WebScrapers\test2.py”,第 23 行,在 pag = soup.find("span", attrs={"class":"pageDistribution"}).find_all("a") AttributeError: 'NoneType' 对象没有属性 'find_all'
【解决方案2】:

我认为您的问题是您希望网络服务器使用 "http://tweakers.net/pricewatch/[^.]*/ 解析 URL 中的通配符,而您没有检查我怀疑是 404 的返回代码。

您需要识别产品 ID(如果已修复)或使用表单发布方法发布搜索请求。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-07-06
    • 2022-01-08
    • 2013-06-27
    • 1970-01-01
    • 2015-07-13
    相关资源
    最近更新 更多