【发布时间】:2017-08-11 03:04:16
【问题描述】:
现在,我一直在尝试使用以下代码抓取 Google 图片:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
import os
import time
import requests
import re
import urllib2
import re
from threading import Thread
import json
#Assuming I have a folder named Pictures1, the images are downloaded there.
def threaded_func(url,i):
raw_img = urllib2.urlopen(url).read()
cntr = len([i for i in os.listdir("Pictures1") if image_type in i]) + 1
f = open("Pictures1/" + image_type + "_"+ str(total), 'wb')
f.write(raw_img)
f.close()
driver = webdriver.Firefox()
driver.get("https://images.google.com/")
elem = driver.find_element_by_xpath('/html/body/div/div[3]/div[3]/form/div[2]/div[2]/div[1]/div[1]/div[3]/div/div/div[2]/div/input[1]')
elem.clear()
elem.send_keys("parrot")
elem.send_keys(Keys.RETURN)
image_type = "parrot_defG"
images=[]
total=0
time.sleep(10)
for a in driver.find_elements_by_class_name('rg_meta'):
link =json.loads(a.text)["ou"]
thread = Thread(target = threaded_func, args = (link,total))
thread.start()
thread.join()
total+=1
我尝试使用 Selenium 打开 google 的图像结果页面,然后注意到每个 div 都有类 'rg-meta' 并且后面是 JSON 代码。
我尝试使用 .text 访问它。 JSON 的“ou”索引包含我要下载的图像的来源。我正在尝试使用“rg-meta”类获取所有此类 div 并下载图像。但它显示错误 " NO JSON OBJECT CAN BE DECODED" 我不知道该怎么做。
编辑: 这就是我要说的:
<div class="rg_meta">{"cl":3,"id":"FqCGaup9noXlMM:","isu":"kids.britannica.com","itg":false,"ity":"jpg","oh":600,"ou":"http://media.web.britannica.com/eb-media/89/89689-004-4C85E0F0.jpg","ow":380,"pt":"grain weevil -- Kids Encyclopedia | Children\u0026#39;s Homework Help ...","rid":"EusB0pk_sLg7vM","ru":"http://kids.britannica.com/comptons/art-143712/grain-or-granary-weevil","s":"grain weevil","sc":1,"st":"Kids Britannica","th":282,"tu":"https://encrypted-tbn2.gstatic.com/images?q\u003dtbn:ANd9GcQPbgXbRVzOicvPfBRtAkLOpJwy_wDQEC6a2q0BuTsUx-s0-h4b","tw":179}</div>
检查 JSON 的“ou”索引。请帮我提取它。
请原谅我的无知。
我通过以下更新解决了这个问题:
for a in driver.find_elements_by_xpath('//div[@class="rg_meta"]'):
atext = a.get_attribute('innerHTML')
link =json.loads(atext)["ou"]
print link
thread = Thread(target = threaded_func, args = (link,total))
thread.start()
thread.join()
total+=1
【问题讨论】:
-
您是否尝试过打印 a.text 的上下文。如果是,请提供该输出。您还尝试过什么来调试您遇到的问题
-
是的@GiannisSpiliopoulos,试图打印 .text 的内容。在我的终端中显示为空白。我假设文本可能是 unicode,所以我尝试使用 json.dumps() 将其转换为 JSON。它也没有工作。
-
a打印什么? -
我不清楚将 unicode 字符串转换为 json 将如何帮助您...无论如何尝试打印整个页面的源 (
print(driver.page_source)) 并检查您的假设是否正确。 -
@SatishGarg ,如下:
标签: python json selenium web-scraping urllib2