【问题标题】:Scraping Google Images using Selenium in Python在 Python 中使用 Selenium 抓取 Google 图片
【发布时间】:2017-08-11 03:04:16
【问题描述】:

现在,我一直在尝试使用以下代码抓取 Google 图片:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys 
import os
import time
import requests
import re
import urllib2
import re
from threading import Thread
import json
#Assuming I have a folder named Pictures1, the images are downloaded there. 
def threaded_func(url,i):
     raw_img = urllib2.urlopen(url).read()
     cntr = len([i for i in os.listdir("Pictures1") if image_type in i]) + 1
     f = open("Pictures1/" + image_type + "_"+ str(total), 'wb')
     f.write(raw_img)
     f.close()
driver = webdriver.Firefox()
driver.get("https://images.google.com/")
elem = driver.find_element_by_xpath('/html/body/div/div[3]/div[3]/form/div[2]/div[2]/div[1]/div[1]/div[3]/div/div/div[2]/div/input[1]')
elem.clear()
elem.send_keys("parrot")
elem.send_keys(Keys.RETURN)
image_type = "parrot_defG"
images=[]
total=0
time.sleep(10)
for a in driver.find_elements_by_class_name('rg_meta'):
     link =json.loads(a.text)["ou"]
     thread = Thread(target = threaded_func, args = (link,total))
     thread.start()
     thread.join()
     total+=1

我尝试使用 Selenium 打开 google 的图像结果页面,然后注意到每个 div 都有类 'rg-meta' 并且后面是 JSON 代码。

我尝试使用 .text 访问它。 JSON 的“ou”索引包含我要下载的图像的来源。我正在尝试使用“rg-meta”类获取所有此类 div 并下载图像。但它显示错误 " NO JSON OBJECT CAN BE DECODED" 我不知道该怎么做。

编辑: 这就是我要说的:

    <div class="rg_meta">{"cl":3,"id":"FqCGaup9noXlMM:","isu":"kids.britannica.com","itg":false,"ity":"jpg","oh":600,"ou":"http://media.web.britannica.com/eb-media/89/89689-004-4C85E0F0.jpg","ow":380,"pt":"grain weevil -- Kids Encyclopedia | Children\u0026#39;s Homework Help ...","rid":"EusB0pk_sLg7vM","ru":"http://kids.britannica.com/comptons/art-143712/grain-or-granary-weevil","s":"grain weevil","sc":1,"st":"Kids Britannica","th":282,"tu":"https://encrypted-tbn2.gstatic.com/images?q\u003dtbn:ANd9GcQPbgXbRVzOicvPfBRtAkLOpJwy_wDQEC6a2q0BuTsUx-s0-h4b","tw":179}</div>

检查 JSON 的“ou”索引。请帮我提取它。

请原谅我的无知。

我通过以下更新解决了这个问题:

    for a in driver.find_elements_by_xpath('//div[@class="rg_meta"]'):
        atext = a.get_attribute('innerHTML')
        link =json.loads(atext)["ou"]
        print link
        thread = Thread(target = threaded_func, args = (link,total))
        thread.start()
        thread.join()
        total+=1

【问题讨论】:

  • 您是否尝试过打印 a.text 的上下文。如果是,请提供该输出。您还尝试过什么来调试您遇到的问题
  • 是的@GiannisSpiliopoulos,试图打印 .text 的内容。在我的终端中显示为空白。我假设文本可能是 unicode,所以我尝试使用 json.dumps() 将其转换为 JSON。它也没有工作。
  • a 打印什么?
  • 我不清楚将 unicode 字符串转换为 json 将如何帮助您...无论如何尝试打印整个页面的源 (print(driver.page_source)) 并检查您的假设是否正确。
  • @SatishGarg ,如下:

标签: python json selenium web-scraping urllib2


【解决方案1】:

替换:

driver.find_elements_by_class_name('rg_meta') 与driver.find_element_by_xpath('//div[@class="rg_meta"]/text()')

和a.text 和a

将解决您的问题。

结果代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys 
import os
import time
import requests
import re
import urllib2
import re
from threading import Thread
import json
#Assuming I have a folder named Pictures1, the images are downloaded there. 
def threaded_func(url,i):
     raw_img = urllib2.urlopen(url).read()
     cntr = len([i for i in os.listdir("Pictures1") if image_type in i]) + 1
     f = open("Pictures1/" + image_type + "_"+ str(total), 'wb')
     f.write(raw_img)
     f.close()
driver = webdriver.Firefox()
driver.get("https://images.google.com/")
elem = driver.find_element_by_xpath('/html/body/div/div[3]/div[3]/form/div[2]/div[2]/div[1]/div[1]/div[3]/div/div/div[2]/div/input[1]')
elem.clear()
elem.send_keys("parrot")
elem.send_keys(Keys.RETURN)
image_type = "parrot_defG"
images=[]
total=0
time.sleep(10)
for a in driver.find_element_by_xpath('//div[@class="rg_meta"]/text()'):
     link =json.loads(a)["ou"]
     thread = Thread(target = threaded_func, args = (link,total))
     thread.start()
     thread.join()
     total+=1

打印链接会导致:

http://media.web.britannica.com/eb-media/89/89689-004-4C85E0F0.jpg

【讨论】:

  • 不适用于我,请提供您进行更改的 sn-p。这将非常有帮助。谢谢。
  • 即使在将其更正后,对于 driver.find_elements_by_xpath(...) 中的 a,也会出现错误:TypeError: expected string or buffer
  • 在哪一行?
  • 抱歉,您的代码不工作 :( :( 行是:link =json.loads(a)["ou"] 错误是:TypeError: expected string or buffer
  • 您可以添加回溯到您收到的错误吗?
猜你喜欢
  • 2020-05-22
  • 1970-01-01
  • 1970-01-01
  • 2017-04-27
  • 2022-10-05
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多