【问题标题】:When appending dictionaries to a list through for loop, I get only the last dictionary通过 for 循环将字典附加到列表时,我只得到最后一个字典
【发布时间】:2016-07-22 23:42:18
【问题描述】:

我试图通过浏览所有不同的页面来抓取一个职业搜索网站,当我尝试使用 for 循环将字典附加到列表中时,我一直遇到问题。当我在 Python 3.4 中执行下面的代码时,代码会将每个页面中的所有相关数据提取到字典中(我已使用 print() 检查过)并附加到“FullJobDetails”中,但在 for 循环结束时我仅从最后一页获取一个充满字典的列表。字典的数量与列表“ListofJobs”中的页数完全相同。 “ListofJobs”是我要报废的每个页面的 html 链接列表。

我刚开始学习代码,所以我知道下面的代码并不是最有效或最好的方式。任何建议,将不胜感激。提前致谢!

FullJobDetails = []
browser = webdriver.Chrome()
dictionary = {}

for jobs in ListofJobs:
  browser.get(jobs)
  dictionary["Web Page"] = jobs
  try:
    dictionary["Views"] = browser.find_element_by_class_name('job-viewed-item-count').text
  except NoSuchElementException:
    dictionary["Views"] = 0

  try:
    dictionary['Applicants'] = browser.find_element_by_class_name('job-applied-item-count').text
  except NoSuchElementException:
    dictionary["Applicants"] = 0

  try:
    dictionary["Last Application"] = browser.find_element_by_class_name('last-application-time-digit').text
  except NoSuchElementException:
    dictionary["Last Application"] = "N/A"

  try:
    dictionary["Job Title"] = browser.find_element_by_class_name('title').text
  except NoSuchElementException:
    dictionary["Job Title"] = "N/A"

  try:
    dictionary['Company'] = browser.find_element_by_xpath('/html/body/div[3]/article/section[2]/div/ul/li[4]/span/span').text
  except NoSuchElementException:
    dictionary['Company'] = "Not found"

  try:
    dictionary['Summary'] = browser.find_element_by_class_name('summary').text
  except NoSuchElementException:
    dictionary['Summary'] = "Not found"

  FullJobDetails.append(dictionary)

【问题讨论】:

  • 等一下。您使用真正的 HTML 解析器解析 job.content,然后立即 unparse 它并使用正则表达式搜索原始文本?
  • 您确定您显示的代码就是您正在运行的代码吗?如果dict = {} 行在循环之外而不是在您显示它的位置,您描述的问题正是我所期望的。 (与您的问题无关的一点:使用dict 作为变量名是一个非常糟糕的主意。它会隐藏内置dict 类的名称,这可能会在以后导致非常混乱的错误。)
  • 是的,显示的代码与正在运行的代码完全相同,“缩进”等等。如果它正在重置自己,我想列表中只有一个字典(最后一个),而不是多个都对应于最后一个字典的字典。感谢您关于重命名 dict 的建议,我将其更改为另一个变量。

标签: list python-3.x for-loop dictionary web-scraping


【解决方案1】:

问题在于您只创建了一个字典 - 字典是可变对象 - 相同的字典一遍又一遍地附加到您的列表中,并且在 for 循环的每次传递中,您都会更新其内容。因此,最后,您将拥有同一个词典的多个副本,所有副本都显示最后一页上的信息。

只需为for 循环的每次运行创建一个新的字典对象。该新字典将保存在列表中,变量名称dictionary 可以保存您的新对象而不会发生冲突。

for jobs in ListofJobs:
  dictionary = {} 
  browser.get(jobs)
  ...

【讨论】:

  • 成功了!非常感谢您抽出宝贵时间回答问题。
猜你喜欢
  • 2021-05-19
  • 1970-01-01
  • 2022-12-13
  • 1970-01-01
  • 2014-07-06
  • 1970-01-01
  • 1970-01-01
  • 2018-12-01
相关资源
最近更新 更多