【问题标题】:URL not updating when iterating through URLs遍历 URL 时 URL 不更新
【发布时间】:2020-11-21 01:44:35
【问题描述】:

以下代码执行以下操作:

1 打开一个特定的 URL(对于第一个日期 YYYY-MM-DD);

2 getURL() 生成所有日期在特定日期范围内的所有 URL(从第二天开始);

3 打开新标签页,其中第一个日期由 getURL() 生成;

4 返回上一个标签并关闭它;

5 重复步骤 3 和 4。

from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
from datetime import datetime, timedelta

# Load Chrome driver and movement.uber.com/cities website
PATH = 'C:\Program Files (x86)\chromedriver.exe'
driver = webdriver.Chrome(PATH)


# Attributing the city name and the center-most zone code (or origin) to variables so they can be inserted in the URL later
city = 'atlanta'
origin_code = '1074'
coordinates = '&lat.=33.7489&lng.=-84.4234622&z.=12'


# Open URL for the first day in the desired city (change coordinates depending on city)
driver.get('https://movement.uber.com/explore/' + city + '/travel-times/query?si' + origin_code + '&ti=&ag=taz&dt[tpb]=ALL_DAY&dt[wd;]=1,2,3,4,5,6,7&dt[dr][sd]=' + 
           '2016-01-02' + '&dt[dr][ed]=' + '2016-01-02' + '&cd=&sa;=&sdn=' + coordinates + '&lang=en-US')


# Generating the correct URLs for each date
def getURL():
    date = datetime(2016,1,4)
    while date <= datetime(2020,3,31):
        yield ('https://movement.uber.com/explore/' + city + '/travel-times/query?si' + origin_code + '&ti=&ag=taz&dt[tpb]=ALL_DAY&dt[wd;]=1,2,3,4,5,6,7&dt[dr][sd]=' +
               date.strftime('%Y-%m-%d') + '&dt[dr][ed]=' + date.strftime('%Y-%m-%d') + '&cd=&sa;=&sdn=&lat.=33.7489&lng.=-84.4234622&z.=12&lang=en-US')
        date += timedelta(days=1)

# Open new tab
i = 0
for url in getURL():
    i += 1
    if i < 3:
        driver.execute_script("window.open(url)")

        # Switch to previous tab and close it (leaving us with the newly above opened tab)
        tabs = driver.window_handles

        if len(tabs) > 1:
            driver.switch_to.window(tabs[0])
            driver.close()
            driver.switch_to.window(tabs[1])

问题:每次打开新标签/“窗口”时,代码都会打开第一个日期为 YYYY-MM-DD 的 URL,完全忽略了 getURL() 生成的 URL。

问题:如何打开下一个日期的新标签,关闭前一个,重复?

我的最终目标:下载每个不同 URL 内的数据集(但其代码与这里的问题无关)。 Obs.:我为此使用 Selenium 库。

【问题讨论】:

  • 代码执行后i的值是多少?
  • 如果我在最后一个 if 语句中添加一个 print(i),它首先打印 1,然后是 2。如果我在 for 循环中添加,它会给我 1549。它从 1 迭代到 1549 for不管什么原因。也许这就是生成的 URL 总数?
  • 这条线对吗? driver.execute_script("window.open(url)") - 这是传递一个字符串 url 而不是值。我看不出这会如何打开一个标签到任何位置
  • 但我想那是唯一真正打开新标签的代码行,不是吗?如果我运行代码,它实际上会打开一个新选项卡,只是它总是以 2016-01-02 打开,这是第一个日期。它不会移动到 2016-01-03。
  • 在execute_script 中修复了 JS 错误后,您的代码运行良好。如果发布的答案有效并且您现在在线,那么很棒,请忽略此评论。如果你还在苦苦挣扎,请告诉我。告诉我你是如何执行脚本的——我的 VSCode 捕获了 JS 错误,现在它无需其他更改即可运行......我将 i 提升到 10,我得到了 10 个带有迭代日期的 URL。你的代码很棒:-)

标签: python html selenium loops url


【解决方案1】:

错误在于您启动标签的方式。

当我改变 FROM 时:

driver.execute_script("window.open(url)")

到:

driver.execute_script("window.open('"+url+"','_blank')")

脚本在我的机器上完美执行。

看起来你的方法中的url 没有在迭代中更新。如果你把它变成for循环的一个参数,它每次都会被解析。

查看here 以获取有关如何打开窗口的 javascript 的更多信息(仅供参考 - 您也可以使用 _self 而不是 _blank 来替换当前窗口 - 这可能会减轻您对标签管理的需求) .

我的测试结果...

这是第一次迭代:

这是第二次迭代:


供参考这是我运行的整个脚本:(请注意,我更新了i 以进行更多迭代,我的机器的chromedriver PATH + 添加了几个prints)

from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
from datetime import datetime, timedelta

# Load Chrome driver and movement.uber.com/cities website
#PATH = ''# - mine lives local -> 'C:\Program Files (x86)\chromedriver.exe'
#driver = webdriver.Chrome(PATH)
driver = webdriver.Chrome()

# Attributing the city name and the center-most zone code (or origin) to variables so they can be inserted in the URL later
city = 'atlanta'
origin_code = '1074'
coordinates = '&lat.=33.7489&lng.=-84.4234622&z.=12'


# Open URL for the first day in the desired city (change coordinates depending on city)
driver.get('https://movement.uber.com/explore/' + city + '/travel-times/query?si' + origin_code + '&ti=&ag=taz&dt[tpb]=ALL_DAY&dt[wd;]=1,2,3,4,5,6,7&dt[dr][sd]=' + 
           '2016-01-02' + '&dt[dr][ed]=' + '2016-01-02' + '&cd=&sa;=&sdn=' + coordinates + '&lang=en-US')


# Generating the correct URLs for each date
def getURL():
    date = datetime(2016,1,4)
    while date <= datetime(2020,3,31):
        yield ('https://movement.uber.com/explore/' + city + '/travel-times/query?si' + origin_code + '&ti=&ag=taz&dt[tpb]=ALL_DAY&dt[wd;]=1,2,3,4,5,6,7&dt[dr][sd]=' +
               date.strftime('%Y-%m-%d') + '&dt[dr][ed]=' + date.strftime('%Y-%m-%d') + '&cd=&sa;=&sdn=&lat.=33.7489&lng.=-84.4234622&z.=12&lang=en-US')
        date += timedelta(days=1)

# Open new tab
i = 0
print("urls: %i", len(list(getURL())))
for url in getURL():
    i += 1
    if i < 10:
        driver.execute_script("window.open('"+url+"','_blank')")
        print (url)
        # Switch to previous tab and close it (leaving us with the newly above opened tab)
        tabs = driver.window_handles

        if len(tabs) > 1:
            driver.switch_to.window(tabs[0])
            driver.close()
            driver.switch_to.window(tabs[1])

【讨论】:

  • 这太好了,非常感谢!!不过还有一个问题:正如我在帖子底部提到的,我的目标是编写更多代码,也就是说,在该循环内再单击几下以获取我需要的数据。我点击按钮的代码是: time.sleep(3) download_button = driver.find_element_by_css_selector('div.f5 button') download_button.click() 我得到的错误是:“元素点击被拦截”。我在下面插入这段代码: driver.execute_script("window.open('"+url+"','_blank')") print(url) 为什么我不能再次单击下一个选项卡中的按钮?
  • 点击被拦截意味着你的按钮上面有东西。 - 当我转到您的基本 URL 时,会出现一个“了解来源”弹出窗口,阻止访问它后面的页面。你会处理吗? (按跳过)- & 处理后,您是否在等待按钮可用(即使用 webdriverwait 是正确的方法,但 sleep 是验证它是否与同步相关的快速方法)
  • 哦,真的!我忘记了跳过按钮。我将尝试添加我必须单击跳过按钮的代码。谢谢!页面加载速度如此之快,我没有注意到弹出窗口。让我试试。
  • 有趣。当我打开新选项卡时,“了解来源”窗口/弹出窗口永远不会出现。即使通过粘贴代码,我也必须在 download_button 之前关闭窗口,它也不起作用:/
  • 选项卡更难,因为它们具有共享状态。您在第一个选项卡上的操作会影响下一个选项卡。我没有检查过,但你可能只会在第一次访问时点击按钮......但如果它是另一个你需要支持的问题,然后发布一个新问题,人们可以加入:-)
【解决方案2】:

您也许可以尝试将您创建的所有网址放在一个列表中,然后返回它 像这样:

def getURL():
    tab = []
    date = datetime(2016,1,4)
    while date <= datetime(2020,3,31):
        url ='https://movement.uber.com/explore/' + city + '/travel-times/query?si' + origin_code + '&ti=&ag=taz&dt[tpb]=ALL_DAY&dt[wd;]=1,2,3,4,5,6,7&dt[dr][sd]=' +
               date.strftime('%Y-%m-%d') + '&dt[dr][ed]=' + date.strftime('%Y-%m-%d') + '&cd=&sa;=&sdn=&lat.=33.7489&lng.=-84.4234622&z.=12&lang=en-US'
        tab.append(url)
        date += timedelta(days=1)
    return tab

【讨论】:

  • 太棒了!但是有一个问题:此代码是否考虑到例如二月有时有 28 或 29 天这一事实?还是它完全忽略它并在第二天跳到 31 点?
  • 我觉得既然你使用了datetime类,它考虑到了2月份的28或29天,你可以尝试打印所有的日期看看结果
猜你喜欢
  • 2016-05-27
  • 1970-01-01
  • 1970-01-01
  • 2021-12-20
  • 2018-08-27
  • 1970-01-01
  • 2022-09-23
  • 2018-04-15
  • 1970-01-01
相关资源
最近更新 更多