【问题标题】:Firestore batch insert 375 documents per commit, but not 500 documents. Why?Firestore 每次提交批量插入 375 个文档,但不是 500 个文档。为什么?
【发布时间】:2021-06-29 20:27:22
【问题描述】:

我正在尝试使用以下代码通过云功能(超时 540 秒)在我的 firestore 数据库中插入更多的 1400 个对象:


...

const response = await fetch(url)
if (response.ok) {
    const json = await response.json()
    if (json.hasOwnProperty('data')) {
        const teams = json[`data`]
        
        var players = teams.flatMap((team) => {
            return team.squad.data
        })
        
        var playersBatch = []
        while (players.length > 0) {
            const playerBatch = players.splice(0, 375)
            playersBatch.push(playerBatch)
        }

        for (playerBatch of playersBatch) {
            const batch = database.batch()

            for (player of playerBatch) {
                const reference = database
                    .collection(`players`)
                    .doc(`${player.player_id}`)

                batch.set(reference, player, { merge: true })
            }

            await batch.commit()
        }
    } else {
        ...
    }
} else {
    ...
}

...

上面的代码对我有用,但是当我每批插入 375 个文档时工作,当我尝试插入 500 个文档时,批量提交在第一个循环中不起作用并给我一个超时异常。

函数执行耗时 540005 毫秒,完成状态为:'timeout'

批处理可以产生超时吗?插入大型文档时批处理有任何限制吗?为什么每次都能插入 375 而不能插入 500?

【问题讨论】:

    标签: javascript firebase google-cloud-platform google-cloud-firestore google-cloud-functions


    【解决方案1】:

    TLDR:不是批处理超时,而是需要超过 9 分钟才能完成,但是您的函数在 9 分钟后超时,这是超时的上限,这这就是您收到此错误的原因。

    这种情况下的问题是云函数本身,而不是批量写入。

    正如您在documentation 中看到的:

    批量写入最多可包含 500 个操作

    但是,云函数documentation for timeout 表示以下内容:

    函数执行时间受超时时间限制,您可以在函数部署时指定。默认情况下,函数会在 1 分钟后超时,但您可以将此时间延长至 9 分钟。

    如果您转换540005 ms,您将获得大约 9 分钟,这是超时前 Cloud Functions 执行的上限。所以这就是为什么您不能操作 500 条记录的批次,但您可以处理 375 条记录,因为它低于 Cloud Function 的 9 分钟超时限制。

    【讨论】:

    • 我认为这是因为 await batch.commit() 在 for 循环内(意味着它正在尝试编写 375 个文档,等待,编写下一个 375 个文档,等待,重复)你可能在哪里使用await Promise.all(batches.map((b) => b.commit())) 而不是在循环之外同时运行所有这些。
    【解决方案2】:

    如果我理解您的操作正确,这些是您尝试执行的步骤:

    1. 获取网址
    2. 对于响应中的每个团队,提取所有玩家的列表
    3. 对于每个玩家,更新他们在数据库中的数据

    稍微改组你的代码并将await batch.commit()移出for循环(这使你的代码在移动到下一个之前等待每个批处理完成),给出:

    const response = await fetch(url)
    
    if (!response.ok || response.status === 204) {
      // A 204 code will break response.json() with a parsing error
      // You might want to check for the 429 status here
      throw new Error(`Unexpected status code ${response.status}!`)
    }
    
    const json = await response.json() // note: empty bodies will throw a parsing error
    
    if (!json.hasOwnProperty("data")) {
      throw new Error(`Unexpected response body!`, json)
    }
    
    const teams = json["data"]
    
    const players = teams.flatMap((team) => {
        return team.squad.data // the array of players in this team
    })
    
    const playersInBatches = [];
    while (players.length > 0) {
        const thisPlayerBatch = players.splice(0, 500)
        playersInBatches.push(thisPlayerBatch)
    }
    
    const batches = playersInBatches.map((playersInThisBatch) => {
        const dbBatch = database.batch()
    
        for (let player of playersInThisBatch) {
            const reference = database
                .collection("players")
                .doc(`${player.player_id}`)
    
            dbBatch.set(reference, player, { merge: true })
        }
    
        return dbBatch
    }
    
    // commit all batches in parallel and wait for them to finish
    await Promise.all(batches.map((b) => b.commit())) 
    
    console.log("Synced successfully!")
    

    注意事项:

    • 作为对您所做工作的扩展,您可能希望存储响应的缓存标头,例如 ETag 或 Last-Modified。这使您可以在下载数据之前询问第三方服务器数据是否发生了变化。
    • 我已将您的if (condition) { /* do lots of work */ } else { /* do small amount of work to handle error */ } 替换为if (!condition) { /* do small amount of work to handle error */ return; } /* do lots of work */。这被称为“快速失败”,用于防止大型和/或嵌套的 if-else 树,同时还在导致错误的原因旁边显示您的错误处理。
    • 如果任何一个批次失败,那么其他批次可能无法成功写入数据库,但这个错误也存在于您的原始代码中。您可以将最后几行更改为以下内容,以使它们不会杀死其他批次:
    // commit all batches in parallel and wait for them to finish
    const results = await Promise.all(batches.map(
        (b) => b.commit().then(
            () => ({success: true}),
            (error) => ({success: false, error})
        )
    ))
    
    let succeeded = 0, failed = 0
    
    results.forEach(result => result.success ? succeeded++ : failed++)
    
    if (failed > 0) {
        console.log(`Synced ${succeeded}/${results.length} batches of players successfully!`)
        return
    }
    
    console.log("Synced all players successfully!")
    

    您也不是第一个遇到批量写入数据库的问题的人。有一个 MultiBatch 实用程序类为您处理批处理是很常见的。

    const response = await fetch(url)
    
    if (!response.ok || response.status === 204) {
      // A 204 code will break response.json() with a parsing error
      // You might want to check for the 429 status here
      throw new Error(`Unexpected status code ${response.status}!`)
    }
    
    const json = await response.json() // note: empty bodies will throw a parsing error
    
    if (!json.hasOwnProperty("data")) {
      throw new Error(`Unexpected response body!`, json)
    }
    
    const teams = json["data"]
    const multiBatch = new MultiBatch(database)
    const playersColRef = database.collection("players")
    
    teams.forEach((team) => {
         team.squad.data // the array of players in this team
             .forEach(player => {
                 const reference = playersColRef.doc(`${player.player_id}`)
                 multiBatch.set(reference, player, { merge: true })
             })
    })
    
    await multiBatch.commit(/* pass true here to suppress errors */)
    
    console.log("Synced successfully!")
    

    【讨论】:

      猜你喜欢
      • 2020-12-21
      • 1970-01-01
      • 2021-06-21
      • 2018-08-02
      • 1970-01-01
      • 2020-08-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多