【问题标题】:Performance: how to insert CLOB fast using cx_Oracle and executemany()?性能:如何使用 cx_Oracle 和 executemany() 快速插入 CLOB?
【发布时间】:2014-05-03 12:21:37
【问题描述】:

在我尝试使用 CLOB 值之前,cx_Oracle API 对我来说非常快。

我是这样做的:

import time
import cx_Oracle

num_records = 100
con = cx_Oracle.connect('user/password@sid')
cur = con.cursor()
cur.prepare("insert into table_clob (msg_id, message) values (:msg_id, :msg)")
cur.bindarraysize = num_records
msg_arr = cur.var(cx_Oracle.CLOB, arraysize=num_records)
text = '$'*2**20    # 1 MB of text
rows = []

start_time = time.perf_counter()
for id in range(num_records):
    msg_arr.setvalue(id, text)
    rows.append( (id, msg_arr) )    # ???

print('{} records prepared, {:.3f} s'
    .format(num_records, time.perf_counter() - start_time))
start_time = time.perf_counter()
cur.executemany(None, rows)
con.commit()
print('{} records inserted, {:.3f} s'
    .format(num_records, time.perf_counter() - start_time))

cur.close()
con.close()
  1. 让我担心的主要问题是性能:

    100 records prepared, 25.090 s - Very much for copying 100MB in memory!
    100 records inserted, 23.503 s - Seems to be too much for 100MB over network.
    

    有问题的步骤是msg_arr.setvalue(id, text)。如果我评论它,脚本只需几毫秒即可完成(当然,将 null 插入 CLOB 列)。

  2. 其次,在rows 数组中添加对 CLOB 变量的相同引用似乎很奇怪。我在互联网上找到了这个例子,它工作正常,但我做对了吗?

  3. 在我的情况下,有什么方法可以提高性能?

更新:测试的网络吞吐量:一个 107 MB 的文件在 11 秒内通过 SMB 复制到同一台主机。但同样,网络传输不是主要问题。数据准备需要异常多的时间。

【问题讨论】:

  • 第一件事是衡量你的网络有多快。您可以从简单地使用 ftp 传输 100MB 开始。这会给你一个想法,但是 oracle 使用不同的协议。接下来是跟踪会话。
  • @steve,添加网络速度信息:它可能比这里快两倍,但主要瓶颈是准备 CLOB 变量。
  • 行“rows.append((id,msg_arr))#???”是循环的一部分,这似乎不对。如果这没有帮助,是时候获取函数的代码覆盖率性能数据了。我不认为这个问题与oracle有关,而是与cx_oracle库的实现有关。
  • @steve,executemany() 采用数组,其长度表示操作重复的次数。如果我的代码不正确,有人知道如何正确操作吗?它 100% 与 Oracle 无关,因为类似的 Java 代码运行速度非常快。我也这么认为,这是由于 cx_Oracle 的实现。如果是这样,哪个库更适合我的情况?
  • 我看到的唯一选择是您自己修复库....

标签: performance python-3.x cx-oracle oracle12c


【解决方案1】:

奇怪的解决方法(感谢 cx_Oracle 邮件列表中的 Avinash Nandakumar),但它是在插入 CLOB 时大大提高性能的真正方法:

import time
import cx_Oracle
import sys

num_records = 100
con = cx_Oracle.connect('user/password@sid')
cur = con.cursor()
cur.bindarraysize = num_records
text = '$'*2**20    # 1 MB of text
rows = []

start_time = time.perf_counter()
cur.executemany(
    "insert into table_clob (msg_id, message) values (:msg_id, empty_clob())",
    [(i,) for i in range(1, 101)])
print('{} records prepared, {:.3f} s'
      .format(num_records, time.perf_counter() - start_time))

start_time = time.perf_counter()
selstmt = "select message from table_clob " +
          "where msg_id between 1 and :1 for update"
cur.execute(selstmt, [num_records])
for id in range(num_records):
    results = cur.fetchone()
    results[0].write(text)
con.commit()
print('{} records inserted, {:.3f} s'
      .format(num_records, time.perf_counter() - start_time))

cur.close()
con.close()

从语义上讲,这与我原来的帖子中的不完全相同,我想尽可能简单地举例说明原理。关键是你应该插入emptyclob(),然后选择它并写入它的内容。

【讨论】:

    猜你喜欢
    • 2011-10-01
    • 2012-11-12
    • 2019-05-02
    • 2018-08-05
    • 2012-12-26
    • 2013-12-22
    • 2023-03-13
    • 2021-07-17
    • 2018-01-16
    相关资源
    最近更新 更多