【发布时间】:2015-02-05 10:40:17
【问题描述】:
我在 while 循环中从 UDP 套接字读取数据。我需要最有效的方法
1) 读取数据 (*)(这有点解决了,但赞赏 cmets)
2) 定期将(操纵的)数据转储到文件中 (**)(问题)
我预计 numpy 的“tostring”方法会出现瓶颈。让我们考虑以下一段(不完整的)代码:
import socket
import numpy
nbuf=4096
buf=numpy.zeros(nbuf,dtype=numpy.uint8) # i.e., an array of bytes
f=open('dump.data','w')
datasocket=socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
# ETC.. (code missing here) .. the datasocket is, of course, non-blocking
while True:
gotsome=True
try:
N=datasocket.recv_into(buf) # no memory-allocation here .. (*)
except(socket.error):
# do nothing ..
gotsome=False
if (gotsome):
# the bytes in "buf" will be manipulated in various ways ..
# the following write is done frequently (not necessarily in each pass of the while loop):
f.write(buf[:N].tostring()) # (**) The question: what is the most efficient way to do this?
f.close()
现在,在 (**),据我所知:
1) buf[:N] 为一个长度为 N+1 的新数组对象分配内存,对吧? (也许不是)
.. 然后:
2) buf[:N].tostring() 为一个新的字符串分配内存,来自buf的字节被复制到这个字符串中
这似乎有很多内存分配和交换。在同一个循环中,以后我会读取几个socket,写入几个文件。
有没有办法只告诉 f.write 直接访问“buf”的内存地址从 0 到 N 字节并将它们写入磁盘?
即,本着缓冲区接口的精神这样做并避免这两个额外的内存分配?
P。 S. f.write(buf[:N].tostring()) 等价于 buf[:N].tofile(f)
【问题讨论】:
-
你分析过你的代码吗?我可以想象,在写入文件时,一些额外的分配不会造成太大的伤害。
标签: python arrays numpy buffer