【发布时间】:2023-03-21 02:55:02
【问题描述】:
有没有人知道什么是更好地使用考虑速度和资源?链接到一些受信任的来源将不胜感激。
if key not in dictionary.keys():
或
if not dictionary.get(key):
【问题讨论】:
标签: python performance dictionary
有没有人知道什么是更好地使用考虑速度和资源?链接到一些受信任的来源将不胜感激。
if key not in dictionary.keys():
或
if not dictionary.get(key):
【问题讨论】:
标签: python performance dictionary
首先,你会这样做
if key not in dictionary:
因为字典是通过键迭代的。
其次,这两个语句是不等价的 - 如果相应的值是虚假的(0、""、[] 等),则第二个条件为真,不仅在密钥不存在的情况下。
最后,第一种方法肯定更快,更pythonic。函数/方法调用很昂贵。如果您不确定,timeit。
【讨论】:
try/except 块也可能是实现这一目标的最佳方式。
try/except 可能是最好的方法?
根据我的经验,使用in 比使用get 更快,尽管可以通过缓存get 方法来提高get 的速度,因此不必每次都查找它。以下是一些timeit 测试:
''' in vs get speed test
Comparing the speed of cache retrieval / update using `get` vs using `in`
http://stackoverflow.com/a/35451912/4014959
Written by PM 2Ring 2015.12.01
Updated for Python 3 2017.08.08
'''
from __future__ import print_function
from timeit import Timer
from random import randint
import dis
cache = {}
def get_cache(x):
''' retrieve / update cache using `get` '''
res = cache.get(x)
if res is None:
res = cache[x] = x
return res
def get_cache_defarg(x, get=cache.get):
''' retrieve / update cache using defarg `get` '''
res = get(x)
if res is None:
res = cache[x] = x
return res
def in_cache(x):
''' retrieve / update cache using `in` '''
if x in cache:
return cache[x]
else:
res = cache[x] = x
return res
#slow to fast.
funcs = (
get_cache,
get_cache_defarg,
in_cache,
)
def show_bytecode():
for func in funcs:
fname = func.__name__
print('\n%s' % fname)
dis.dis(func)
def time_test(reps, loops):
''' Print timing stats for all the functions '''
for func in funcs:
fname = func.__name__
print('\n%s: %s' % (fname, func.__doc__))
setup = 'from __main__ import data, ' + fname
cmd = 'for v in data: %s(v)' % (fname,)
times = []
t = Timer(cmd, setup)
for i in range(reps):
r = 0
for j in range(loops):
r += t.timeit(1)
cache.clear()
times.append(r)
times.sort()
print(times)
datasize = 1024
maxdata = 32
data = [randint(1, maxdata) for i in range(datasize)]
#show_bytecode()
time_test(3, 500)
典型输出在我运行 Python 2.6.6 的 2Ghz 机器上:
get_cache: retrieve / update cache using `get`
[0.65624237060546875, 0.68499755859375, 0.76354193687438965]
get_cache_defarg: retrieve / update cache using defarg `get`
[0.54204297065734863, 0.55032730102539062, 0.56702113151550293]
in_cache: retrieve / update cache using `in`
[0.48754477500915527, 0.49125504493713379, 0.50087881088256836]
【讨论】:
TLDR:使用if key not in dictionary。这是惯用的、健壮的和快速的。
与此问题相关的版本有四个:问题中提出的 2 个版本,以及它们的最佳变体:
key not in dictionary.keys() # inA
key not in dictionary # inB
not dictionary.get(key) # getA
sentinel = object()
dictionary.get(key, sentinel) is not sentinel # getB
A 两种变体都有缺点,意味着您不应该使用它们。 inA 不必要地在键上创建一个 dict 视图 - 这增加了一个间接步骤。 getA 查看值的真值 - 这会导致 '' 或 0 等值的错误结果。
至于使用inB 而不是getB:两者都做同样的事情,即查看key 是否存在值。但是,getB 也 返回该值或默认值,并且必须将其与哨兵进行比较。因此,使用get 会慢很多:
$ PREPARE="
> import random
> data = {a: True for a in range(0, 512, 2)}
> sentinel=object()"
$ python3 -m perf timeit -s "$PREPARE" '27 in data'
.....................
Mean +- std dev: 33.9 ns +- 0.8 ns
$ python3 -m perf timeit -s "$PREPARE" 'data.get(27, sentinel) is not sentinel'
.....................
Mean +- std dev: 105 ns +- 5 ns
请注意,一旦 JIT 预热,pypy3 对两种变体的性能几乎相同。
【讨论】:
好的,我已经在 python 3.4.3 上对其进行了测试,所有三种方法都在 0.00001 秒左右给出了相同的结果。
import random
a = {}
for i in range(0, 1000000):
a[str(random.random())] = random.random()
import time
t1 = time.time(); 1 in a.keys(); t2 = time.time(); print("Time=%s" % (t2 - t1))
t1 = time.time(); 1 in a; t2 = time.time(); print("Time=%s" % (t2 - t1))
t1 = time.time(); not a.get(1); t2 = time.time(); print("Time=%s" % (t2 - t1))
【讨论】:
timeit module。
a[1]。
xrange()进行迭代