【问题标题】:Precision of numpy array lost after tolisttolist后numpy数组的精度丢失
【发布时间】:2021-05-28 07:48:28
【问题描述】:

我有一个 numpy 数组,其中每个数字都有一定的指定精度(使用 around(x,1)。

[[     3.   15294.7  32977.7   4419.5    978.4    504.4    123.6]
 [     4.   14173.8  31487.2   3853.9    967.8    410.2    107.1]
 [     5.   15323.5  34754.5   3738.7   1034.7    376.1    105.5]
 [     6.   17396.7  41164.5   3787.4   1103.2    363.9    109.4]
 [     7.   19665.5  48967.6   3900.9   1161.     362.1    115.8]
 [     8.   21839.8  56922.5   4037.4   1208.2    365.9    123.5]
 [     9.   23840.6  64573.8   4178.1   1247.     373.2    131.9]
 [    10.   25659.9  71800.2   4314.8   1279.5    382.7    140.5]
 [    11.   27310.3  78577.7   4444.3   1307.1    393.7    149.1]
 [    12.   28809.1  84910.4   4565.8   1331.     405.5    157.4]]

我正在尝试将每个数字转换为字符串,以便我可以使用 python-docx 将它们写入单词表。但是 tolist() 函数的结果是一团糟。数字的精度会丢失,导致输出很长。

[['3.0',
  '15294.7001953',
  '32977.6992188',
  '4419.5',
  '978.400024414',
  '504.399993896',
  '123.599998474'],
 ['4.0',
  '14173.7998047',
  '31487.1992188',
  '3853.89990234',
  '967.799987793',
  '410.200012207',
  '107.099998474'],
.......

除了 tolist() 函数,我还尝试了 [[str(e) for e in a] for a in m]。结果是一样的。这很烦人。如何在保持精度的同时轻松转换为字符串?谢谢!

【问题讨论】:

  • 你的数组是单精度(np.float32)吗?
  • 是的,它是 float32。有问题吗?
  • 查看我的答案,或@HenryGomersall 的答案

标签: numpy


【解决方案1】:

转换为字符串时出现问题。只有数字:

>>> import numpy as np
>>> a = np.random.random(10)*30
>>> a
array([ 27.30713434,  10.25895255,  19.65843272,  23.93161555,
        29.08479175,  25.69713898,  11.90236158,   5.41050686,
        18.16481691,  14.12808414])
>>> 
>>> b = np.round(a, decimals=1)
>>> b
array([ 27.3,  10.3,  19.7,  23.9,  29.1,  25.7,  11.9,   5.4,  18.2,  14.1])
>>> b.tolist()
[27.3, 10.3, 19.7, 23.9, 29.1, 25.7, 11.9, 5.4, 18.2, 14.1]

请注意np.round 不能就地工作:

>>> a
array([ 27.30713434,  10.25895255,  19.65843272,  23.93161555,
        29.08479175,  25.69713898,  11.90236158,   5.41050686,
        18.16481691,  14.12808414])

如果您只需要将数字转换为字符串:

>>> " ".join(str(_) for _ in np.round(a, 1)) 
'27.3 10.3 19.7 23.9 29.1 25.7 11.9 5.4 18.2 14.1'

编辑:显然,np.round 与 float32 不匹配(其他答案给出了原因)。一个简单的解决方法是将数组显式转换为 np.float 或 np.float64 或只是 float:

>>> # prepare an array of float32 values
>>> a32  = (np.random.random(10) * 30).astype(np.float32)
>>> a32.dtype
dtype('float32')
>>> 
>>> # notice the use of .astype(np.float32)
>>> np.round(a32.astype(np.float64), 1)
array([  5.5,   8.2,  29.8,   8.6,  15.5,  28.3,   2. ,  24.5,  18.4,   8.3])
>>> 

EDIT2:正如 Warren 在他的回答中所证明的,字符串格式实际上可以正确地四舍五入(尝试"%.1f" % (4.79,))。因此不需要在浮点类型之间进行转换。我将留下我的答案主要是为了提醒您在这种情况下使用np.around 不是正确的做法。

【讨论】:

  • 感谢您的回复,但我仍然无法做到这一点。我简单地使用了 np.around(x, 1)。但是我在每个浮点数上都有很长的尾巴。喜欢:数组([ 448.3999939 , 521.59997559, 581.70001221, 635.40002441, 688.79998779, 746. , 808. , 872.40002441, 935.90002441, 935.90002441, 996.40002]4)4, 996.40002]3
【解决方案2】:

精度没有“丢失”;你一开始就没有精确度。 值 15294.7 不能用单精度精确表示(即 np.float32);最佳近似值 是 15294.70019...:

In [1]: x = np.array([15294.7], dtype=np.float32)

In [2]: x
Out[2]: array([ 15294.70019531], dtype=float32)

见http://floating-point-gui.de/

使用 np.float64 可以得到更好的近似值,但它仍然不能准确地表示 15294.7。

如果您想要使用单个十进制数字格式化的文本输出,请使用专为格式化文本输出设计的函数,例如np.savetxt:

In [56]: x = np.array([[15294.7, 32977.7],[14173.8, 31487.2]], dtype=np.float32) 

In [57]: x
Out[57]: 
array([[ 15294.70019531,  32977.69921875],
       [ 14173.79980469,  31487.19921875]], dtype=float32)

In [58]: np.savetxt("data.txt", x, fmt="%.1f", delimiter=",")

In [59]: !cat data.txt
15294.7,32977.7
14173.8,31487.2

如果你真的需要一个格式良好的 numpy 数组,你可以这样做:

In [63]: def myfmt(r):
   ....:     return "%.1f" % (r,)
   ....: 

In [64]: vecfmt = np.vectorize(myfmt)

In [65]: vecfmt(x)
Out[65]: 
array([['15294.7', '32977.7'],
       ['14173.8', '31487.2']], 
      dtype='|S64')

如果您使用其中任何一种方法,则无需先通过around 传递数据;舍入将作为格式化过程的一部分进行。

【讨论】:

  • (+1) 我的印象是字符串格式会截断,而不是舍入。谢谢!
  • 感谢您的解释。我的矩阵由几个一维数组合并而成,每个数组都可能有不同的精度要求。这就是为什么在最终显示过程中不能使用 "%.1f" % (r,) 来中继它们的原因。我现在切换到 float64,它工作正常,但我担心可能需要比 float32 更多的内存,因为我的数据可能很大。
  • can not 通常写成cannot :)
【解决方案3】:

浮点数非常擅长存储具有明确定义的相对精度的大范围。在 32 位浮点数的情况下,这大约是 7 个有效数字。正如您所注意到的,您在进行四舍五入时得到的实际数字并不完全是您希望的数字,而是接近 7 位有效数字。

获得所需内容的一种方法可能是使用decimal.Decimal type。您可以通过将 dtype 设置为该类型来构造一个 numpy 数组:

import decimal
a = numpy.array(original_array, dtype=decimal.Decimal)

注意,结果数组只是一个 python 对象数组,而不是一个“正确的”numpy 数组,所以你可能需要滚动你自己的舍入函数,也许还有其他一些不起作用的东西。

最好只处理内置的python结构以获得你想要的。

【讨论】:

    【解决方案4】:

    即使您一开始就无法控制 numpy float32 数组中的数据,您也可以将类型更改为更高的精度,然后在调用 tolist 之前进行舍入。实际上,您甚至可以使用astype 进行字符串转换。例如:

    >>> import numpy as np
    >>> a = np.array([[    3.0, 15294.7, 32977.7],
                      [ 4419.5,   978.4,   504.4]])
    >>> a.astype(float).round(1).astype(str).tolist()
    [['3.0', '15294.7', '32977.7'], ['4419.5', '978.4', '504.4']]
    

    【讨论】:

      【解决方案5】:

      所有答案都正确地讨论了浮点精度和输出,但我想补充一点,您首先不需要从 np.array 转换为使用 tolist 的列表。事实上,您很少需要执行该操作,因为 numpy 数组的行为通常非常相似,如下例所示:

      import docx
      import numpy as np
      
      # Your values from above
      raw_data = np.array([[ 3., 15294.7, 32977.7, 4419.5,  978.4, 504.4, 123.6],
                           [ 4., 14173.8, 31487.2, 3853.9,  967.8, 410.2, 107.1],
                           [ 5., 15323.5, 34754.5, 3738.7, 1034.7, 376.1, 105.5],
                           [ 6., 17396.7, 41164.5, 3787.4, 1103.2, 363.9, 109.4],
                           [ 7., 19665.5, 48967.6, 3900.9, 1161.0, 362.1, 115.8],
                           [ 8., 21839.8, 56922.5, 4037.4, 1208.2, 365.9, 123.5],
                           [ 9., 23840.6, 64573.8, 4178.1, 1247.0, 373.2, 131.9],
                           [10., 25659.9, 71800.2, 4314.8, 1279.5, 382.7, 140.5],
                           [11., 27310.3, 78577.7, 4444.3, 1307.1, 393.7, 149.1],
                           [12., 28809.1, 84910.4, 4565.8, 1331.0, 405.5, 157.4]],
                          dtype=np.float32)
      
      # This conversion is just for comparison purposes, both tables will be printed.
      pyt_data = raw_data.tolist()
      
      def create_table(document, values, heading):
          """Creates a docx table inside the document.
      
          This function takes a docx.Document, a two-dimensional data structure, e.g.
          numpy arrays or a list of lists, and fills the table with it.
          The table is also prefixed with a heading.
          """
          document.add_heading(heading)
          table = document.add_table(rows=0, cols=len(values[0]))
          for row in values:
              cells = table.add_row().cells
              for i, value in enumerate(row):
                  # Use `str` for any types, but the format string 
                  # only if you expect numerical types exclusively
                  cells[i].text = str(value)  # f'{value:.1f}'
      
      document = docx.Document()
      create_table(document, raw_data, 'Raw table')
      create_table(document, pyt_data, 'tolist table')
      document.save('table_demo.docx')
      

      如果您将注释行 cells[i].text = str(value) 更改为 cells[i].text = f'{value:.1f'}(或者如果使用 Python cells[i].text = '{:.1f}'.format(value)),则两个表都可以正常工作,因为您使用自定义格式格式化浮点值。如果只使用字符串表示,numpy 的值已经是正确的了。

      注意,如果你使用np.float64,两个版本都是正确的!

      使用字符串表示,生成的 docx 呈现如下:

      使用格式字符串/格式化字符串文字,生成的 docx 如下所示:

      【讨论】:

        猜你喜欢
        • 2015-07-29
        • 2016-10-26
        • 2015-05-16
        • 2013-06-17
        • 1970-01-01
        • 2018-06-17
        • 2012-05-13
        • 2013-07-25
        • 1970-01-01
        相关资源
        最近更新 更多