【问题标题】:Better way to convert file sizes in Python [closed]在 Python 中转换文件大小的更好方法 [关闭]
【发布时间】:2011-07-08 19:41:39
【问题描述】:

我正在使用一个库来读取文件并以字节为单位返回其大小。

这个文件大小然后显示给最终用户;为了让他们更容易理解,我通过将文件大小除以1024.0 * 1024.0 将其显式转换为MB。这当然可行,但我想知道在 Python 中是否有更好的方法来做到这一点?

更好的是,我的意思可能是一个 stdlib 函数,它可以根据我想要的类型来操作大小。就像我指定MB 一样,它会自动将其除以1024.0 * 1024.0。这些线路上的东西。

【问题讨论】:

  • 所以写一个。另请注意,许多系统现在使用 MB 来表示 10^6 而不是 2^20。
  • @A A, @tc:请记住,SI 和 IEC 标准是 kB (Kilo) for 1.000 Byte 和 KiB (Kibi) for 1.024 Byte。见en.wikipedia.org/wiki/Kibibyte。
  • @Bobby:kB 实际上的意思是“千贝”,等于 10000 dB。字节没有 SI 单位。 IIRC,IEC 推荐 KiB,但没有定义 kB 或 KB。
  • @tc。前缀kilo由SI定义为1000。IEC定义kB等使用SI前缀代替2^10。
  • 我的意思是前缀通常由SI定义,但数据大小的缩写不是:physics.nist.gov/cuu/Units/prefixes.html。这些由 IEC 定义:physics.nist.gov/cuu/Units/binary.html

标签: python filesize


【解决方案1】:

hurry.filesize 会以字节为单位获取大小并生成一个不错的字符串。

>>> from hurry.filesize import size
>>> size(11000)
'10K'
>>> size(198283722)
'189M'

或者如果你想要 1K == 1000(这是大多数用户的假设):

>>> from hurry.filesize import size, si
>>> size(11000, system=si)
'11K'
>>> size(198283722, system=si)
'198M'

它也有 IEC 支持(但没有记录):

>>> from hurry.filesize import size, iec
>>> size(11000, system=iec)
'10Ki'
>>> size(198283722, system=iec)
'189Mi'

因为它是由 Awesome Martijn Faassen 编写的,所以代码小、清晰且可扩展。编写自己的系统非常容易。

这是一个:

mysystem = [
    (1024 ** 5, ' Megamanys'),
    (1024 ** 4, ' Lotses'),
    (1024 ** 3, ' Tons'), 
    (1024 ** 2, ' Heaps'), 
    (1024 ** 1, ' Bunches'),
    (1024 ** 0, ' Thingies'),
    ]

这样使用:

>>> from hurry.filesize import size
>>> size(11000, system=mysystem)
'10 Bunches'
>>> size(198283722, system=mysystem)
'189 Heaps'

【讨论】:

  • 嗯,现在我需要一个去另一条路。从“1 kb”到1024(一个整数)。
  • 仅适用于 python 2
  • 这个包可能很酷,但奇怪的许可证以及没有在线可用源代码的事实使我很乐意避免使用它。而且它似乎只支持python2。
  • @AlmogCohen 源代码在线,可直接从 PyPI 获得(有些软件包没有 Github 存储库,只有一个 PyPI 页面)并且许可证并不那么晦涩,ZPL 是 Zope 公共许可证,它据我所知,它类似于 BSD。我同意许可本身很奇怪:没有标准的“LICENSE.txt”文件,每个源文件的顶部也没有序言。
  • 为了获得兆字节,我使用按位移位运算符做了以下等式:MBFACTOR = float(1 << 20); mb= int(size_in_bytes) / MBFACTOR@LennartRegebro
【解决方案2】:

您可以使用<< bitwise shifting operator 代替大小除数1024 * 1024,即1<<20 获取兆字节,1<<30 获取千兆字节等。

在最简单的情况下,您可以拥有例如一个常量MBFACTOR = float(1<<20),然后可以与字节一起使用,即:megas = size_in_bytes/MBFACTOR。

兆字节通常是您所需要的,或者可以使用其他类似的东西:

# bytes pretty-printing
UNITS_MAPPING = [
    (1<<50, ' PB'),
    (1<<40, ' TB'),
    (1<<30, ' GB'),
    (1<<20, ' MB'),
    (1<<10, ' KB'),
    (1, (' byte', ' bytes')),
]


def pretty_size(bytes, units=UNITS_MAPPING):
    """Get human-readable file sizes.
    simplified version of https://pypi.python.org/pypi/hurry.filesize/
    """
    for factor, suffix in units:
        if bytes >= factor:
            break
    amount = int(bytes / factor)

    if isinstance(suffix, tuple):
        singular, multiple = suffix
        if amount == 1:
            suffix = singular
        else:
            suffix = multiple
    return str(amount) + suffix

print(pretty_size(1))
print(pretty_size(42))
print(pretty_size(4096))
print(pretty_size(238048577))
print(pretty_size(334073741824))
print(pretty_size(96995116277763))
print(pretty_size(3125899904842624))

## [Out] ###########################
1 byte
42 bytes
4 KB
227 MB
311 GB
88 TB
2 PB

【讨论】:

  • 不是&gt;&gt;吗?
  • @Tjorriemorrie:它必须是左移,右移将丢弃唯一的位并导致0。
  • 出色的答案。谢谢。
  • 我知道这是旧的,但这是正确的用法吗? def convert_to_mb(data_b): print(data_b/(1
【解决方案3】:

这是我使用的:

import math

def convert_size(size_bytes):
   if size_bytes == 0:
       return "0B"
   size_name = ("B", "KB", "MB", "GB", "TB", "PB", "EB", "ZB", "YB")
   i = int(math.floor(math.log(size_bytes, 1024)))
   p = math.pow(1024, i)
   s = round(size_bytes / p, 2)
   return "%s %s" % (s, size_name[i])

注意:大小应该以字节为单位发送。

【讨论】:

  • 如果您以字节为单位发送大小,那么只需添加“B”作为 size_name 的第一个元素。
  • 当你有 0 大小的文件字节时,它会失败。 log(0, 1024) 未定义!您应该在此语句之前检查 0 字节大小写 i = int(math.floor(math.log(size,1024)))。
  • genclik - 你是对的。我刚刚提交了一个小修改,它将解决这个问题,并启用字节转换。谢谢,Sapam,原版
  • 嗨 @WHK,因为 tuxGurl 提到它很容易解决。
  • 实际上尺寸名称需要是 ("B", "KiB", "MiB", "GiB", "TiB", "PiB", "EiB", "ZiB", "乙”)。请参阅en.wikipedia.org/wiki/Mebibyte 了解更多信息。
【解决方案4】:

这是计算大小的紧凑函数

def GetHumanReadable(size,precision=2):
    suffixes=['B','KB','MB','GB','TB']
    suffixIndex = 0
    while size > 1024 and suffixIndex < 4:
        suffixIndex += 1 #increment the index of the suffix
        size = size/1024.0 #apply the division
    return "%.*f%s"%(precision,size,suffixes[suffixIndex])

更详细的输出反之操作请参考:http://code.activestate.com/recipes/578019-bytes-to-human-human-to-bytes-converter/

【讨论】:

  • while 语句应改为while size &gt;= 1024 and index &lt; len(suffixes):,否则函数将返回1024.0KB 而不是1.0MB。
【解决方案5】:

以防万一有人在寻找这个问题的反面(我确实这样做了),这对我有用:

def get_bytes(size, suffix):
    size = int(float(size))
    suffix = suffix.lower()

    if suffix == 'kb' or suffix == 'kib':
        return size << 10
    elif suffix == 'mb' or suffix == 'mib':
        return size << 20
    elif suffix == 'gb' or suffix == 'gib':
        return size << 30

    return False

【讨论】:

  • 您没有处理像 1.5GB 这样的十进制数字的情况。要修复它,只需将&lt;&lt; 10 更改为* 1024,&lt;&lt; 20 更改为* 1024**2 和&lt;&lt; 30 更改为* 1024**3。
【解决方案6】:

这是我的两分钱,它允许上下投射,并增加了可定制的精度:

def convertFloatToDecimal(f=0.0, precision=2):
    '''
    Convert a float to string of decimal.
    precision: by default 2.
    If no arg provided, return "0.00".
    '''
    return ("%." + str(precision) + "f") % f

def formatFileSize(size, sizeIn, sizeOut, precision=0):
    '''
    Convert file size to a string representing its value in B, KB, MB and GB.
    The convention is based on sizeIn as original unit and sizeOut
    as final unit. 
    '''
    assert sizeIn.upper() in {"B", "KB", "MB", "GB"}, "sizeIn type error"
    assert sizeOut.upper() in {"B", "KB", "MB", "GB"}, "sizeOut type error"
    if sizeIn == "B":
        if sizeOut == "KB":
            return convertFloatToDecimal((size/1024.0), precision)
        elif sizeOut == "MB":
            return convertFloatToDecimal((size/1024.0**2), precision)
        elif sizeOut == "GB":
            return convertFloatToDecimal((size/1024.0**3), precision)
    elif sizeIn == "KB":
        if sizeOut == "B":
            return convertFloatToDecimal((size*1024.0), precision)
        elif sizeOut == "MB":
            return convertFloatToDecimal((size/1024.0), precision)
        elif sizeOut == "GB":
            return convertFloatToDecimal((size/1024.0**2), precision)
    elif sizeIn == "MB":
        if sizeOut == "B":
            return convertFloatToDecimal((size*1024.0**2), precision)
        elif sizeOut == "KB":
            return convertFloatToDecimal((size*1024.0), precision)
        elif sizeOut == "GB":
            return convertFloatToDecimal((size/1024.0), precision)
    elif sizeIn == "GB":
        if sizeOut == "B":
            return convertFloatToDecimal((size*1024.0**3), precision)
        elif sizeOut == "KB":
            return convertFloatToDecimal((size*1024.0**2), precision)
        elif sizeOut == "MB":
            return convertFloatToDecimal((size*1024.0), precision)

根据需要添加TB等。

【讨论】:

  • 我会投票赞成,因为它可以通过 python 标准库来解决
【解决方案7】:

如果您已经知道自己想要的单位尺寸,这里有一些易于复制的单衬纸。如果您正在寻找具有一些不错选项的更通用的功能,请参阅我的 2021 年 2 月更新...

字节

print(f"{os.path.getsize(filepath):,} B") 

千比特

print(f"{os.path.getsize(filepath)/float(1<<7):,.0f} kb")

千字节

print(f"{os.path.getsize(filepath)/float(1<<10):,.0f} KB")

兆比特

print(f"{os.path.getsize(filepath)/float(1<<17):,.0f} mb")

兆字节

print(f"{os.path.getsize(filepath)/float(1<<20):,.0f} MB")

千兆

print(f"{os.path.getsize(filepath)/float(1

千兆字节

print(f"{os.path.getsize(filepath)/float(1<<30):,.0f} GB")

太字节

print(f"{os.path.getsize(filepath)/float(1<<40):,.0f} TB")

2021 年 2 月更新 这是我更新和充实的函数,用于 a) 获取文件/文件夹大小,b) 转换为所需的单位:

from pathlib import Path

def get_path_size(path = Path('.'), recursive=False):
    """
    Gets file size, or total directory size

    Parameters
    ----------
    path: str | pathlib.Path
        File path or directory/folder path

    recursive: bool
        True -> use .rglob i.e. include nested files and directories
        False -> use .glob i.e. only process current directory/folder

    Returns
    -------
    int:
        File size or recursive directory size in bytes
        Use cleverutils.format_bytes to convert to other units e.g. MB
    """
    path = Path(path)
    if path.is_file():
        size = path.stat().st_size
    elif path.is_dir():
        path_glob = path.rglob('*.*') if recursive else path.glob('*.*')
        size = sum(file.stat().st_size for file in path_glob)
    return size


def format_bytes(bytes, unit, SI=False):
    """
    Converts bytes to common units such as kb, kib, KB, mb, mib, MB

    Parameters
    ---------
    bytes: int
        Number of bytes to be converted

    unit: str
        Desired unit of measure for output


    SI: bool
        True -> Use SI standard e.g. KB = 1000 bytes
        False -> Use JEDEC standard e.g. KB = 1024 bytes

    Returns
    -------
    str:
        E.g. "7 MiB" where MiB is the original unit abbreviation supplied
    """
    if unit.lower() in "b bit bits".split():
        return f"{bytes*8} {unit}"
    unitN = unit[0].upper()+unit[1:].replace("s","")  # Normalised
    reference = {"Kb Kib Kibibit Kilobit": (7, 1),
                 "KB KiB Kibibyte Kilobyte": (10, 1),
                 "Mb Mib Mebibit Megabit": (17, 2),
                 "MB MiB Mebibyte Megabyte": (20, 2),
                 "Gb Gib Gibibit Gigabit": (27, 3),
                 "GB GiB Gibibyte Gigabyte": (30, 3),
                 "Tb Tib Tebibit Terabit": (37, 4),
                 "TB TiB Tebibyte Terabyte": (40, 4),
                 "Pb Pib Pebibit Petabit": (47, 5),
                 "PB PiB Pebibyte Petabyte": (50, 5),
                 "Eb Eib Exbibit Exabit": (57, 6),
                 "EB EiB Exbibyte Exabyte": (60, 6),
                 "Zb Zib Zebibit Zettabit": (67, 7),
                 "ZB ZiB Zebibyte Zettabyte": (70, 7),
                 "Yb Yib Yobibit Yottabit": (77, 8),
                 "YB YiB Yobibyte Yottabyte": (80, 8),
                 }
    key_list = '\n'.join(["     b Bit"] + [x for x in reference.keys()]) +"\n"
    if unitN not in key_list:
        raise IndexError(f"\n\nConversion unit must be one of:\n\n{key_list}")
    units, divisors = [(k,v) for k,v in reference.items() if unitN in k][0]
    if SI:
        divisor = 1000**divisors[1]/8 if "bit" in units else 1000**divisors[1]
    else:
        divisor = float(1 << divisors[0])
    value = bytes / divisor
    return f"{value:,.0f} {unitN}{(value != 1 and len(unitN) > 3)*'s'}"


# Tests 
>>> assert format_bytes(1,"b") == '8 b'
>>> assert format_bytes(1,"bits") == '8 bits'
>>> assert format_bytes(1024, "kilobyte") == "1 Kilobyte"
>>> assert format_bytes(1024, "kB") == "1 KB"
>>> assert format_bytes(7141000, "mb") == '54 Mb'
>>> assert format_bytes(7141000, "mib") == '54 Mib'
>>> assert format_bytes(7141000, "Mb") == '54 Mb'
>>> assert format_bytes(7141000, "MB") == '7 MB'
>>> assert format_bytes(7141000, "mebibytes") == '7 Mebibytes'
>>> assert format_bytes(7141000, "gb") == '0 Gb'
>>> assert format_bytes(1000000, "kB") == '977 KB'
>>> assert format_bytes(1000000, "kB", SI=True) == '1,000 KB'
>>> assert format_bytes(1000000, "kb") == '7,812 Kb'
>>> assert format_bytes(1000000, "kb", SI=True) == '8,000 Kb'
>>> assert format_bytes(125000, "kb") == '977 Kb'
>>> assert format_bytes(125000, "kb", SI=True) == '1,000 Kb'
>>> assert format_bytes(125*1024, "kb") == '1,000 Kb'
>>> assert format_bytes(125*1024, "kb", SI=True) == '1,024 Kb'

【讨论】:

  • 这是一个非常聪明的方法。我想知道您是否可以将这些放入一个函数中,您可以在其中传入是否需要 kb。 mb之类的。你甚至可以有一个输入命令来询问你想要哪个,如果你经常这样做会很方便。
  • 见上文,Hildy...您还可以自定义字典行,如上面概述的@lennart-regebro...这可能对存储管理很有用,例如“分区”、“集群”、“4TB 磁盘”、“DVD_RW”、“蓝光光盘”、“1GB 记忆棒”等等。
  • 我还刚刚添加了 Kb (Kilobit)、Mb (Megabit) 和 Gb (Gigabit) - 用户经常对网络或文件传输速度感到困惑,所以认为可能是方便。
  • 我喜欢单线,考虑用 f-strings 压缩,例如:f'{os.path.getsize(filepath)/float(1&lt;&lt;20):.0f} MB'
  • 非常感谢@pan0ramic!
【解决方案8】:

这是一个与 ls -lh 的输出相匹配的版本。

def human_size(num: int) -> str:
    base = 1
    for unit in ['B', 'K', 'M', 'G', 'T', 'P', 'E', 'Z', 'Y']:
        n = num / base
        if n < 9.95 and unit != 'B':
            # Less than 10 then keep 1 decimal place
            value = "{:.1f}{}".format(n, unit)
            return value
        if round(n) < 1000:
            # Less than 4 digits so use this
            value = "{}{}".format(round(n), unit)
            return value
        base *= 1024
    value = "{}{}".format(round(n), unit)
    return value

【讨论】:

    【解决方案9】:

    这是我的实现:

    from bisect import bisect
    
    def to_filesize(bytes_num, si=True):
        decade = 1000 if si else 1024
        partitions = tuple(decade ** n for n in range(1, 6))
        suffixes = tuple('BKMGTP')
    
        i = bisect(partitions, bytes_num)
        s = suffixes[i]
    
        for n in range(i):
            bytes_num /= decade
    
        f = '{:.3f}'.format(bytes_num)
    
        return '{}{}'.format(f.rstrip('0').rstrip('.'), s)
    

    它将打印最多三位小数,并去除尾随的零和句点。布尔参数 si 将切换使用基于 10 与基于 2 的大小大小。

    这是它的对应物。它允许编写干净的配置文件,如{'maximum_filesize': from_filesize('10M')。它返回一个近似于预期文件大小的整数。我没有使用位移,因为源值是一个浮点数(它会接受from_filesize('2.15M') 就好了)。将其转换为整数/小数会起作用,但会使代码更加复杂,而且它已经按原样工作了。

    def from_filesize(spec, si=True):
        decade = 1000 if si else 1024
        suffixes = tuple('BKMGTP')
    
        num = float(spec[:-1])
        s = spec[-1]
        i = suffixes.index(s)
    
        for n in range(i):
            num *= decade
    
        return int(num)
    

    【讨论】:

      【解决方案10】:

      这里是:

      def convert_bytes(size):
          for x in ['bytes', 'KB', 'MB', 'GB', 'TB']:
              if size < 1024.0:
                  return "%3.1f %s" % (size, x)
              size /= 1024.0
      
          return size
      

      输出

      >>> convert_bytes(1024)
      '1.0 KB'
      >>> convert_bytes(102400)
      '100.0 KB'
      

      【讨论】:

      • 那个 MiB,而不是 MB 等等......
      【解决方案11】:
      UNITS = {1000: ['KB', 'MB', 'GB'],
                  1024: ['KiB', 'MiB', 'GiB']}
      
      def approximate_size(size, flag_1024_or_1000=True):
          mult = 1024 if flag_1024_or_1000 else 1000
          for unit in UNITS[mult]:
              size = size / mult
              if size < mult:
                  return '{0:.3f} {1}'.format(size, unit)
      
      approximate_size(2123, False)
      

      【讨论】:

      • 这在很多环境中都可以使用。很高兴我看到了这个评论。非常感谢。
      • 是的,这很可爱,不需要外部库
      【解决方案12】:

      我想要 2 路转换,并且我想使用 Python 3 format() 支持来最 Pythonic。也许尝试数据大小库模块? https://pypi.org/project/datasize/

      $ pip install -qqq datasize
      $ python
      ...
      >>> from datasize import DataSize
      >>> 'My new {:GB} SSD really only stores {:.2GiB} of data.'.format(DataSize('750GB'),DataSize(DataSize('750GB') * 0.8))
      'My new 750GB SSD really only stores 558.79GiB of data.'
      

      【讨论】:

        猜你喜欢
        • 2013-05-24
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2021-04-17
        • 1970-01-01
        • 1970-01-01
        • 2011-12-04
        相关资源
        最近更新 更多