【问题标题】:Shannon-Fano code as max-heap in pythonShannon-Fano 代码作为 python 中的最大堆
【发布时间】:2015-02-11 08:49:52
【问题描述】:

我有一个用于 Huffman 编码的最小堆代码,您可以在此处查看:http://rosettacode.org/wiki/Huffman_coding#Python

我正在尝试制作一个类似于 min-heap 的最大堆 Shannon-Fano 代码。

这是一个代码:

from collections import defaultdict, Counter
import heapq, math

def _heappop_max(heap):
"""Maxheap version of a heappop."""
lastelt = heap.pop()    # raises appropriate IndexError if heap is empty
if heap:
    returnitem = heap[0]
    heap[0] = lastelt
    heapq._siftup_max(heap, 0)
    return returnitem
return lastelt

def _heappush_max(heap, item):
    """Push item onto heap, maintaining the heap invariant."""
    heap.append(item)
    heapq._siftdown_max(heap, 0, len(heap)-1)

def sf_encode(symb2freq):
heap = [[wt, [sym, ""]] for sym, wt in symb2freq.items()]
heapq._heapify_max(heap)
while len(heap) > 1:
    lo = _heappop_max(heap)
    hi = _heappop_max(heap)
    for pair in lo[1:]:
        pair[1] = '0' + pair[1]
    for pair in hi[1:]:
        pair[1] = '1' + pair[1]
    _heappush_max(heap, [lo[0] + hi[0]] + lo[1:] + hi[1:])
print heap
return sorted(_heappop_max(heap)[1:], key=lambda p: (len(p[1]), p))

但我有这样的输出:

Symbol  Weight  Shannon-Fano Code
!   1   1
3   1   01
:   1   001
J   1   0001
V   1   00001
z   1   000001
E   3   0000001
L   3   00000001
P   3   000000001
N   4   0000000001
O   4   00000000001

我是否正确使用 heapq 来实现 Shannon-Fano 编码?这个字符串的问题:

_heappush_max(heap, [lo[0] + hi[0]] + lo[1:] + hi[1:])

我不明白如何解决它。

期望输出类似于 Huffman 编码

Symbol  Weight  Huffman Code
    2875    01
a   744 1001
e   1129    1110
h   606 0000
i   610 0001
n   617 0010
o   668 1000
t   842 1100
d   358 10100
l   326 00110

添加:

好吧,我尝试在没有 heapq 的情况下执行此操作,但递归无法停止:

def sf_encode(iA, iB, maxP):
global tupleList, total_sf
global mid
maxP = maxP/float(2)
sumP = 0    
for i in range(iA, iB):
    tup = tupleList[i]
    if sumP < maxP or i == iA: # top group
        sumP += tup[1]/float(total_sf)
        tupleList[i] = (tup[0], tup[1], tup[2] + '0')
        mid = i           
    else: # bottom group
        tupleList[i] = (tup[0], tup[1], tup[2] + '1')
print tupleList
if mid - 1 > iA:
    sf_encode(iA, mid - 1, maxP)
if iB - mid > 0:
    sf_encode(mid, iB, maxP)
return tupleList

【问题讨论】:

  • 你希望你的输出是什么?
  • @ScottHunter,我已将其添加到我的问题中。

标签: python algorithm encoding heap binary-tree


【解决方案1】:

在 Shannon-Fano 编码中,您需要以下 steps:

Shannon–Fano 树是根据设计的规范构建的 定义一个有效的代码表。实际算法很简单:

  1. 对于给定的符号列表,制定相应的符号列表 概率或频率计数,以便每个符号的相对 发生频率是已知的。
  2. 根据符号列表排序 频率,最常出现的符号在左边 最不常见的在右边。
  3. 将列表分为两部分, 左侧部分的总频率计数与 尽可能的总权利。
  4. 分配列表的左侧部分 二进制数字 0,右侧部分分配数字 1。这 表示第一部分中的符号代码将全部开始 0,第二部分的代码都以1开头。
  5. 递归地将步骤 3 和 4 应用于两半中的每一个, 细分组并将位添加到代码中,直到每个符号都有 成为树上对应的代码叶。

因此,您将需要代码进行排序(您的输入似乎已经排序,因此您可以跳过此步骤),加上一个递归函数,该函数选择最佳分区,然后在列表的前半部分和后半部分进行递归。

一旦列表被排序,元素的顺序永远不会改变,所以不需要使用 heapq 来做这种编码风格。

【讨论】:

  • 谢谢。我知道这个算法,但我正在尝试用 heapq lib 来做到这一点。我试过没有它,但输出代码就像统一编码一样。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2016-01-06
  • 1970-01-01
  • 1970-01-01
  • 2019-01-12
  • 1970-01-01
  • 2023-04-10
相关资源
最近更新 更多