【问题标题】:Haskell Program Out of Memory (Infinite Recursion? Loop? Something?)Haskell 程序内存不足(无限递归?循环?什么?)
【发布时间】:2014-03-24 12:15:45
【问题描述】:

编辑:更新以包含整个代码。

我对 Haskell 还是很陌生,我编写的一个程序有问题,用于为课程作业做一些熵计算(作业就是计算,使用 Haskell 是一种选择,所以我我不要求有人为我做作业,用 Python 做这件事会花费我微不足道的时间和精力)。代码采用一维数组:

--- first input (length 2): 
---     0,0   0,1   1,0   1,1
---    [.48,  .02,  .02,  .48]
--- or:
---     0    1   
---    .48  .02  0
---               
---    .02  .48  1

然后我定义了几个通用函数:

log2 :: Float -> Float
log2 x =
  logBase 2 x

entropy :: [Float] -> Float
entropy probArray =
  sum(map (\i -> (i * (log2 (1/i)))) probArray)

以及每个特定计算的函数:

-- calculate joint entropy
jointEntropy :: [Float] -> Float
jointEntropy probArray =
  entropy probArray

-- calculate entropy of X
splitByCol :: Int -> [Float] -> [[Float]]
splitByCol length probArray =
  [(take length probArray)] ++ (splitByCol length (drop length probArray))

xEntropy :: Int -> [Float] -> Float
xEntropy length probArray =
  entropy (map sum (splitByCol length probArray))

-- calculate entropy of Y
ithElements :: Int -> Int -> [Float] -> [Float]
ithElements level length matrixArray =
  let indexArray = zip [0..(length^2 - 1)] matrixArray
  in [snd x | x <- indexArray, fst x `mod` length == level]

splitByRow :: Int -> Int -> [[Float]] -> [[Float]]
splitByRow level length lists =
  if level == length
  then
    tail lists -- return list sans full matrix array which was being carried at the front
  else
    splitByRow (level+1) length (lists ++ [(ithElements level length (lists !! 0))]) 

yEntropy :: Int -> [Float] -> Float
yEntropy length probArray =
  entropy (map sum (splitByRow 0 length [probArray]))

--calculate mutual information
mutualInfo :: Float -> Float -> Float
mutualInfo xEnt yEnt =
  xEnt - yEnt

-- calculate conditional of X given Y - (X|Y)
xCond :: Float -> Float -> Float
xCond xEnt mInfo =
  xEnt - mInfo

-- calculate conditional of Y given X - (Y|X)
yCond :: Float -> Float -> Float
yCond yEnt mInfo =
  yEnt - mInfo

然后将它们全部链接在一起以返回一个包含我想要执行的每个计算的数组:

-- caller functions -> resArray ends up looking like [H(X,Y), H(X), H(Y), I(X;Y), H(X|Y), H(Y|X)]
calcJointEnt :: [Float] -> [Float]
calcJointEnt probArray =
  calcVarEnt probArray [(jointEntropy probArray)]

calcVarEnt :: [Float] -> [Float] -> [Float]
calcVarEnt probArray resArray =
  let len = floor (sqrt (fromIntegral (length probArray)))
  in calcMutual probArray (resArray ++ [(xEntropy len probArray), (yEntropy len probArray)])

calcMutual :: [Float] -> [Float] -> [Float]
calcMutual probArray resArray =
  calcCond probArray (resArray ++ [(mutualInfo (resArray !! 1) (resArray !! 2))])

calcCond :: [Float] -> [Float] -> [Float]
calcCond probArray resArray =
  resArray ++ [(xCond (resArray !! 1) (resArray !! 3)), (yCond (resArray !! 2) (resArray !! 3))]

等等...然后我有一些函数来格式化打印字符串,还有一个 main 函数将它们组合在一起:

-- prepare printout
statString :: (String, String) -> String
statString t =  
  (fst t) ++ ": " ++ (snd t)

printOut :: [Float] -> String
printOut resArray =
  let statArray = zip ["H(X,Y)", "H(X)", "H(Y)", "H(X;Y)", "H(X|Y)", "H(Y|X)"] (map show resArray)
  in "results:\n\t" ++ intercalate "\n\t" (map statString statArray) ++ "\n\n---\n"

-- main
main :: IO()
main = 
  let inputs = [[0.48,  0.02,  0.02,  0.48], [0.31,  0.02,  0.00,  0.02,  0.32,  0.02,  0.00,  0.02,  0.29]]
  in putStrLn (intercalate "" (map printOut (map calcJointEnt inputs)))

所以我确信有更好的方法来做很多这件事,但在我看来,从我最小的 haskell 经验和我稍微更广泛但仍然有限的函数式风格的编程经验来看,它应该可以工作。

我的问题是,当我编译和运行时,我得到这个输出:

bash-4.2$ ./noise 
results:
    H(X,Y): 1.2422923
noise: out of memory (requested 1048576 bytes)

在打印出一个结果和内存错误消息之间有很长的时间。当我在 ghci 调试器(我第一次使用它)中弹出它时,如果我尝试在 printOut 函数中强制使用 resArray,它会执行相同的操作,并且当我尝试在链接函数的最低级别:

calcCond :: [Float] -> [Float] -> [Float]
calcCond probArray resArray =
  resArray ++ [(xCond (resArray !! 1) (resArray !! 3)), (yCond (resArray !! 2) (resArray !! 3))]

我得到以下信息:

[noise.hs:101:3-96] *Main> seq _t1 ()
()
[noise.hs:101:3-96] *Main> :print resArray
resArray = (_t2::Float) : (_t3::[Float])
[noise.hs:101:3-96] *Main> seq _t2 ()
()
[noise.hs:101:3-96] *Main> :print resArray
resArray = 1.2422923 : (_t4::[Float])
[noise.hs:101:3-96] *Main> seq _t3 ()
()
[noise.hs:101:3-96] *Main> :print resArray
resArray = 1.2422923 : (_t5::Float) : (_t6::[Float])
[noise.hs:101:3-96] *Main> seq _t5 ()
^C^C^C^C^CInterrupted.
[noise.hs:101:3-96] *Main> 

我研究了 RTS 调试工具,它似乎是推荐的工具,用于在网站上类似提出的问题中弹出类似问题的引擎盖,但是当我使用 +RTS -xc 运行它时,什么也没发生。我认为这是因为 RTS 似乎要求它实际抛出异常,而不是操作系统介入?

我认为我自己的主要问题来自命令式背景,即程序可以通过某种无限循环过程到达 IO 语句的概念仍然在逻辑上的某个地方进行,这是一个陌生的概念。当然,我可能完全不正确,这就是正在发生的事情,但在我看来就是这样。非常感谢大家可以提供的任何帮助(不仅在此代码上,而且在我对 Haskell 的方法上也能提供一般帮助)。

【问题讨论】:

  • 1.您能否在问题中包含您如何编译代码(请务必尝试使用-O2)。 2. 请包含完整的、可编译的示例。您发布的代码缺少功能。
  • @ThomasM.DuBuisson - 我只是在使用 >>> ghc noise.hs 并尝试使用 -O2 并没有改变任何东西。我不想发布整个代码以使其可读,但我现在会添加它。
  • 因为H(X) 从未被打印出来,所以看看它的计算位置是有意义的,即xEntropy。 xEntropy 调用 splitByCol 有一个明显的错误。它永远不会终止!
  • 奥古瓦尔特。哇。那与Haskell没有任何关系,那只是我很愚蠢。谢谢。
  • @TomEllis 既然您回答了这个问题,也许您可​​以发布该评论作为答案?

标签: haskell infinite


【解决方案1】:

由于从未打印过H(X),因此查看它的计算位置是有意义的,即xEntropy。 xEntropy 调用 splitByCol 有一个明显的错误。它返回一个无限列表!这意味着entropy 永远不会终止,因为它会尝试在无限列表上调用sum。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2020-01-10
    • 1970-01-01
    • 2016-10-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多