【发布时间】:2018-02-10 15:24:15
【问题描述】:
我创建一个记录类型如下:
type CombEmp =
{
empid:int
empname:string
email:string
}
let defCombEmp:CombEmp =
{
empid = 0
empname = ""
email = ""
}
然后我创建记录实例:
let chrE1 = {defCombEmp with empid = 100; empname = "Wayne Rooney"; email = "wroo@mun.com"}
let chrE2 = {defCombEmp with empid = 100; empname = "Wayne R"; email = "war@mun.com"}
let chrE3 = {defCombEmp with empid = 100; empname = "Wayne R"; email = "rooney@mun.com"}
然后我使用上述实例创建一个包含 65-70 条记录的列表:
let hrLst = [|chrE1;chrE2;chrE3;chrE1;chrE2;chrE3;chrE1; ...... |] |> Array.toList
我现在已经编写了如下所示的代码。
函数GetByteCount获取序列化数据的大小(这里使用NewtonSoft)。
函数loop1 使用上面的hrLst 创建一个长度约为500k 的长列表。
函数loop2在每次迭代中将输入序列减少1000,并为剩余的列表调用GetByteCount
let GetByteCount data = //can improve this algorithm?
let stopWatch = System.Diagnostics.Stopwatch.StartNew()
let x = data
|> JsonConvert.SerializeObject
|> Encoding.UTF8.GetByteCount
stopWatch.Stop()
Console.WriteLine("time reqd: " + stopWatch.Elapsed.TotalMilliseconds.ToString() + " milliseconds")
x
let PerfTestLoop() =
let rec loop1 ctr l =
if ctr%100 = 0 then Console.WriteLine("" + ctr.ToString())
match ctr with
|500 -> l
|_ ->
let l2 = l @ hrLst
loop1 (ctr+1) l2
let l = loop1 1 hrLst
let len = l.Length
let s = l @ l |> List.toSeq
Console.WriteLine("input length: " + (Seq.length s).ToString())
Console.WriteLine("Start Measuring ..")
let stopWatch = System.Diagnostics.Stopwatch.StartNew()
let rec loop2 ctr s =
match ctr with
|100 -> s |> GetByteCount
|_ ->
let news = Seq.skip 2000 s
let size = news |> GetByteCount
loop2 (ctr+1) news
let x = loop2 1 s
let res = PerfTestLoop() // becomes slow gradually
观察到在loop2 的每次迭代中执行GetByteCount 所花费的时间继续增加,即使序列的大小正在减少!为什么会这样?在任务管理器中,CPU 和内存使用率保持稳定。有没有其他方法可以找到数据的字节数或以某种方式减少执行GetByteCount 所需的时间?
如果在loop2 中删除Seq.skip 行并在每次迭代中使用相同的序列,则每次迭代所需的时间相似并且变化不大。
【问题讨论】:
-
Seq.skip不会像您认为的那样做。它创建了一个新的序列,在枚举时会引起底层序列的枚举,但会跳过前 N 个元素。当您一遍又一遍地执行此操作时,您会创建一个包含Seq.skip调用的“嵌套娃娃”,每个调用都会枚举前一个,最终总是枚举整个原始 500K 序列。 -
哇,谢谢,不知道这个。那么我如何真正“拆分”一个序列?
标签: performance serialization f# json.net