【发布时间】:2020-03-23 06:31:48
【问题描述】:
在尝试使用 fread 在 R 中读取相当大的文件时,我认为读取速度比预期的要慢得多。
该文件约为 60m 行 x 147 列,其中我只选择 27 列,直接在使用 select 的 fread 调用中;在实际文件中仅找到 27 个中的 23 个。 (可能我输入了一些错误的字符串,但我想这不太重要。)
data.table::fread("..\\TOI\\TOI_RAW_APextracted.csv",
verbose = TRUE,
select = cols2Select)
所使用的系统是一个 Azure VM,它具有 16 核 Intel Xeon 和 114 GB RAM,运行 Windows 10。 我也在使用 R 3.5.2、RStudio 1.2.1335 和 data.table 1.12.0
我还应该补充一点,该文件是我已传输到 VM 本地驱动器上的 csv 文件,因此不涉及网络/以太网。我不确定 Azure VM 的工作原理以及它们使用的驱动器,但我认为它相当于 SSD。虚拟机上没有同时运行/处理其他任何内容。
请在fread 的verbose 输出下方找到:
omp_get_max_threads() = 16 omp_get_thread_limit() = 2147483647 DTthreads = 0 RestoreAfterFork = true Input contains no \n. Taking this to be a filename to open [01] Check arguments Using 16 threads (omp_get_max_threads()=16, nth=16) NAstrings = [<<NA>>] None of the NAstrings look like numbers. show progress = 1 0/1 column will be read as integer [02] Opening the file Opening file ..\TOI\TOI_RAW_APextracted.csv File opened, size = 49.00GB (52608776250 bytes). Memory mapped ok [03] Detect and skip BOM [04] Arrange mmap to be \0 terminated \n has been found in the input and different lines can end with different line endings (e.g. mixed \n and \r\n in one file). This is common and ideal. [05] Skipping initial rows if needed Positioned on line 1 starting: <<"POLNO","ProdType","ProductCod>> [06] Detect separator, quoting rule, and ncolumns Detecting sep automatically ... sep=',' with 100 lines of 147 fields using quote rule 0 Detected 147 columns on line 1. This line is either column names or first data row. Line starts as: <<"POLNO","ProdType","ProductCod>> Quote rule picked = 0 fill=false and the most number of columns found is 147 [07] Detect column types, good nrow estimate and whether first row is column names Number of sampling jump points = 100 because (52608776248 bytes from row 1 to eof) / (2 * 85068 jump0size) == 309216 Type codes (jump 000) : A5AA5555A5AA5AAAA57777777555555552222AAAAAA25755555577555757AA5AA5AAAAA5555AAA2A...2222277555 Quote rule 0 Type codes (jump 001) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777555577555777AA5AA5AAAAA7555AAAAA...2222277555 Quote rule 0 Type codes (jump 002) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777775577555777AA5AA5AAAAA7555AAAAA...2222277555 Quote rule 0 Type codes (jump 003) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777775577555777AA5AA5AAAAA7555AAAAA...2222277775 Quote rule 0 Type codes (jump 010) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777775577555777AA5AA5AAAAA7555AAAAA...2222277775 Quote rule 0 Type codes (jump 031) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777775577555777AA7AA5AAAAA7555AAAAA...2222277775 Quote rule 0 Type codes (jump 098) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777775577555777AA7AA5AAAAA7555AAAAA...2222277775 Quote rule 0 Type codes (jump 100) : A5AA5555A5AA5AAAA5777777757777775A5A5AAAAAAA7777775577555777AA7AA5AAAAA7555AAAAA...2222277775 Quote rule 0 'header' determined to be true due to column 2 containing a string on row 1 and a lower type (int32) in the rest of the 10045 sample rows ===== Sampled 10045 rows (handled \n inside quoted fields) at 101 jump points Bytes from first data row on line 2 to the end of last row: 52608774311 Line length: mean=956.51 sd=35.58 min=823 max=1063 Estimated number of rows: 52608774311 /
956.51 = 55000757 Initial alloc = 60500832 rows (55000757 + 9%) using bytes/max(mean-2*sd,min) clamped between [1.1*estn, 2.0*estn]
===== [08] Assign column names [09] Apply user overrides on column types After 0 type and 124 drop user overrides : 05000005A0005AA0A0000770000077000A000A00000000770700000000000000A00A000000000000...0000000000 [10] Allocate memory for the datatable Allocating 23 column slots (147 - 124 dropped) with 60500832 rows [11] Read the data jumps=[0..50176), chunk_size=1048484, total_size=52608774311 |--------------------------------------------------| |==================================================| jumps=[0..50176), chunk_size=1048484, total_size=52608774311 |--------------------------------------------------| |==================================================| Read 54964696 rows x 23 columns from 49.00GB (52608776250 bytes) file in 30:26.810 wall clock time [12] Finalizing the datatable Type counts:
124 : drop '0'
3 : int32 '5'
7 : float64 '7'
13 : string 'A'
=============================
0.000s ( 0%) Memory map 48.996GB file
0.035s ( 0%) sep=',' ncol=147 and header detection
0.001s ( 0%) Column type detection using 10045 sample rows
6.000s ( 0%) Allocation of 60500832 rows x 147 cols (9.466GB) of which 54964696 ( 91%) rows used
1820.775s (100%) Reading 50176 chunks (0 swept) of 1.000MB (each chunk 1095 rows) using 16 threads + 1653.728s ( 91%) Parse to row-major thread buffers (grown 32 times) + 22.774s ( 1%) Transpose +
144.273s ( 8%) Waiting
24.545s ( 1%) Rereading 1 columns due to out-of-sample type exceptions
1826.810s Total Column 2 ("ProdType") bumped from 'int32' to 'string' due to <<"B810">> on row 14
基本上,我想知道这是否正常,或者我是否可以做些什么来提高这些阅读速度。根据我看到的各种基准测试以及我自己对使用较小文件的 fread 的经验和直觉,我希望这会更快地被读取。
我还想知道多核功能是否得到了充分利用,因为我听说在 Windows 下这可能并不总是那么简单。不幸的是,我对这个主题的了解非常有限,但从verbose 输出中确实可以看出fread 正在检测16 个内核。
【问题讨论】:
-
不确定是否有帮助,但也许用
colClasses指定所有列类可以节省时间。还有一个nThread选项 -
Jonny,我相信
nThread默认为getDTthreads应该没问题,但我会尝试覆盖它。我刚刚尝试了几件事,这种行为似乎变得更加陌生。更改要选择的列的名称以便找到所有 27 个列后,速度显着提高到约 180 秒。然而,必须对 2 个列类型进行碰撞(重新分配),所以我厌倦了通过 colClasses 为这 2 个列指定列类来进一步加速它。然而,这显着减慢了读取过程(我没有等到确切地看到) -
我上面提到的加速可能是由于
fread的缓存。我现在重新启动了会话,无法再复制 180 秒的读取速度。所以问题仍然存在...... -
您是在本地运行,还是在服务器上运行?我们已经看到高并行性在繁忙的服务器上可能不是一件好事
-
迈克尔,我不确定是否诚实,但我倾向于根据同事告诉我的内容在服务器上说。我现在已经尝试了这里给出的大部分建议,但我看不到任何实质性的改进。我的原始数据量现在实际上已经增加(大约 80GB,80m 行 x 147 列,其中只有 27 个被选中)
标签: r data.table fread