【问题标题】:Reading CSV with mixed type data使用混合类型数据读取 CSV
【发布时间】:2012-09-06 19:22:05
【问题描述】:

我需要在 MATLAB 中阅读以下csv 文件:

2009-04-29 01:01:42.000;16271.1;16271.1
2009-04-29 02:01:42.000;2.5;16273.6
2009-04-29 03:01:42.000;2.599609;16276.2
2009-04-29 04:01:42.000;2.5;16278.7
...

我想要三列:
时间戳;value1;value2

我尝试了这里描述的方法:
Reading date and time from CSV file in MATLAB
修改为:

filename = 'prova.csv';  
fid = fopen(filename, 'rt');  
a = textscan(fid, '%s %f %f', ...  
        'Delimiter',';', 'CollectOutput',1);  
fclose(fid);

但它返回一个 1x2 单元格,其第一个元素是 a{1}='ÿþ2',其他为空。

我也曾尝试根据我的情况调整这些问题的答案:
importing data with time in MATLAB
Read data files with specific format in matlab and convert date to matal serial time
但我没有成功。

如何导入 csv 文件?

编辑在@macduff 的回答之后,我尝试将上面报告的数据复制粘贴到一个新文件中并使用:

a = textscan(fid, '%s %f %f','Delimiter',';');  

它有效。 不幸的是,这并没有解决问题,因为我必须处理自动生成的 csv 文件,这似乎是导致 MATLAB 奇怪行为的原因。

【问题讨论】:

  • 也许 Matlab 出于某种原因在第一行卡住了?您是否对生成的文件和您使用复制粘贴制作的文件进行了比较?你能以编程方式从 Matlab 复制/粘贴并使其工作吗?

标签: matlab csv import


【解决方案1】:

试试怎么样:

a = textscan(fid, '%s %f %f','Delimiter',';');

对我来说,我得到:

a = 

{4x1 cell}    [4x1 double]    [4x1 double]

因此,a 的每个元素都对应于 csv 文件中的一列。这是你需要的吗?

谢谢!

【讨论】:

  • 感谢您的回答,但我从您的代码中得到:a = {1x1 cell} [0x1 double] [0x1 double]
【解决方案2】:

看来你的做法是正确的。您提供的示例在这里没有问题,我得到了您想要的输出。 1x2 单元格中有什么?

如果我是你,我会再次尝试使用文件的较小子集,比如 10 行,然后查看输出是否更改。如果是,则尝试 100 行等,直到找到 4x1 单元 + 4x2 阵列分解为 1x2 单元的位置。可能是有一个空行或一个空字段或其他任何东西,这会迫使textscan 收集额外级别的单元格中的数据。

请注意,'CollectOutput',1 会将最后两列收集到一个数组中,因此您最终会得到 1 个包含字符串的 4x1 元胞数组和 1 个包含双精度的 4x2 数组。这真的是你想要的吗?否则,请参阅@macduff 的帖子。

【讨论】:

  • 感谢您的回答。单元格的第一个元素是a{1}='ÿþ2' 其他都是空的....我尝试向原始文件添加更多数据(10 行,50,100...),但没有任何变化...。如果我复制粘贴整个数据集(大约 2 万行)在另一个 csv它完美地工作
【解决方案3】:

我不得不解析这样的大文件,但我发现我不喜欢 textscan 来完成这项工作。我只是使用一个基本的 while 循环来解析文件,并使用 datevec 将时间戳组件提取到一个 6 元素时间向量中。

%% Optional: initialize for speed if you have large files
n = 1000  %% <# of rows in file - if known>
timestamp = zeros(n,6);
value1 = zeros(n,1);
value2 = zeros(n,1);

fid = fopen(fname, 'rt');
if fid < 0
    error('Error opening file %s\n', fname); % exit point
end

cntr = 0
while true
    tline = fgetl(fid);  %% get one line
    if ~ischar(tline), break; end; % break out of loop at end of file
    cntr = cntr + 1;

    splitLine = strsplit(tline, ';');  %% split the line on ; delimiters
    timestamp(cntr,:) = datevec(splitLine{1}, 'yyyy-mm-dd HH:MM:SS.FFF'); %% using datevec to parse time gives you a standard timestamp vector
    value1(cntr) = splitLine{2};
    value2(cntr) = splitLine{3};
 end

%% Concatenate at the end if you like
result = [timestamp  value1  value2];

【讨论】:

    猜你喜欢
    • 2012-03-11
    • 2021-08-25
    • 1970-01-01
    • 2014-03-26
    • 1970-01-01
    • 2020-09-17
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多