【问题标题】:Read a Comma Delimited Text File into a Struct C将逗号分隔的文本文件读入结构 C
【发布时间】:2021-05-24 12:50:11
【问题描述】:

我有一个逗号分隔的船列表及其规格,我需要将其读入结构。每行包含不同的船及其规格,因此我必须逐行阅读文件。

示例输入文件(我将使用的文件有 20 多行):

pontoon,Crest,Carribean RS 230 SLC,1,Suzuki,115,Blue,26,134595.00,135945.00,1,200,0,250,450,450,0
fishing,Key West,239 FS,1,Mercury,250,Orange,24,86430.00,87630.00,0,0,250,200,500,250,0
sport boat,Tahoe,T16,1,Yamaha,300,Yellow,22,26895.00,27745.00,0,250,0,0,350,250,0

我有一个链表 watercraft_t:

typedef struct watercraft {
    char type[15];     // e.g. pontoon, sport boat, sailboat, fishing, 
                       //      canoe, kayak, jetski, etc.
    char make[20];
    char model[30];
    int propulsion;    // 0 = none; 1 = outBoard; 2 = inBoard; 
    char engine[15];   // Suzuki, Yamaha, etc.
    int hp;             // horse power  
    char color[25];
    int length;        // feet
    double base_price;
    double total_price;
    accessories_t extras;
    struct watercraft *next;
} watercraft_t;

我的主函数打开文件并将其存储在指针中:

FILE * fp = fopen(argv[1], "r"); // Opens file got from command line arg

然后将该文件传递给一个函数,该函数应该准确解析 1 行,然后返回该节点以放置在链表中。

 // Create watercrafts from the info in file
watercraft_t *new_waterCraft( FILE *inFile )
{
    watercraft_t *newNode;

    newNode = (watercraft_t*)malloc(sizeof(watercraft_t));

    fscanf(inFile, "%s %s %s %d %s %d %s %d %lf %lf", newNode->type, newNode->make, newNode->model, &(newNode->propulsion), newNode->engine, &(newNode->hp), newNode->color, &(newNode->length), &(newNode->base_price), &(newNode->total_price));

    return newNode;
}

当调用函数仅打印每艘船的类型时,结果如下:

1. pontoon,Crest,CRS
2. SLC,1,Suzuki,11fishing,Key
3. SLC,1,Suzuki,11fishing,Key
4. SLC,1,Suzuki,11fishing,Key
5. SLC,1,Suzuki,11fishing,Key
6. SLC,1,Suzuki,11fishing,Key
7. SLC,1,Suzuki,11fishing,Key
8. SLC,1,Suzuki,11fishing,Key
9. SLC,1,Suzuki,11fishing,Key
10. SLC,1,Suzuki,11fishing,Key
11. SLC,1,Suzuki,11fishing,Key
12. SLC,1,Suzuki,11fishing,Key
13. SLC,1,Suzuki,11fishing,Key
14. SLC,1,Suzuki,11fishing,Key
15. SLC,1,Suzuki,11fishing,Key
16. SLC,1,Suzuki,11fishing,Key
17. SLC,1,Suzuki,11fishing,Key

我已将问题缩小到如何使用 fscanf 从文件中读取值。

我尝试的第一件事是在所有占位符之间使用 %*c,但在运行之后,我的输出看起来完全一样。接下来我意识到我将无法使用 fscanf,因为文本文件将包含需要读取的空格。

我的下一个想法是使用 fgets,但我认为我也不能使用它,因为我不确定每次必须读取多少个字符。我只需要它在行尾停止读取,同时用逗号分隔值。

我已经寻找了几个小时的答案,但到目前为止似乎没有任何效果。

【问题讨论】:

标签: c struct io scanf text-parsing


【解决方案1】:

当您使用%s 时,文本将被解析,直到找到空格或换行符为止,例如,对于文件的第一行,fscanf 会将"pontoon,Crest,Carribean" 存储在make 中,解析找到空格时停止。

fscanf 说明符必须匹配文件中的行,包括逗号,所以你需要这样的东西:

" %14[^,], %19[^,], %29[^,], %d , %14[^,], %d , %24[^,], %d , %lf , %lf /*...*/"

(注意格式说明符开头的空格,这样可以避免解析之前读取的剩余空白)

格式说明符[^,] 使fscanf 读取直到找到逗号或达到限制大小,它还将解析空格而不是%s,此外,使用%14[^,] 避免了潜在的未定义行为通过缓冲区溢出,因为它将读取限制为14 字符加上匹配缓冲区大小的空终止符15

使用fgets 解析行似乎是个好主意,然后您可以使用sscanf 转换值,它的工作原理类似于fscanf

我建议您验证 *scanf 的返回值,以确保读取了正确数量的字段。

【讨论】:

    【解决方案2】:

    我的下一个想法是使用 fgets,但我认为我也不能使用它,因为我不确定每次必须读取多少个字符。我只需要它在行尾停止读取,同时用逗号分隔值。

    不错的方法。这可以通过这种方式完成:

    int main(void) {
        FILE *fp = fopen("in.txt", "r");
    
        while(1) {
            int length = 0;
            int ch;
            long offset= ftell(fp);
    
            // Calculate length of next line
            while((ch = fgetc(fp)) != '\n' && ch != EOF)
                length++;
            
            // Go back to beginning of line
            fseek(fp, offset, SEEK_SET);
    
            // If EOF, it's the last line, and if the length is zero, we're done
            if(ch == EOF && length == 0)
                break;
        
            // Allocate space for line plus BOTH \n AND \0
            char *buffer = malloc(length+2);
    
            // Read the line
            fgets(buffer, length+2, fp);
    
            // Do something with the buffer, for instance with sscanf
    
            // Cleanup
            free(buffer);
            if(ch == EOF) break;
        }
    }
    

    我省略了所有错误检查以保持代码简短。

    【讨论】:

      【解决方案3】:

      我不推荐使用fscanf 的方法。容易出错且不灵活。

      这里是一个解析 csv 文件的例子:

      int main(int argc, char *argv[])
      {
          #define LINE_MAX 1024
          char line_buf[LINE_MAX];
          char *line = line_buf;
      
          char *delim;
          int sep = ',';
      
          #define FIELD_MAX 128
          char field[FIELD_MAX];
      
          FILE *fp = fopen(argv[1], "r");
      
          //foreach line
          while ((line = fgets(line_buf, LINE_MAX, fp))) {
      
              //iterate over the line
              for (char *line_end = line + strlen(line) - 1; line < line_end;) {
      
                  //search for a separator
                  delim = strchr(line, sep);
      
                  //strchr returns NULL if no separator was found
                  if (delim == NULL) {
                      //we set delim to the end of line,
                      //because we want to process the remaining chars (field)
                      delim = line_end;
                  }
      
                  //the first character of field is at 'line'
                  //the last character of field is at delim - 1
                  //delim points to the separator
      
                  //e.g.
                  size_t len = delim - line;
                  memcpy(field, line, len);
                  field[len] = '\0';
                  printf("%s ", field);
                  //end of example
      
                  //set the position to the next character after the separator
                  line = delim + 1;
      
              }
      
              printf("\n");
      
          }
      
          return EXIT_SUCCESS;
      }
      

      注意:在csv文件末尾添加一个空行,否则不考虑最后一行的最后一个字符(原因:line_end = line + strlen(line) - 1)。

      【讨论】:

        猜你喜欢
        • 2016-08-19
        • 1970-01-01
        • 2013-03-25
        • 1970-01-01
        • 2023-02-02
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2017-06-09
        相关资源
        最近更新 更多