【问题标题】:How to read caret notations from a binary file in C?如何从 C 中的二进制文件中读取插入符号?
【发布时间】:2015-04-03 15:53:31
【问题描述】:

我需要一次读取一个字节的二进制文件中的字符,并在满足特定条件时将它们连接起来。我在读取空字符时遇到问题,即^@,如插入符号中所示。 snprintf 和 strcpy 都没有帮助我将这个空字符与其他字符连接起来。这很奇怪,因为当我使用

打印这个字符时
printf("%c",char1);

它打印出插入符号中的空字符,即^@。所以我的理解是,即使是 snprintf 也应该成功连接。

谁能告诉我如何实现这样的连接?

谢谢

【问题讨论】:

  • 顺便说一句,printf("%c",char); 无效C。 char 是保留关键字。您介意向我们展示您的实际代码吗?
  • 对不起。我只是举个例子。会改的

标签: c file-io printf concatenation string-concatenation


【解决方案1】:

C 字符串以空字符结尾。如果您的输入数据可以包含空字节,则您不能安全地使用字符串函数。相反,考虑分配一个足够大的缓冲区(或根据需要动态调整大小)并将每个传入字节写入该缓冲区中的正确位置。

【讨论】:

    【解决方案2】:

    由于您不使用原始 ANSI 字符串,因此您不能使用旨在用于原始 ANSI 字符串的函数,因为解释字符串的方式。

    在 C(和 C++)中,字符串通常以 null 结尾,即最后一个字符是 \0(值 0x00)。至少对于字符串操作和输入/输出的标准函数是这样的(比如printf() 或strcpy())。

    例如行

    const char *text = "Hello World";
    

    幕后变成

    const char *text = "Hello World\0";
    

    因此,当您从文件中读取 \0 并将其放入您的字符串时,您基本上会得到一个基本上为空的字符串。

    为了让问题更清楚,只是一个简单的例子:

    // Let's just assume the sequence 0x00, 0x01 is some special encoding
    const char *input = "Hello\0\1World!";
    char output[256];
    
    strcpy(output, input);
    // strncpy() is for string manipulation, as such it will stop once it encounters a null terminator
    
    printf("%s\n", output); // This will print 'Hello'
    
    memcpy(output, input, 14); // 14 is the string length above plus null terminator
    
    printf("%s\n", output); // This will again print 'Hello' (since it stops at the terminator)
    
    printf("%s\n", output + 7); // This will print "World" (you're skipping the terminator using the offset)
    

    以下是我整理的一个简单示例。它不一定展示最佳实践,也可能存在一些错误,但它应该向您展示一些可能的概念,如何处理原始字节数据,尽可能避免使用标准字符串函数。

    #include <stdio.h>
    
    #define WIDTH 16
    
    int main (int argc, char **argv) {
        int offset = 0;
        FILE *fp;
        int byte;
        char buffer[WIDTH] = ""; // This buffer will store the data read, essentially concatenating it
    
        if (argc < 2)
            return 1;
    
        if (fp = fopen(argv[1], "rb")) {
            for(;;) {
                byte = fgetc(fp); // get the next byte
    
                if (byte == EOF) { // did we read over the end of the file?
                    if (offset % WIDTH)
                        printf("%*s %*.*s", 3 * (WIDTH - offset % WIDTH), "", offset % WIDTH, offset % WIDTH, buffer);
                    else
                        printf("\n");
                    return 0;
                }
    
                if (offset % WIDTH == 0) { // should we print the offset?
                    if (offset)
                        printf(" %*.*s", WIDTH, WIDTH, buffer); // print the char representation of the last line
                    printf("\n0x%08x", offset);
                }
    
                // print the hex representation of the current byte
                printf(" %02x", byte);
    
                // add printable characters to our buffer
                if (byte >= ' ')
                    buffer[offset % WIDTH] = byte;
                else
                    buffer[offset % WIDTH] = '.';
    
                // move the offset
                ++offset;
            }
            fclose(fp);
        }
        return 0;
    }
    

    编译后,将任何文件作为第一个参数传递以查看其内容(不应太大以免破坏格式)。

    【讨论】:

    • 是否可以检查从文件中读取的值是否为 '\0'?
    • @user2220232 如果您只想读取单个字符,甚至不要开始尝试读取字符串:char byte = fgetc(yourhandle); 操作系统将为您缓冲文件,因此不应该有任何显着的性能影响。
    • 顺便说一句,你说 printf("%s\n", output);将打印你好。那么为什么在我的输出中显示空字符,即 ^@ 在给出 printf("%c",char1); 时被打印出来。 ?
    • fgetc 不适合我的工作,因为我不仅使用 ASCII,还使用字节值。
    • @user2220232 因为您正在打印实际的字符表示(使用%c 而不是%s)。这正是“将其解释为字符串或原始数据”的全部内容。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-09-07
    • 2020-10-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多