【问题标题】:What is the correct way of reading from a TCP socket in C/C++?在 C/C++ 中从 TCP 套接字读取的正确方法是什么?
【发布时间】:2010-10-14 13:47:53
【问题描述】:

这是我的代码:

// Not all headers are relevant to the code snippet.
#include <stdio.h>
#include <sys/types.h>
#include <sys/socket.h>
#include <netinet/in.h>
#include <netdb.h>
#include <cstdlib>
#include <cstring>
#include <unistd.h>

char *buffer;
stringstream readStream;
bool readData = true;

while (readData)
{
    cout << "Receiving chunk... ";

    // Read a bit at a time, eventually "end" string will be received.
    bzero(buffer, BUFFER_SIZE);
    int readResult = read(socketFileDescriptor, buffer, BUFFER_SIZE);
    if (readResult < 0)
    {
        THROW_VIMRID_EX("Could not read from socket.");
    }

    // Concatenate the received data to the existing data.
    readStream << buffer;

    // Continue reading while end is not found.
    readData = readStream.str().find("end;") == string::npos;

    cout << "Done (length: " << readStream.str().length() << ")" << endl;
}

如您所知,这有点像 C 和 C++。 BUFFER_SIZE 是 256 - 我应该增加大小吗?如果是这样,该怎么办?有关系吗?

我知道如果出于某种原因没有收到“end”,这将是一个无限循环,这很糟糕 - 所以如果你能提出更好的方法,也请这样做。

【问题讨论】:

  • 感谢您的贡献。请注意,我的代码实现了 read() 方法,该方法可以在 sys/socket.h 库中找到,它是“GNU C 库的一部分”,而不是 C++ 库。

标签: c++ c tcp


【解决方案1】:

在不了解您的完整应用程序的情况下,很难说解决问题的最佳方法是什么,但一种常见的技术是使用以固定长度字段开头的标头,该字段表示消息其余部分的长度.

假设您的标头仅包含一个 4 字节整数,表示消息其余部分的长度。然后只需执行以下操作。

// This assumes buffer is at least x bytes long,
// and that the socket is blocking.
void ReadXBytes(int socket, unsigned int x, void* buffer)
{
    int bytesRead = 0;
    int result;
    while (bytesRead < x)
    {
        result = read(socket, buffer + bytesRead, x - bytesRead);
        if (result < 1 )
        {
            // Throw your error.
        }

        bytesRead += result;
    }
}

然后在后面的代码中

unsigned int length = 0;
char* buffer = 0;
// we assume that sizeof(length) will return 4 here.
ReadXBytes(socketFileDescriptor, sizeof(length), (void*)(&length));
buffer = new char[length];
ReadXBytes(socketFileDescriptor, length, (void*)buffer);

// Then process the data as needed.

delete [] buffer;

这做了一些假设:

  • 发送方和接收方的整数大小相同。
  • 发送方和接收方的字节序相同。
  • 您可以控制双方的协议
  • 发送消息时,您可以预先计算长度。

由于通常希望明确知道您通过网络发送的整数的大小,因此将它们定义在头文件中并明确使用它们,例如:

// These typedefs will vary across different platforms
// such as linux, win32, OS/X etc, but the idea
// is that a Int8 is always 8 bits, and a UInt32 is always
// 32 bits regardless of the platform you are on.
// These vary from compiler to compiler, so you have to 
// look them up in the compiler documentation.
typedef char Int8;
typedef short int Int16;
typedef int Int32;

typedef unsigned char UInt8;
typedef unsigned short int UInt16;
typedef unsigned int UInt32;

这会将上面的内容更改为:

UInt32 length = 0;
char* buffer = 0;

ReadXBytes(socketFileDescriptor, sizeof(length), (void*)(&length));
buffer = new char[length];
ReadXBytes(socketFileDescriptor, length, (void*)buffer);

// process

delete [] buffer;

我希望这会有所帮助。

【讨论】:

  • Ori Pessach 的评论是对这个评论的一个很好的补充。
  • 迟到了,但由于你不知道通信对方的字节序,长度可能应该是网络字节顺序,所以在你的例子中:ReadXBytes(socketFileDescriptor, sizeof(length), (void*)(&amp;length)); length=::ntohl(length); buffer = new char[length]; ReadXBytes(socketFileDescriptor, length, (void*)buffer);
【解决方案2】:

几个指针:

你需要处理一个返回值0,它告诉你远程主机关闭了套接字。

对于非阻塞套接字,您还需要检查错误返回值 (-1) 并确保 errno 不是 EINPROGRESS,这是预期的。

您肯定需要更好的错误处理 - 您可能会泄漏 'buffer' 指向的缓冲区。我注意到,你没有在这段代码 sn-p 中分配任何地方。

如果你的 read() 填满了整个缓冲区,那么其他人就你的缓冲区不是一个空终止的 C 字符串提出了一个很好的观点。这确实是一个问题,而且是一个严重的问题。

您的缓冲区有点小,但只要您不尝试读取超过 256 个字节或您为其分配的任何内容,就应该可以使用。

如果您担心在远程主机向您发送格式错误的消息(潜在的拒绝服务攻击)时进入无限循环,那么您应该使用 select() 并在套接字上设置超时以检查可读性,并且仅在数据可用时读取,并在 select() 超时时退出。

这样的事情可能对你有用:

fd_set read_set;
struct timeval timeout;

timeout.tv_sec = 60; // Time out after a minute
timeout.tv_usec = 0;

FD_ZERO(&read_set);
FD_SET(socketFileDescriptor, &read_set);

int r=select(socketFileDescriptor+1, &read_set, NULL, NULL, &timeout);

if( r<0 ) {
    // Handle the error
}

if( r==0 ) {
    // Timeout - handle that. You could try waiting again, close the socket...
}

if( r>0 ) {
    // The socket is ready for reading - call read() on it.
}

根据您期望接收的数据量,您重复扫描整个消息以寻找“结尾”的方式;令牌效率很低。最好使用状态机(状态为 'e'->'n'->'d'->';')来完成,这样您只需查看每个传入字符一次。

说真的,你应该考虑找一个图书馆来为你做这一切。做对了并不容易。

【讨论】:

  • 不是 EINPROGRESS。 EAGAIN 或 EWOULDBLOCK。
【解决方案3】:

如果你真的按照 dirks 的建议创建了缓冲区,那么:

  int readResult = read(socketFileDescriptor, buffer, BUFFER_SIZE);

可能会完全填满缓冲区,可能会覆盖您在提取到字符串流时所依赖的终止零字符。你需要:

  int readResult = read(socketFileDescriptor, buffer, BUFFER_SIZE - 1 );

【讨论】:

    【解决方案4】:

    1) 其他人(尤其是急切地)注意到缓冲区需要分配一些内存空间。对于较小的 N 值(例如,N

    #define BUFFER_SIZE 4096
    char buffer[BUFFER_SIZE]
    

    这让您不必担心在抛出异常时确保delete[] 缓冲区。

    但请记住,堆栈 的大小是有限的(堆也是如此,但堆栈是有限的),所以你不想放太多。

    2) 在 -1 返回代码上,您不应该简单地立即返回(立即抛出异常更加粗略。)如果您的代码不仅仅是简短的家庭作业。例如,如果非阻塞套接字上当前没有可用数据,则 EAGAIN 可能会在 errno 中返回。查看 read(2) 的手册页。

    【讨论】:

    • 好点,将打开的套接字句柄留在周围是不好的;以后会考虑扔的。
    • 其实我没有解决打开的套接字句柄,因为它没有在你发布的sn-p中打开。但我很高兴你想到了 :-)
    【解决方案5】:

    您在哪里为您的buffer 分配内存?调用bzero 的行调用了未定义的行为,因为缓冲区没有指向任何有效的内存区域。

    char *buffer = new char[ BUFFER_SIZE ];
    // do processing
    
    // don't forget to release
    delete[] buffer;
    

    【讨论】:

      【解决方案6】:

      这是我在使用套接字时经常参考的一篇文章..

      THE WORLD OF SELECT()

      它将向您展示如何可靠地使用“select()”,并在底部包含一些其他有用的链接,以获取有关套接字的更多信息。

      【讨论】:

      • 虽然这在理论上可以回答这个问题,it would be preferable 在此处包含答案的基本部分,并提供链接以供参考。
      【解决方案7】:

      只是从上面的几个帖子中添加内容:

      read() -- 至少在我的系统上 -- 返回 ssize_t。这类似于 size_t,但已签名。在我的系统上,它是一个长整数,而不是整数。如果使用 int,您可能会收到编译器警告,具体取决于您的系统、编译器以及您打开了哪些警告。

      【讨论】:

        【解决方案8】:

        对于任何重要的应用程序(即应用程序必须接收和处理不同长度的不同类型的消息),针对您的特定问题的解决方案不一定只是一种编程解决方案 - 它是一种约定,即 I.E.一个协议。

        为了确定您应该将多少字节传递给您的read 调用,您应该建立一个您的应用程序接收的公共前缀或标头。这样,当一个套接字第一次读取可用时,您就可以决定期望什么。

        二进制示例可能如下所示:

        #include <stdint.h>
        #include <stdlib.h>
        #include <stdio.h>
        #include <unistd.h>
        #include <arpa/inet.h>
        
        enum MessageType {
            MESSAGE_FOO,
            MESSAGE_BAR,
        };
        
        struct MessageHeader {
            uint32_t type;
            uint32_t length;
        };
        
        /**
         * Attempts to continue reading a `socket` until `bytes` number
         * of bytes are read. Returns truthy on success, falsy on failure.
         *
         * Similar to @grieve's ReadXBytes.
         */
        int readExpected(int socket, void *destination, size_t bytes)
        {
            /*
            * Can't increment a void pointer, as incrementing
            * is done by the width of the pointed-to type -
            * and void doesn't have a width
            *
            * You can in GCC but it's not very portable
            */
            char *destinationBytes = destination;
            while (bytes) {
                ssize_t readBytes = read(socket, destinationBytes, bytes);
                if (readBytes < 1)
                    return 0;
                destinationBytes += readBytes;
                bytes -= readBytes;
            }
            return 1;
        }
        
        int main(int argc, char **argv)
        {
            int selectedFd;
        
            // use `select` or `poll` to wait on sockets
            // received a message on `selectedFd`, start reading
        
            char *fooMessage;
            struct {
                uint32_t a;
                uint32_t b;
            } barMessage;
        
            struct MessageHeader received;
            if (!readExpected (selectedFd, &received, sizeof(received))) {
                // handle error
            }
            // handle network/host byte order differences maybe
            received.type = ntohl(received.type);
            received.length = ntohl(received.length);
        
            switch (received.type) {
                case MESSAGE_FOO:
                    // "foo" sends an ASCII string or something
                    fooMessage = calloc(received.length + 1, 1);
                    if (readExpected (selectedFd, fooMessage, received.length))
                        puts(fooMessage);
                    free(fooMessage);
                    break;
                case MESSAGE_BAR:
                    // "bar" sends a message of a fixed size
                    if (readExpected (selectedFd, &barMessage, sizeof(barMessage))) {
                        barMessage.a = ntohl(barMessage.a);
                        barMessage.b = ntohl(barMessage.b);
                        printf("a + b = %d\n", barMessage.a + barMessage.b);
                    }
                    break;
                default:
                    puts("Malformed type received");
                    // kick the client out probably
            }
        }
        

        您可能已经看到使用二进制格式的一个缺点 - 对于每个大于您读取的char 的属性,您必须使用ntohl 或ntohs 函数确保其字节顺序正确。

        另一种方法是使用字节编码的消息,例如简单的 ASCII 或 UTF-8 字符串,这完全避免了字节顺序问题,但需要额外的努力来解析和验证。

        C 中网络数据有两个最终考虑因素。

        首先是一些 C 类型没有固定宽度。例如,不起眼的int被定义为处理器的字长,所以32位处理器会产生32位ints,而64位处理器会产生64位ints。好的、可移植的代码应该让网络数据使用固定宽度的类型,就像在stdint.h 中定义的那样。

        第二个是结构填充。具有不同宽度成员的结构将在某些成员之间添加数据以保持内存对齐,从而使结构在程序中使用起来更快,但有时会产生令人困惑的结果。

        #include <stdio.h>
        #include <stdint.h>
        
        int main()
        {
            struct A {
                char a;
                uint32_t b;
            } A;
        
            printf("sizeof(A): %ld\n", sizeof(A));
        }
        

        在这个例子中,它的实际宽度不会是 1 char + 4 uint32_t = 5 字节,而是 8:

        mharrison@mharrison-KATANA:~$ gcc -o padding padding.c
        mharrison@mharrison-KATANA:~$ ./padding 
        sizeof(A): 8
        

        这是因为在 char a 之后添加了 3 个字节,以确保 uint32_t b 与内存对齐。

        因此,如果您write 和struct A,然后尝试在另一侧读取char 和uint32_t,您将得到char a 和一个uint32_t,其中前三个字节是垃圾最后一个字节是你写的实际整数的第一个字节。

        将您的数据格式显式记录为 C 结构类型,或者更好的是,记录它们可能包含的任何填充字节。

        【讨论】:

        • 来吧。我写的是 ReadXBytes,而不是 Duncan。 ;) 严重的是,如果您有多个读取,您的 ReadExpected 将覆盖目标缓冲区的前面,因为您并没有抵消您已经从套接字读取的内容。此外,如果 MessageHeader 结构不是单字节对齐的,您可能读取的内容超出了您的预期。它必须匹配发送方的字节对齐方式。
        • 这很尴尬,对于错误的归属感到抱歉。我发誓我以前写过网络代码,遇到过这个问题并修复过,这次我只是忘记了。一方面,如果您使用自定义二进制消息格式,我会假设您使用相同的库客户端和服务器端进行数据表示,因此 struct padding 之类的内容最终会共享。另一方面,考虑到不同的编译器、平台字节序、与自定义客户端的互操作性……这可能是很多人只使用字节字符串的原因。
        • 我也只记得 C 标准没有为枚举指定标准大小,因此枚举的大小可能会根据它包含的值的数量而改变......或您正在使用哪个编译器。试图专注于答案的重点而不费力地了解很多细节是一项艰巨的任务。
        猜你喜欢
        • 2021-05-04
        • 2013-09-05
        • 2021-02-09
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2015-03-31
        • 2011-08-08
        • 1970-01-01
        相关资源
        最近更新 更多