【问题标题】:Extract words from a string (words are separated with spaces and tabs, possibly multiple)从字符串中提取单词(单词用空格和制表符分隔,可能有多个)
【发布时间】:2023-03-08 18:36:01
【问题描述】:

我正在尝试在 C 中创建一个从文件读取输入的程序,让它成为 Input.inp,其中包含带有用空格和制表符分隔的单词的字符串,可能是多个,然后写入文件 Output.out , 每个单词在一行。例如,输入文件包含

Hi  my name         is Yang

那么输出文件将如下所示

Hi
my
name 
is 
Yang

此外,如果程序到达文件结尾或到达“#”,程序将停止读取。

下面是我的代码。我从文件中获取字符,然后检查它是“#”还是文件结尾。如果不是,那么它将检查字符是空格、制表符还是行尾。如果不是,则该字符将被放入字符串“word”中。现在,如果我们到达空格、制表符或行尾,那么我将打印字符串“word”,将pos 设置回 0 并继续执行此操作。但这不起作用。有人可以解释为什么我的代码会失败并为我提供如何解决这个问题的方向吗?

#include <stdio.h>
#include <string.h>
#include <stdlib.h>

#define maxn 300

int main(){
    FILE *fin, *fout;
    fin = fopen("splitwords.inp", "r");
    fout = fopen("splitwords.txt", "w");
    char buffer[maxn], word[maxn], ch, d;
    int i, pos = 0;

    while((ch = fgetc(fin)) != EOF && ch != '#'){
        while(ch != ' ' && ch != '\t' && ch != '\0'){
            word[pos] = ch;
            pos++;
            if((d = fgetc(fin)) == ' ' || d == '\t' || d == '\0'){
                word[pos] = '\0';
                fputs(word, fout);
                printf("%s", word);
                pos = 0;
            }
        }
        if(ch == ' ' || ch == '\t' || ch == '\0') continue;
    }

    fclose(fin);
    fclose(fout);
}

【问题讨论】:

  • 使用strtok,任务会简单很多
  • 在您的 while 循环中,您使用 d 存储读取的字符,但条件测试 ch,它永远不会改变。
  • ch 应该是int
  • buffer 在您的程序中未使用,您不需要将读取的字符与空字符进行比较,它不存在于有效的文本文件中
  • 为什么不使用scanf() 和"%s" 格式?它会自动处理空间跳过。您所要做的就是寻找#。如果输入是 abcd#efgh 会发生什么——该哈希算作终止输入吗?

标签: c string file-io


【解决方案1】:

关于你的提案的一些评论

正如评论中所说,当您读取一个字符时,使用 int 来保存它而不是 char,您可能会收到来自编译器的警告:表示while((ch = fgetc(fin)) != EOF 上的问题,例如 comparison 由于数据类型范围有限 始终为真,这是因为 EOF 无法保存在 char。所以在你的代码中 ch 和 d 必须是一个 int

检查 fopen 的结果以确保您打开了文件。

最好加上(),避免可能出现的运算符之间的优先级问题,所以替换

while((ch = fgetc(fin)) != EOF && ch != '#')

while(ch != ' ' && ch != '\t' && ch != '\0'){

if((d = fgetc(fin)) == ' ' || d == '\t' || d == '\0'){

if(ch == ' ' || ch == '\t' || ch == '\0')

by(不考虑其他可能的问题)

while(((ch = fgetc(fin)) != EOF) && (ch != '#'))

while((ch != ' ') && (ch != '\t') && (ch != '\0')){

if(((d = fgetc(fin)) == ' ') || (d == '\t') || (d == '\0')){

if((ch == ' ') || (ch == '\t') || (ch == '\0'))

正如在备注中所说,如果您输入这两个时间:

while((ch = fgetc(fin)) != EOF && ch != '#'){
   while(ch != ' ' && ch != '\t' && ch != '\0'){

你永远无法出去,因为 ch 内部没有变化,所以你在 word 中写得越来越多,最后以未定义的行为退出它(通常崩溃)。

您不需要检查空字符的大小写,它不存在于文本文件中。

您错过了管理换行符('\n' 和 '\r')的情况

与问题无关,因为 ch 不变,您永远不会检查读取的单词是否不够长,无法放置在 word 中,您不能认为它在任何情况下都会发生。

在

if((d = fgetc(fin)) == ' ' || d == '\t' || d == '\0'){

您错过了管理换行符的大小写,并且您不必管理空字符的大小写。

线

if(ch == ' ' || ch == '\t' || ch == '\0') continue;

没用,它在 while 块的末尾,所以即使没有它你也要重新循环


用 C 语言创建一个从文件读取输入的程序,让它成为 Input.inp,其中包含带有用空格和制表符分隔的单词的字符串,可能是多个,然后写入文件 Output.out,每个单词都打开一行。

你的程序也太复杂了,你不需要把单词保存在内存中(这还有一个好处是可以管理超过 299 个单词),你的目标是在输出中将每个单词放在单独的行中文件,所以一个简单的解决方案是:

#include <stdio.h>

int main()
{
  FILE *fin, *fout;
  
  if ((fin = fopen("splitwords.inp", "r")) == NULL)
    puts("cannot open splitwords.inp");
  else {
    if ((fout = fopen("splitwords.txt", "w"))  == NULL)
      puts("cannot open splitwords.txt");
    else {
      int word = 0; /* not inside a word */
      int c; /* an int to manage EOF */
      
      while (((c = fgetc(fin)) != EOF) && (c != '#')) {
        if ((c == ' ') || (c == '\t') ||
            (c == '\n') || (c == '\r')) { /* can use isspace() */
          if (word) {
            /* the space finishes a word, add the new line */
            fputc('\n', fout);
            word = 0; /* not in a word now */
          }
        }
        else {
          fputc(c, fout); /* char of word are placed in output file */
          word = 1; /* we are in a word */
        }
      }
      
      if (word) {
        /* we was reading a word, need to add the final newline */
        fputc('\n', fout);
      }
      
      fclose(fout);
    }
    
    fclose(fin);
  }
}

编译和执行:

/tmp % gcc -pedantic -Wextra f.c
/tmp % cat splitwords.inp
Hi  my name         is Yang
/tmp % ./a.out
/tmp % cat splitwords.txt 
Hi
my
name
is
Yang

一些解释和备注:

  • 打开文件后,我检查结果以确保 fopen 成功
  • 当我读取一个 char 时,我不会将它保存在 char 而是一个 int 中,以管理 EOF 的情况
  • 在上面的代码中,我比较了空格和制表符等,以便您轻松查看我所做的事情,但是有一个 lib 函数可以完美地做到这一点:isspace 看看它和其他有用的功能(isalpha isdigit ...)。您可以更改相应的行以添加任何其他字符作为分隔符,例如'-'或标点符号(',' ';')等

上面的代码只是在输出文件中写入了非空格/制表符/换行符,更多的是它只需要检测一个单词的结尾来添加一个换行符,这就是我的变量word的目标 当先前管理的字符不是空格/制表符/换行符时,值为 1,否则为 0

【讨论】:

    【解决方案2】:

    好吧,我对此有很多错误并添加了我的 cmets:

        while(ch != EOF && ch != '#') {
                word[pos] = ch;
                pos++;
                if(ch == ' ' || ch == '\t' || ch == '\0') {
                    word[pos] = '\0';
                    fputs(word, fout);
                    printf("%s\n", word);
                    memset(word, '\0', maxn); //flush word
                    pos = 0;
    
                    while (ch == ' ' || ch == '\t' || ch == '\0') { // handle multiple whitespaces
                        ch = fgetc(fin);
                    }
                } else {
                    ch = fgetc(fin);
                }
        }
    

    这可行,但是:
    1.检查pos &lt; maxn,因为可能是内存故障。
    2. 创建函数bool isWhitespace(char c);,因为使用or的多重使用条件很难看。
    3.检查文件是否正确打开fin != NULL &amp;&amp; fout != NULL

    【讨论】:

      猜你喜欢
      • 2015-01-29
      • 2019-08-10
      • 2019-12-02
      • 1970-01-01
      • 2010-12-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-09-07
      相关资源
      最近更新 更多