【问题标题】:Finding common words in given two strings [closed]在给定的两个字符串中查找常用词[关闭]
【发布时间】:2017-11-27 19:36:23
【问题描述】:

我遇到了一个问题,我们需要在给定的两个字符串中找到常用词。问题描述:

概述:给定两个字符串,找出两个字符串共有的单词。 例如:输入:“一二三”、“二三五”。输出:“二”、“三”。

输入:两个字符串。

OUTPUT:两个给定字符串中的常用词,返回字符串的二维数组。

错误情况:无效输入返回 NULL。

注意:如果没有常用词,返回NULL。

我已经解决了这个问题,并开始在 Visual Studio 2013 中解决它。 它有很多测试用例。

这是我的代码:

int ispresent(char*a, char *b)
{
    int m = strlen(a);
    int n = strlen(b);
    int i, j;
    for ( i = 0; i <= n - m; i++){             
        for ( j = 0; j<m; j++){
            if (a[j] != b[i + j]) break;
        }
        if (j == m)
            return 1;
   }
    return 0;
}
char ** commonWords(char *str1, char *str2)
{
    if (str1 != NULL&&str2 != NULL)
    {
        char **res = (char**)malloc(10 * sizeof(char*));
        for (int i = 0; i < 10; i++)
            res[i] = (char*)malloc(31 * sizeof(char));
        char *a = (char*)malloc(31 * sizeof(char));
        int k = 0, j = 0;
        for (int i = 0; str1[i] != '\0'; i++, j++)
        {
            if (str1[i] != ' ')
                a[j] = str1[i];
            if (str1[i] == ' ' || str1[i] == '\0')
            {
                a[j] = '\0';
                if (ispresent(a, str2))
                    res[k++] = a;
                j = 0;
            }
        }
        return res;
    }
    return NULL;
}

但是我在这里遇到了运行时错误: 程序意外终止

请问有什么解决办法吗?

我有一个建议,它可能是重复的 但是我的代码和那个不一样..我不是在问问题>我是在向我的代码提出建议

【问题讨论】:

  • 使用调试器,这是最好的方法。
  • 即使我一直在尝试,但它停止了
  • 调试器有望在您的代码不正确或显示未正确初始化的数组附近停止。
  • 这个家伙的作业副本:P stackoverflow.com/questions/36043210/…
  • 那个人是我们课程的成员之一@Spikolynn

标签: c string char malloc


【解决方案1】:

最后,我得到了稍加改动的解决方案。并感谢为此做出贡献的人们。 该解决方案已针对我的所有测试用例运行:

解决方案:

 #include<stdio.h>
    #include<string.h>
    #include<stdlib.h>
    #define SIZE 10
    #define WORD_SIZE 31 
    int ispresent(char*a, char *b)
    {
        int m = strlen(a);
        int n = strlen(b);
        int i, j;
        for (i = 0; i <= n - m; i++){
            for (j = 0; j<m; j++){
                if (a[j] != b[i + j]) break;
            }
            if (j == m)
                return 1;
        }
        return 0;
    }
    char ** commonWords(char *str1, char *str2)
    {
        if (str1 != NULL&&str2 != NULL)
        {
            char **res = (char**)malloc(SIZE * sizeof(char*));
            for (int p = 0; p < 3; p++)
                res[p] = (char*)malloc(WORD_SIZE * sizeof(char));
            char *a = (char*)malloc(WORD_SIZE * sizeof(char));
            int k = 0, j = 0;
            for (int i = 0; str1[i] != '\0'; i++)
            {
                if (str1[i] != ' ')
                    a[j++] = str1[i];
                if (str1[i] == ' ' || str1[i] == '\0')
                {
                    a[j++] = '\0';
                    if (ispresent(a, str2)&&a!="\0")
                        res[k++] = strdup(a);
                    j = 0;
                }
            }
            for (int x = 0; x <=k; x++)
                if (res[x] == "\0"||k==0||strlen(res[x])<=1)
                    return NULL;
            return res;
        }
        return NULL;
    }

我在 Visual Studio 中运行的程序的测试用例是

test.spec file:

#include "stdafx.h"
#include "CppUnitTest.h"
#include "../src/commonWords.cpp"

using namespace Microsoft::VisualStudio::CppUnitTestFramework;

namespace spec
{
    TEST_CLASS(commonWordsSpec)
    {
    public:

        bool strcmp(char *str1, char *str2) {
            while (*str1 && *str2) {
                if (*str1 != *str2) {
                    return false;
                }
                str1++;
                str2++;
            }
            return !*str1 && !*str2;
        }

        bool compare(char expected[][31], int count, char **actual) {
            for (int i = 0; i < count; ++i) {
                bool found = false;
                for (int j = 0; j < count; ++j) {
                    if (strcmp(expected[i], actual[j])) {
                        found = true;
                        break;
                    }
                }
                if (!found) {
                    return false;
                }
            }
            return true;
        }

        TEST_METHOD(nullInput)
        {
            Assert::IsNull(commonWords(NULL, NULL), L"Common Words null 
       check failed.", LINE_INFO());
        }

        TEST_METHOD(stringsWithSpaces)
        {
            char *str1 = "       ";
            char *str2 = " who what";
            Assert::IsNull(commonWords(str1, str2), L"No common words check failed.", LINE_INFO());
        }

        TEST_METHOD(noCommonWordsInput)
        {
            char *str1 = "the are all is well";
            char *str2 = " who what";
            Assert::IsNull(commonWords(str1, str2), L"No common words check failed.", LINE_INFO());
        }

        TEST_METHOD(commonWordsInput)
        {
            char *str1 = "the are all is well";
            char *str2 = "is who the";
            char expected[2][31] = { { "the" }, { "is" } };
            char **res = commonWords(str1, str2);
            Assert::IsTrue(compare(expected, 2, res), L"Common Words positive check failed.", LINE_INFO());
        }

    };
}

【讨论】:

    【解决方案2】:

    继续我的评论,一种简单且合乎逻辑的方法就是使用strtok(或您喜欢的任何方法)标记第一个字符串中的单词。为找到的每个单词分配存储空间,并将这些单词存储在一个临时的 char **tmp 变量中。 tokenize 第二个字符串,并将第一个字符串中的每个单词与第二个字符串中的每个标记进行比较。如果它们匹配,则将指向第一个数组中单词的指针添加为results 数组中的指针。 (您无需分配额外的存储空间,因为您的所有char **results 变量将保存指向已在tmp 中分配的内存的指针。(在您完成使用result 之前不要释放tmp)

    既然你知道有多少比较会导致匹配,那么输出result 中的每个单词就很简单了。

    要将所有部分放在一起,您可以执行以下操作:

    #include <stdio.h>
    #include <stdlib.h>
    #include <string.h>
    
    enum { NPT = 10, MAXC = 32 };
    
    int main (int argc, char **argv) {
    
        char *s1 = argc > 1 ? argv[1] : (char[]) { 
                            "My dog has fleas and my cat has none." },
             *s2 = argc > 2 ? argv[2] : (char[]) { 
                            "My frog likes the dog and the rat like the cat. "},
             *delim = " \t\n.,",
             **result = malloc (NPT * sizeof *result),
             **tmp = malloc (NPT * sizeof *tmp);
        int n = 0, r = 0;
    
        if (!result || !tmp) { /*... handle mem error ...*/ }
    
        printf ("finding common words between:\n s1: %s\n s2: %s\n", s1, s2);
    
        /* tokenize all words in s1, save in tmp */
        for (char *p = strtok (s1, delim); p && n < NPT; p = strtok (NULL, delim))
            tmp[n++] = strdup (p);
    
        /* tokenize all words in s2, compare against words in tmp, save result */
        for (char *p = strtok (s2, delim); p && r < n; p = strtok (NULL, delim)) {
            for (int i = 0; i < n; i++)
                if (strcmp (p, tmp[i]) == 0)
                    result[r++] = tmp[i];
        }
    
        if (r) printf ("\ncommon words:\n"); else printf ("\nno common words.\n");
        for (int i = 0; i < r; i++)
            printf ("  %s\n", result[i]);
    
        free (result);      /* free all allocated memory */
        for (int i = 0; i < n; i++)
            free (tmp[i]);
        free (tmp);
    
        return 0;
    }
    

    (注意: strtok 修改作为参数传递的字符串,因此使用复合文字来确保 s1 和 s2 可以修改字符数组。

    使用/输出示例

    $ ./bin/str_common_words
    finding common words between:
     s1: My dog has fleas and my cat has none.
     s2: My frog likes the dog and the rat like the cat.
    
    common words:
      My
      dog
      and
      cat
    
    
    $ ./bin/str_common_words "a quick brown fox jumps over the lazy dog" \
    "a quick greed frog jumps over the lazy bog"
    finding common words between:
     s1: a quick brown fox jumps over the lazy dog
     s2: a quick greed frog jumps over the lazy bog
    
    common words:
      a
      quick
      jumps
      over
      the
      lazy
    

    一个没有常用词的例子。

    $ ./bin/str_common_words "my flea and bat" "your snake & rat"
    finding common words between:
     s1: my flea and bat
     s2: your snake & rat
    
    no common words.
    

    检查一下,如果您有任何其他问题,请告诉我。

    【讨论】:

    • 这很好。实际上我们被赋予了这个任务,他们告诉我们不要创建任何额外的字符串,除了结果字符串,也不要使用字符串内置函数。即使我的大多数朋友都像说的那样解决了由你,所以我尽量不标记作品..感谢您为我的问题付出努力@DavidC.Rankin
    • 但是您以不同的方式做到了。我从您的解决方案中了解了很多事情。谢谢
    • 当然,您可以使用两个指针来遍历字符串,就像使用 strtok 一样简单。 (使用完全相同的逻辑布局)您所做的只是定位空格并获取单词。使用startp 和endp 指针,只需使用endp 向前搜索,直到找到' ',然后使用tmp[i] = malloc (endp - startp); strncpy (tmp[i], startp, endp - 1 - startp); tmp[i][endp - 1 - startp] = 0; 我提供的字符串只是默认值,如果你在命令行上给出任何两个字符串,这就是比较的内容。
    【解决方案3】:

    我认为通过这个练习你应该学习这些东西:

    1. 不要使用幻数,而是使用#define
    2. 如果函数可以返回 NULL,则始终检查指针是否为 NULL
    3. 检查边界,如果输入太长,请确保您的代码行为合理
    4. 不要忘记您分配的空闲内存
    5. 尽可能使用库函数

    这里有一个例子,你可以怎么做。

    #include <stdio.h>
    #include <stdlib.h>
    #include <string.h>
    #define MAX_WORD_LEN    30
    #define MAX_WORDS       10
    
    
    char ** commonWords(char *str1, char *str2)
    {
    if (str1 != NULL && str2 != NULL)
    {
        char a[MAX_WORD_LEN+1];
        char **res = (char**)calloc(MAX_WORDS+1, sizeof(char*));
        if(!res) return NULL;
    
        int k = 0, j = 0;
        for (int i = 0; str1[i] != '\0' && k < MAX_WORDS; i++)
        {
            if (str1[i] != ' ' && j < MAX_WORD_LEN)
                a[j++] = str1[i];
            if (str1[i] == ' ' || str1[i] == '\0') {
                a[j] = '\0';
                if (strstr(str2,a)) {
                    res[k++] = strdup(a);
                }
                j = 0;
            }
        }
        return res;
    }
    return NULL;
    }
    
    int main(void)
    {
    char **cw = commonWords("one two three", "two three five");
    if(cw) {
        for(int i=0; cw[i]; i++) {
            printf("%s ",cw[i]);
            free(cw[i]);
        }
        free(cw);
    }
    printf("\n");
    return 0;
    }
    

    请记住,只有练习才能学习;)

    【讨论】:

    • 感谢您的解决方案。如果某些输入工作正常....如果输入为 str1=" ", str2="adc sbgj aka" 。这将不起作用,并且对于没有常用词的字符串也不起作用,我可以做吧。谢谢你的辛勤工作。
    猜你喜欢
    • 1970-01-01
    • 2018-02-28
    • 2023-03-18
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-10-12
    相关资源
    最近更新 更多