【问题标题】:What is Best way for Replace with More Performance Using C? [closed]使用 C 替换为更高性能的最佳方法是什么? [关闭]
【发布时间】:2017-12-07 05:34:54
【问题描述】:

我想编写一个支持 utf8 的好函数,以实现更好的性能。

我进行了深入研究,发现了以下Replace() 候选人:

char* replace(char* orig, char* rep, char* with)
{
  //33-34
    char* result; // the return string
    char* ins;    // the next insert point
    char* tmp;    // varies
    size_t len_rep;  // length of rep
    size_t len_with; // length of with
    size_t len_front; // distance between rep and end of last rep
    int count;    // number of replacements
    /* char* strstr(char const* s1, char const* s2);
    retourne un pointeur vers la première occurrence de s2 dans s1
    (ou NULL si s2 n’est pas incluse dans s1). */
    if (!orig)
        return NULL;
    if (!rep || !(len_rep = strlen(rep)))
        return NULL;
    if (!(ins = strstr(orig, rep)))
        return NULL;
    if (!with)
        with = "";
    len_with = strlen(with);
    /*  {   initialisation;
            while (condition) {
                Instructions
                mise_à_jour;
        } }*/
    // compte le nombre d'occurences de la chaîne à remplacer
    for (count = 0; (tmp = strstr(ins, rep)); ++count) {
        ins = tmp + len_rep;
    }
    // allocation de mémoire pour la nouvelle chaîne
    tmp = result = malloc(strlen(orig) + (len_with - len_rep) * count + 1);
    if (!result)
        return NULL;
    /* char* strcpy(char* dest, char const* src);
    copie la chaîne src dans la chaîne dest. Retourne dest.
    Attention ! aucune vérification de taille n’est effectuée ! */

    /* char* strncpy(char* dest, char const* src, size_t n);
    copie les n premiers caractères de src dans dest. Retourne dest.
    Attention ! n’ajoute pas le '\0' à la fin si src contient plus de n
    caractères !*/
    // from here on,
    //    tmp points to the end of the result string
    //    ins points to the next occurrence of rep in orig
    //    orig points to the remainder of orig after "end of rep"
    while (count--) { // count évaluée, puis incrémentée
    // donc ici tant que count est > 0
        ins = strstr(orig, rep);
        len_front = ins - orig;
        tmp = strncpy(tmp, orig, len_front) + len_front;
        tmp = strcpy(tmp, with) + len_with;
        orig += len_front + len_rep; // move to next "end of rep"
    }
    strcpy(tmp, orig);
    return result;
}


char* replace8(char *str, char *old,char *new)
{
  //1.2 :||||||||||
  int i, count = 0;
  int newlen = strlen(new);
  int oldlen = strlen(old);
  for (i = 0; str[i]; ++i)
    if (strstr(&str[i], old) == &str[i])
      ++count, i += oldlen - 1;
  char *ret = (char *) calloc(i + 1 + count * (newlen - oldlen), sizeof(char));
  if (!ret) return "";
  i = 0;
  while (*str)
    if (strstr(str, old) == str)
      strcpy(&ret[i], new),
      i += newlen,
      str += oldlen;
    else
      ret[i++] = *str++;
  ret[i] = '\0';
  return ret;
}



char *replace7(char *orig, char *rep, char *with)
{
  //33-34-35
    char *result; // the return string
    char *ins;    // the next insert point
    char *tmp;    // varies
    int len_rep;  // length of rep
    int len_with; // length of with
    int len_front; // distance between rep and end of last rep
    int count;    // number of replacements
    if (!orig)
    {
        return NULL;
    }
    if (!rep)
    {
        rep = "";
    }
    len_rep = strlen(rep);
    if (!with)
    {
        with = "";
    }
    len_with = strlen(with);

    ins = orig;
    for (count = 0; tmp = strstr(ins, rep); ++count)
    {
        ins = tmp + len_rep;
    }
    // first time through the loop, all the variable are set correctly
    // from here on,
    //    tmp points to the end of the result string
    //    ins points to the next occurrence of rep in orig
    //    orig points to the remainder of orig after "end of rep"
    tmp = result = malloc(strlen(orig) + (len_with - len_rep) * count + 1);
    if (!result)
    {
        return NULL;
    }
    while (count--)
    {
        ins = strstr(orig, rep);
        len_front = ins - orig;
        tmp = strncpy(tmp, orig, len_front) + len_front;
        tmp = strcpy(tmp, with) + len_with;
        orig += len_front + len_rep; // move to next "end of rep"
    }
    strcpy(tmp, orig);
    return result;
}



char *replace6(char *st, char *orig, char *repl)
{
  //17-18
  static char buffer[4000];
  char *ch;
  if (!(ch = strstr(st, orig)))
   return st;
  strncpy(buffer, st, ch-st);
  buffer[ch-st] = 0;
  sprintf(buffer+(ch-st), "%s%s", repl, ch+strlen(orig));
  return buffer;
}



char* replace3(char* s, char* term, char* new_term)
{
  //error
    char *nw = NULL, *pos;
    char *cur = s;
    while(pos = strstr(cur, term))
    {
        nw = (char*)realloc(nw, pos - cur + strlen(new_term));
        strncat(nw, cur, pos-cur);
        strcat(nw, new_term);
        cur = pos + strlen(term);
    }
    strcat(nw, cur);
    free(s);
    return nw;
}



char *replace2(char *original,char *pattern,char *replacement)
{
  //34-37
  size_t replen = strlen(replacement);
  size_t patlen = strlen(pattern);
  size_t orilen = strlen(original);
  size_t patcnt = 0;
  char * oriptr;
  char * patloc;
  // find how many times the pattern occurs in the original string
  for (oriptr = original; patloc = strstr(oriptr, pattern); oriptr = patloc + patlen)
  {
    patcnt++;
  }
  {
    // allocate memory for the new string
    size_t retlen = orilen + patcnt * (replen - patlen);
    char * const returned = (char *) malloc( sizeof(char) * (retlen + 1) );
    if (returned != NULL)
    {
      // copy the original string,
      // replacing all the instances of the pattern
      char * retptr = returned;
      for (oriptr = original; patloc = strstr(oriptr, pattern); oriptr = patloc + patlen)
      {
        size_t skplen = patloc - oriptr;
        // copy the section until the occurence of the pattern
        strncpy(retptr, oriptr, skplen);
        retptr += skplen;
        // copy the replacement
        strncpy(retptr, replacement, replen);
        retptr += replen;
      }
      // copy the rest of the string.
      strcpy(retptr, oriptr);
    }
    return returned;
  }
}



char *replace4(char *string, char *oldpiece, char *newpiece)
{
  //20-21
   int str_index, newstr_index, oldpiece_index, end,
   new_len, old_len, cpy_len;
   char *c;
   static char newstring[10000];
   if ((c = (char *) strstr(string, oldpiece)) == NULL)
      return string;
   new_len        = strlen(newpiece);
   old_len        = strlen(oldpiece);
   //str            = strlen(string);
   end            = strlen(string) - old_len;
   //int count;
   //for (count = 0; (strstr(oldpiece, newpiece)); ++count){}
   //newstring = malloc(str + (old_len - new_len) * count + 1);

   oldpiece_index = c - string;
   newstr_index = 0;
   str_index = 0;
   while(str_index <= end && c != NULL)
   {
      /* Copy characters from the left of matched pattern occurence */
      cpy_len = oldpiece_index-str_index;
      strncpy(newstring+newstr_index, string+str_index, cpy_len);
      newstr_index += cpy_len;
      str_index    += cpy_len;
      /* Copy replacement characters instead of matched pattern */
      ///*newstring=realloc(newstring,sizeof(newstring)+new_len+old_len+end+newstr_index);
      strcpy(newstring+newstr_index, newpiece);
      newstr_index += new_len;
      str_index    += old_len;
      /* Check for another pattern match */
      if((c = (char *) strstr(string+str_index, oldpiece)) != NULL)
         oldpiece_index = c - string;
   }
   /* Copy remaining characters from the right of last matched pattern */
   strcpy(newstring+newstr_index,
  string+str_index);
  return newstring;
}



char *replace5(char *orig, char *rep, char *with)
{
  //32-33-35
  char *result;
  char *ins;
  char *tmp;
  int len_rep;
  int len_with;
  int len_front;
  int count;
  if (!orig)
      return NULL;
  if (!rep)
      rep = "";
  len_rep = strlen(rep);
  if (!with)
      with = "";
  len_with = strlen(with);
  ins = orig;
  for (count = 0; (tmp = strstr(ins, rep)); ++count) {
      ins = tmp + len_rep;
  }
  tmp = result = malloc(strlen(orig) + (len_with - len_rep) * count + 1);
  if (!result)
      return NULL;
  while (count--) {
      ins = strstr(orig, rep);
      len_front = ins - orig;
      tmp = strncpy(tmp, orig, len_front) + len_front;
      tmp = strcpy(tmp, with) + len_with;
      orig += len_front + len_rep;
  }
  strcpy(tmp, orig);
  return result;
}

replace4() 和 replace6() 比其他快,但不是 malloc、realloc。


Replace() 性能测试

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <inttypes.h>
#include <assert.h>
#include <stdint.h>
#include <sys/stat.h>
#include <stdarg.h>
char *temp;
int main()
{
  for(int current=1;current<=80000;current++)
  {
    temp=/*replace6*//*replace3*/replace4("1234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890","0","black");
    //printf("%s",temp);
  }
  return 0;
}

  • 如何制作更好的replace() 函数?
  • 我可以将 malloc、realloc 添加到 replace4() 或 replace6() 吗?

【问题讨论】:

  • 我投票结束这个问题,因为代码正在运行,这个问题应该在codereview.stackexchange.com 上。我不确定......也许因为我们谈论性能,这是正确的地方。
  • 如果代码知道字符串长度,它不应该使用strlen()或strcat()。另外:(在大多数情况下)你不应该混合使用字符串索引和指针算法。
  • UTF8 应该与字符串替换功能无关...
  • 您的“UTF-8 支持”要求有什么意义?没有任何函数验证它们提供的字符串以确保它们是 UTF-8。如果输入是有效的 UTF-8,则输出将是有效的 UTF-8。虚假的输入会产生虚假的输出。因此,如果您在替换中将多字节字符的最后一个字节切掉,例如(长度减一),那么您可能会产生非常糟糕的输出,它不是有效的 UTF-8,但那是 GIGO在工作中(垃圾进,垃圾出)。
  • 此外,字符串替换性能似乎很重要的情况,几乎可以肯定使用的是劣质算法,并且可以通过切换到更有效的方法来显着增强在算法级别 .例如,状态机可以在单次数据传递中并行执行多个替换(全部应用于原始输入,而不是一个接一个)。然后是正则表达式(包含在 POSIX.1-2001 和更高版本的 C 库中,并在其他地方作为库提供)以获得更大的操纵能力。

标签: c performance replace str-replace


【解决方案1】:

就像您提出的解决方案都不是您可以编写的用于平衡性能比较的最快解决方案。

首先,关于功能 4 和 6,有几点:

  • 它们不可重入,因为它们使用函数静态缓冲区
  • 显然,它们无法在特定输入大小之上工作
  • 如果您不需要可重入函数,它们会提供最快的内存分配,因为它们在第一次调用后不会花费任何成本
  • 如果您不复制结果,下一次调用将覆盖它,尽管指针将保持有效,可能会导致奇怪的错误
  • 此外,函数 6 似乎只替换了一次出现

函数 3 的性能很可能比其他函数更差,因为它会多次分配内存,但如果缓冲区已经足够大而可能不执行任何操作,则它可能会更快。

对于其余的函数,我注意到它们遵循以下结构模式:

  1. 扫描输入以计算模式出现的次数
  2. 为输出分配足够的内存
  3. 在替换要替换的字符时将输入复制到输出
  4. 作为旁注:他们似乎对输入进行了不同的检查,我没有检查哪些检查所有常见的极端情况,但功能 8 例如根本不做 NULL 检查,如果你相信你的输入,这可能是正确的事情,或者它只是平淡无奇。您不想牺牲正确性。

此时可能更容易分而治之并独立解决这些问题,因为内存分配可能会占调用成本的很大一部分。因此不清楚,如果你能获得任何东西,例如迭代部分字符串,对输入的缓存大小的块进行替换,并在第一个块之后使用启发式方法估计总大小。然而,使用启发式分配可能没问题,但您必须在代价高昂的重新分配/复制与浪费内存之间取得平衡。

关于 1. 如果你真的想知道你需要的确切内存量,我会说你无法避免全长扫描,从个人经验和SO 我会说strstr 通常非常快.我希望这里优化的回报很低。

第 3 部分在这方面类似,因为使用 c std-lib 复制函数可能会导致编译器使用一些手工制作的程序集,这会破坏您可以编写的所有内容。

然而,我认为第二步可以很好地优化:

  • 首先请注意,如果替换不长于要替换的序列,则可以进行就地子字符串替换
  • 其次,调用者传入的缓冲区可能大于其包含的字符串,在这种情况下,输入缓冲区可能已经足够大了
  • 第三,如果您允许调用者传入输出缓冲区,调用者可能有其他方法(例如内存池),可以让他进行更快的分配

这意味着如果你提供一个允许调用者传入缓冲区的调用,你最终可能会更快。

因此,您可以将 1. 实现为一个计算字符串中模式出现次数的函数(它本身很可能非常有用),2. 可以由用户完成,3. 可以是一个函数假设它接收到足够大的输出缓冲区。如果您想方便地进行一次调用,则可以在第四个助手内部使用这些函数,尽管您可能希望至少检查输入的大小并在可能的情况下就地工作。

关于您的测试,您目前只进行一项检查,该检查涵盖了一个非常简单的案例。如果您想以一种有意义的方式测试调用的性能,您至少应该考虑以下几点:

  • 查看影响算法性能的维度如:

    • 输入字符串的大小
    • 要替换的序列大小
    • 要替换的序列的大小
    • 特别检查在跨越 L1/L2 缓存限制和页面大小倍数时性能如何变化
  • 在预热时间之前进行几次迭代(分支预测等)

  • 对随机序列进行操作只需确保它们的维度属性以防止影响缓存行为,因为您操作相同的 100000 次
  • 收集足够的样本以消除异常值的影响(我认为您在 80k 次迭代中就可以了)
  • 最初将不同维度的测量值分开,独立查看它们,然后将它们融合以获得整体图片

最终,如果您真的知道自己在做什么,就知道您可以查看程序集并尝试改进它。此外,如果您可以排除某些情况,因为您知道自己在一般问题域的特殊子空间中工作,通常情况下,您可以通过为您的情况制定特殊解决方案来获得最大收益。

【讨论】:

  • 很好,可以说replace()的示例代码函数吗?
猜你喜欢
  • 2010-09-09
  • 1970-01-01
  • 2019-03-30
  • 1970-01-01
  • 1970-01-01
  • 2015-11-27
  • 1970-01-01
  • 2012-08-09
  • 1970-01-01
相关资源
最近更新 更多