【问题标题】:Selective iterator选择性迭代器
【发布时间】:2010-06-15 15:42:30
【问题描述】:

仅供参考:没有提升,是的,它有这个,我想要重新发明轮子;)

C++ 中是否存在某种形式的选择性迭代器(可能)?我想要的是像这样分隔字符串:

some:word{or other

变成这样的形式:

some : word { or other

我可以通过两个循环和 find_first_of(":") 和 ("{") 来做到这一点,但这对我来说似乎(非常)低效。我认为也许有一种方法可以创建/定义/编写一个迭代器,该迭代器将使用 for_each 迭代所有这些值。我担心这会让我为 std::string 编写一个成熟的自定义方式过于复杂的迭代器类。

所以我想也许可以这样做:

std::vector<size_t> list;
size_t index = mystring.find(":");
while( index != std::string::npos )
{
    list.push_back(index);
    index = mystring.find(":", list.back());
}
std::for_each(list.begin(), list.end(), addSpaces(mystring));

这对我来说看起来很乱,我很确定存在一种更优雅的方法。但我想不出来。有人有一个好主意吗?谢谢

PS:我没有测试发布的代码,只是快速写下我会尝试的内容

更新:在考虑了你所有的答案之后,我想出了这个,它很合我的胃口:)。这确实假设最后一个字符是换行符或其他字符,否则结尾 {、} 或 : 将不会得到处理。

void tokenize( string &line )
{
    char oneBack = ' ';
    char twoBack = ' ';
    char current = ' ';
    size_t length = line.size();

    for( size_t index = 0; index<length; ++index )
    {
        twoBack = oneBack;
        oneBack = current;
        current = line.at( index );
        if( isSpecial(oneBack) )
        {
            if( !isspace(twoBack) ) // insert before
            {
                line.insert(index-1, " ");
                ++index;
                ++length;
            }
            if( !isspace(current) ) // insert after
            {
                line.insert(index, " ");
                ++index;
                ++length;
            }
        }
    }

一如既往地欢迎评论:)

【问题讨论】:

  • “C++ 中是否存在某种形式的选择性迭代器(可能)?”好吧,据你说,Boost 有这个。也许我很迂腐,但如果你在引用一个可能的例子后立即问某事是否可能,我会认为你的问题有点愚蠢。您可能会发现阅读 Boost 实现的源代码以了解他们是如何做到的很有启发性。即使您想重新发明它,我相信 Boost 也会提供一些关于如何正确操作的提示。
  • “我担心这会让我编写一个成熟的自定义方式过于复杂的迭代器类” ... Boost 有它的迭代器实用程序正是因为编写你自己的迭代器很烦人。

标签: c++ algorithm stl iterator find


【解决方案1】:

使用 std::istream_iterator 相对容易。

你需要做的是定义你自己的类(比如 Term)。然后使用运算符 >> 定义如何从流中读取单个“单词”(术语)。

我不知道你对一个词的确切定义是,所以我使用以下定义:

  • 任何连续的字母数字字符序列都是术语
  • 任何单个非空白字符且不是字母数字都是单词。

试试这个:

#include <string>
#include <sstream>
#include <iostream>
#include <iterator>
#include <algorithm>

class Term
{
    public:

        // This cast operator is not required but makes it easy to use
        // a Term anywhere that a string can normally be used.
        operator std::string const&() const {return value;}

    private:
        // A term is just a string
        // And we friend the operator >> to make sure we can read it.
        friend std::istream& operator>>(std::istream& inStr,Term& dst);
        std::string     value;
};

现在我们要做的就是定义一个运算符>>,它根据规则读取一个单词:

// This function could be a lot neater using some boost regular expressions.
// I just do it manually to show it can be done without boost (as requested)
std::istream& operator>>(std::istream& inStr,Term& dst)
{
   // Note the >> operator drops all proceeding white space.
   // So we get the first non white space
   char first;
   inStr >> first;

   // If the stream is in any bad state the stop processing.
   if (inStr)
   {
       if(std::isalnum(first))
       {
           // Alpha Numeric so read a sequence of characters
           dst.value = first;

           // This is ugly. And needs re-factoring.
           while((first = insStr.get(), inStr) && std::isalnum(first))
           {
               dst.value += first;
           }

           // Take into account the special case of EOF.
           // And bad stream states.
           if (!inStr)
           {
               if (!inStr.eof())
               {
                   // The last letter read was not EOF and and not part of the word
                   // So put it back for use by the next call to read from the stream.
                   inStr.putback(first);
               }
               // We know that we have a word so clear any errors to make sure it
               // is used. Let the next attempt to read a word (term) fail at the outer if.
               inStr.clear();
           }
       }
       else
       {
           // It was not alpha numeric so it is a one character word.
           dst.value   = first;
       }
  }
  return inStr;
}

所以现在我们可以通过使用 istream_iterator 在标准算法中使用它

int main()
{
    std::string         data    = "some:word{or other";
    std::stringstream   dataStream(data);


    std::copy(  // Read the stream one Term at a time.
                std::istream_iterator<Term>(dataStream),
                std::istream_iterator<Term>(),

                // Note the ostream_iterator is using a std::string
                // This works because a Term can be converted into a string.
                std::ostream_iterator<std::string>(std::cout, "\n")
             );

}

输出:

> ./a.exe
some
:
word
{
or
other

【讨论】:

    【解决方案2】:
    std::string const str = "some:word{or other";
    
    std::string result;
    result.reserve(str.size());
    for (std::string::const_iterator it = str.begin(), end = str.end();
         it != end; ++it)
    {
      if (isalnum(*it))
      {
        result.push_back(*it);
      }
      else
      {
        result.push_back(' '); result.push_back(*it); result.push_back(' ');
      }
    }
    

    插入加速版本

    std::string str = "some:word{or other";
    
    for (std::string::iterator it = str.begin(), end = str.end(); it != end; ++it)
    {
      if (!isalnum(*it))
      {
        it = str.insert(it, ' ') + 2;
        it = str.insert(it, ' ');
        end = str.end();
      }
    }
    

    请注意,std::string::insert 在迭代器传递之前插入,并将迭代器返回到新插入的字符。分配很重要,因为缓冲区可能已在另一个内存位置重新分配(迭代器因插入而无效)。另请注意,您不能为整个循环保留end,每次插入时都需要重新计算它。

    【讨论】:

    • 你给了我我认为最清楚的答案。也许不是最好的,但我喜欢它。我的解决方案与您的解决方案在速度方面的表现如何(我使用插入,但您有字符串的副本)。谢谢
    • 你也可以使用 insert 和我的,迭代器非常灵活,只是为了可视化算法,副本通常更容易开始:) 我将编辑以添加插入版本。
    【解决方案3】:

    存在一种更优雅的方法。

    我不知道 BOOST 是如何实现这一点的,但传统方法是将输入字符串逐个字符地输入到 FSM 中,该FSM 会检测标记(单词、符号)的开始和结束位置。

    我可以通过两个循环和 find_first_of(":") 和 ("{") 来做到这一点

    一个带有 std::find_first_of() 的循环就足够了。

    虽然我仍然非常喜欢 FSM 来完成此类解析任务。

    附: Similar question

    【讨论】:

      【解决方案4】:

      怎么样:

      std::string::const_iterator it, end = mystring.end();
      for(it = mystring.begin(); it != end; ++it) {
        if ( !isalnum( *it ))
          list.push_back(it);
      }
      

      这样,您只需遍历字符串一次,并且 ctype.h 中的 isalnum 似乎可以满足您的需求。当然,上面的代码非常简单和不完整,只是提出了一个解决方案。

      【讨论】:

        【解决方案5】:

        您是否要标记输入字符串,ala strtok?

        如果是这样,您可以使用以下标记化功能。它接受一个输入string 和一个分隔符字符串(字符串中的每个字符都是一个可能的分隔符),并返回一个tokens 的向量。每个token 都是一个带有分隔字符串的元组,以及在这种情况下使用的分隔符:

        #include <cstdlib>
        #include <vector>
        #include <string>
        #include <functional>
        #include <iostream>
        #include <algorithm>
        using namespace std;
        
        //  FUNCTION :      stringtok(char const* Raw, string sToks)
        //  PARAMATERS :    Raw     Pointer to NULL-Terminated string containing a string to be tokenized.
        //                  sToks   string of individual token characters -- each character in the string is a token
        //  DESCRIPTION :   Tokenizes a string, much in the same was as strtok does.  The input string is not modified.  The
        //                  function is called once to tokenize a string, and all the tokens are retuned at once.
        //  RETURNS :       Returns a vector of strings.  Each element in the vector is one token.  The token character is
        //                  not included in the string.  The number of elements in the vector is N+1, where N is the number
        //                  of times the Token character is found in the string.  If one token is an empty string (as with the
        //                  string "string1##string3", where the token character is '#'), then that element in the vector
        //                  is an empty string.
        //  NOTES :         
        //
        typedef pair<char,string> token;    // first = delimiter, second = data
        inline vector<token> tokenize(const string& str, const string& delims, bool bCaseSensitive=false)   // tokenizes a string, returns a vector of tokens
        {
            bCaseSensitive;
        
            // prologue
            vector<token> vRet;
            // tokenize input string
            for( string::const_iterator itA = str.begin(), it=itA; it != str.end(); it = find_first_of(++it,str.end(),delims.begin(),delims.end()) )
            {
                // prologue
                // find end of token
                string::const_iterator itEnd = find_first_of(it+1,str.end(),delims.begin(),delims.end());
                // add string to output
                if( it == itA ) vRet.push_back(make_pair(0,string(it,itEnd)));
                else            vRet.push_back(make_pair(*it,string(it+1,itEnd)));
                // epilogue
            }
            // epilogue
            return vRet;
        }
        
        using namespace std;
        
        int main()
        {
            string input = "some:word{or other";
            typedef vector<token> tokens;
            tokens toks = tokenize(input.c_str(), " :{");
            cout << "Input: '" << input << " # Tokens: " << toks.size() << "'\n";
            for( tokens::iterator it = toks.begin(); it != toks.end(); ++it )
            {
                cout << "  Token : '" << it->second << "', Delimiter: '" << it->first << "'\n";
            }
            return 0;
        
        }
        

        【讨论】:

          猜你喜欢
          • 2019-09-12
          • 1970-01-01
          • 1970-01-01
          • 2020-06-05
          • 1970-01-01
          • 2014-08-12
          • 2011-03-24
          • 2010-10-29
          • 2012-02-01
          相关资源
          最近更新 更多