【问题标题】:How to take formatted input from ifstream如何从 ifstream 获取格式化输入
【发布时间】:2011-12-20 12:06:05
【问题描述】:

我有一个包含一组名称的文本文件,格式如下:

"MARY","PATRICIA","LINDA","BARBARA","ELIZABETH"

等等。我想使用 ifstream 打开文件并将名称读入字符串数组(不带引号、逗号)。我设法通过逐个字符检查输入流来做到这一点。有没有更简单的方法来获取这种格式化输入?

编辑: 我听说你可以使用类似的东西 fscanf (f, "\"%[a-zA-Z]\",", str); 在C中,但是ifstream有这样的方法吗?

【问题讨论】:

  • 您可以查看Boost.Spirit.Qi。它可能需要一些努力来学习,但是一旦你掌握了它,任务就应该很简单。
  • fscanf in C 不会像你想象的那样做:标准fscanf 格式字符串中没有类似正则表达式的东西。

标签: c++


【解决方案1】:

该输入应该可以用std::getline 或std::regex_token_iterator 解析(尽管后者是用大炮射击麻雀)。

例子:


正则表达式

又快又脏,但重量级的解决方案(使用 boost 所以大多数编译器都吃这个)

#include <boost/regex.hpp>
#include <iostream>

int main() {
    const std::string s = "\"MARY\",\"PATRICIA\",\"LINDA\",\"BARBARA\",\"ELIZABETH\"";

    boost::regex re("\"(.*?)\"");
    for (boost::sregex_token_iterator it(s.begin(), s.end(), re, 1), end; 
         it != end; ++it)
    {
        std::cout << *it << std::endl;
    }
}

输出:

MARY
PATRICIA
LINDA
BARBARA
ELIZABETH

或者,您可以使用

boost::regex re(",");
for (boost::sregex_token_iterator it(s.begin(), s.end(), re, -1), end; 

让它沿逗号(还要注意 -1)或其他正则表达式分开。


getline

getline 解决方案(允许空格)

#include <sstream>
#include <iostream>

int main() {
    std::stringstream ss;
    ss.str ("\"MARY\",\"PATRICIA\",\"LINDA\",\"BARBARA\",\"ELIZABETH\"");

    std::string curr;
    while (std::getline (ss, curr, ',')) {
        size_t from = 1 + curr.find_first_of ('"'),
               to   =     curr.find_last_of ('"');
        std::cout << curr.substr (from, to-from) << std::endl;
    }
}

输出是一样的。


getline

getline 解决方案(不允许有空格)

循环变得几乎是微不足道的:

    std::string curr;
    while (std::getline (ss, curr, ',')) {
        std::cout << curr.substr (1, curr.length()-2) << std::endl;
    }

自制解决方案

浪费最少的w.r.t。性能(尤其是当您不存储这些字符串,而是使用迭代器或索引时)

#include <iostream>

int main() {
    const std::string str ("\"MARY\",\"PATRICIA\",\"LINDA\",\"BARBARA\",\"ELIZABETH\"");        

    size_t i = 0;
    while (i != std::string::npos) {
        size_t begin  = str.find ('"', i) + 1, // one behind initial '"'
               end    = str.find ('"', begin),
               comma  = str.find (',', end);
        i = comma;

        std::cout << str.substr(begin, end-begin) << std::endl;
    }
}

【讨论】:

  • 你能告诉我如何用getline解析上面的文本吗?
  • 非常感谢!第二个 getline 解决方案简直太酷了。 :)
【解决方案2】:

据我所知,STL 中没有分词器。但是如果你愿意使用 boost,那里有一个非常好的 tokenizer 类。除此之外,逐个字符是您解决它的最佳 C++ 方式(除非您愿意走 C 路线,并在原始 char * 字符串上使用 strtok_t)。

【讨论】:

    【解决方案3】:

    一个简单的标记器应该可以解决问题;不需要像正则表达式这样重量级的东西。 C++ 没有内置的,但它很容易编写。这是我很久以前从互联网上偷来的,我什至不记得是谁写的,所以对公然的抄袭表示歉意:

    #include <vector>
    #include <string>
    
    std::vector<std::string>
    tokenize(const std::string & str, const std::string & delimiters)
    {
      std::vector<std::string> tokens;
    
      // Skip delimiters at beginning.
      std::string::size_type lastPos = str.find_first_not_of(delimiters, 0);
    
      // Find first "non-delimiter".
      std::string::size_type pos     = str.find_first_of(delimiters, lastPos);
    
      while (std::string::npos != pos || std::string::npos != lastPos)
      {
        // Found a token, add it to the vector.
        tokens.push_back(str.substr(lastPos, pos - lastPos));
    
        // Skip delimiters.  Note the "not_of"
        lastPos = str.find_first_not_of(delimiters, pos);
    
        // Find next "non-delimiter"
        pos = str.find_first_of(delimiters, lastPos);
      }
    
      return tokens;
    }
    

    用法:std::vector&lt;std::string&gt; words = tokenize(line, ",");

    【讨论】:

    • 正则表达式很重。就像这个关于输入的分词器一样,它不会去掉引号。
    • @phresnel:我想可以通过说str.substr(lastPos + 1, pos - lastPos -1) 添加引号剥离,但如果引号应该在标记内转义逗号,那么你是对的,更强大的是需要。
    • 从输入中似乎只是引用了名称。下一个问题是空格是否可能。因此,我的帖子中也存在一些混乱。
    • @phresnel:谢谢! This 看起来真的是源头!我记得修改它以按值返回而不是修改引用参数。
    【解决方案4】:

    实际上,因为我感兴趣,所以我想出了如何使用Boost.Spirit.Qi:

    #include <boost/spirit/include/qi.hpp>
    #include <iostream>
    #include <string>
    #include <vector>
    #include <algorithm>
    #include <iterator>
    
    using namespace boost::spirit::qi;
    
    int main() {
      // our test-string
      std::string data("\"MARY\",\"PATRICIA\",\"LINDA\",\"BARBARA\"");
      // this is where we will store the names
      std::vector<std::string> names;
      // parse the string
      phrase_parse(data.begin(), data.end(), 
               ( lexeme['"' >> +(char_ - '"') >> '"'] % ',' ),
               space, names);
      // print what we have parsed
      std::copy(names.begin(), names.end(), 
                std::ostream_iterator<std::string>(std::cout, "\n"));
    }
    

    要检查解析过程中是否发生错误,只需将字符串上的迭代器存储在变量中,然后进行比较。如果它们相等,则匹配整个字符串,否则,begin-iterator 将指向错误站点。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-08-27
      • 1970-01-01
      • 2021-08-27
      • 2010-11-27
      • 1970-01-01
      • 2020-10-13
      相关资源
      最近更新 更多