【发布时间】:2015-09-10 07:51:33
【问题描述】:
简介
我有一个向量 entities 包含 44 百万个名字。我想将它分成 4 个部分并并行处理每个部分。类Freebase包含函数loadData(),用于分割向量并调用函数multiThread进行处理。
-
loadEntities()读取包含名称的文本文件。我没有把实现放在课堂上,因为它不重要 -
loadData()将构造函数中初始化的向量entities拆分为4 部分,并将@987654329@ 的每一部分添加如下:
threads.push_back(thread(&Freebase::multiThread, this, i, i + right, ref(data)));
- multiThread 是我处理文件的函数
-
i和i+right是多线程 for 循环中用于循环实体的索引 -
returnValues是multiThread的子函数,用于调用外部函数。
问题
cout <<"Entity " << entities[i] << endl; 显示以下结果:
- 实体 m.0rzf6wv(正常)
- 实体 m.0rzf70(正常)
- 实体 m.068s4h9 m.0n_k8bz(错误)
- 实体实体 m.068s5_1(错误)
最后的 2 个输出是错误的。输出应该是:
-
Entity name不是entity entity name也不是entity name name
当输入被发送到函数returnValues 时,这会导致分段错误。我该如何解决?
源代码
#ifndef FREEBASE_H
#define FREEBASE_H
class Freebase
{
public:
Freebase(const std::string &, const std::string &, const std::string &, const std::string &);
void loadData();
private:
std::string _serverURL;
std::string _entities;
std::string _xmlFile;
void multiThread(int,int, std::vector<std::pair<std::string, std::string>> &);
//private data members
std::vector<std::string> entities;
};
#endif
#include "Freebase.h"
#include "queries/SparqlQuery.h"
Freebase::Freebase(const string & url, const string & e, const string & xmlFile, const string & tfidfDatabase):_serverURL(url), _entities(e), _xmlFile(xmlFile), _tfidfDatabase(tfidfDatabase)
{
entities = loadEntities();
}
void Freebase::multiThread(int start, int end, vector<pair<string,string>> & data)
{
string basekb = "PREFIX basekb:<http://rdf.basekb.com/ns/> ";
for(int i = start; i < end; i++)
{
cout <<"Entity " << entities[i] << endl;
vector<pair<string, string>> description = returnValues(basekb + "select ?description where {"+ entities[i] +" basekb:common.topic.description ?description. FILTER (lang(?description) = 'en') }");
string desc = "";
for(auto &d: description)
{
desc += d.first + " ";
}
data.push_back(make_pair(entities[i], desc));
}
}
void Freebase::loadData()
{
vector<pair<string, string>> data;
vector<thread> threads;
int Size = entities.size();
//split database into 4 parts
int p = 4;
int right = round((double)Size / (double)p);
int left = Size % p;
float totalduration = 0;
vector<pair<int, int>> coordinates;
int counter = 0;
for(int i = 0; i < Size; i += right)
{
if(i < Size - right)
{
threads.push_back(thread(&Freebase::multiThread, this, i, i + right, ref(data)));
}
else
{
threads.push_back(thread(&Freebase::multiThread, this, i, Size, ref(data)));
}
}//end outer for
for(auto &t : threads)
{
t.join();
}
}
vector<pair<string, string>> Freebase::returnValues(const string & query)
{
vector<pair<string, string>> data;
SparqlQuery sparql(query, _serverURL);
string result = sparql.retrieveInformations();
istringstream str(result);
string line;
//skip first line
getline(str,line);
while(getline(str, line))
{
vector<string> values;
line.erase(remove( line.begin(), line.end(), '\"' ), line.end());
boost::split(values, line, boost::is_any_of("\t"));
if(values.size() == 2)
{
pair<string,string> fact = make_pair(values[0], values[1]);
data.push_back(fact);
}
else
{
data.push_back(make_pair(line, ""));
}
}
return data;
}//end function
【问题讨论】:
-
我会检查这个链接@ArnonZilca。我仍然想知道我是否将错误的输入传递给 returnValues
-
那么我认为你最好创建 4 个结果向量并在加入后合并它们,或者将 safely (使用互斥体)写入每个线程的结果向量 -我猜对于少量线程使用多个向量会更快。
-
确实如此。元素越多,您支付的互斥量开销就越多。阅读更多关于线程安全和向量的内容后,我发现只要您写入不同的索引,就可以从多个线程安全地写入同一个向量(检查this out)。该解决方案将为您节省互斥锁(解决方案 1)和最终合并的向量(解决方案 2)。
-
你调用
push_back无论如何都会添加元素 - 如果你更新了它们,我认为你应该没问题。 -
但是更新元素需要您提前知道有多少元素确切,总是。
标签: c++ multithreading pass-by-reference