【发布时间】:2019-07-11 23:46:05
【问题描述】:
为了练习,我在 C++ 中实现了 Conways 的“生命游戏”,通过并行处理对“世界”进行了更新。我正在使用SFML 制作图形。
添加多线程确实使它运行得更快(至少在这台 4 核机器上),但我注意到它有问题。如果我在 Visual Studio 2017 的 Debug 配置中运行它,它似乎开始很慢,但在运行 2 秒后它突然变得更快并且运行顺利。但是,如果我在 Release 配置中运行它,那么它的运行速度甚至比 Debug 还要快,但每隔半秒左右它就会“卡顿”或卡顿,并且不会像我预期的那样顺利运行。
什么可能导致这两个行为问题,我该如何解决?
GameOfLife.cpp:
#include "GameOfLife.h"
#include <iostream>
#include <vector>
#include <math.h>
#include <thread>
#include <mutex>
#include <SFML/Graphics.hpp>
class GameOfLife
{
public:
GameOfLife(int sizeX, int sizeY);
uint8_t & getCell(int x, int y);
sf::Vector2i get2D(int i);
void doUpdate(int start, int end);
virtual ~GameOfLife() = default;
void update();
std::vector<sf::Vector2i> getLiveCells();
private:
std::vector<uint8_t> world;
};
std::mutex updateListLock;
std::vector<sf::Vector2i> pendingUpdates;
sf::Vector2i worldSize;
GameOfLife::GameOfLife(int sizeX, int sizeY)
{
worldSize = sf::Vector2i(sizeX, sizeY);
// initialize world to specified size, all starting as dead
world = std::vector<uint8_t>(sizeX * sizeY, 0);
// reserve space for worst case (every cell needs to be updated)
pendingUpdates.reserve(sizeX * sizeY);
// place a glider
getCell(1, 3) = true;
getCell(2, 4) = true;
getCell(3, 2) = true;
getCell(3, 3) = true;
getCell(3, 4) = true;
// place a glider at top-center
int midX = std::floor(worldSize.x / 2);
getCell(midX + 1, 3) = true;
getCell(midX + 2, 4) = true;
getCell(midX + 3, 2) = true;
getCell(midX + 3, 3) = true;
getCell(midX + 3, 4) = true;
}
uint8_t& GameOfLife::getCell(int x, int y)
{
return world[y * worldSize.x + x];
}
sf::Vector2i GameOfLife::get2D(int index)
{
int y = std::floor(index / worldSize.x);
int x = index % worldSize.x;
return sf::Vector2i(x, y);
}
// Update the cells from position start (inclusive) to position end (exclusive).
void GameOfLife::doUpdate(int start, int end)
{
for (int i = start; i < end; i++)
{
auto pos = get2D(i);
// # of alive neighbors
int aliveCount = 0;
// check all 8 surrounding neighbors
for (int nX = -1; nX <= 1; nX++) // nX = -1, 0, 1
{
for (int nY = -1; nY <= 1; nY++) // nY = -1, 0, 1
{
// make sure to skip the current cell!
if (nX == 0 && nY == 0)
continue;
// wrap around to other side if neighbor would be outside world
int newX = (nX + pos.x + worldSize.x) % worldSize.x;
int newY = (nY + pos.y + worldSize.y) % worldSize.y;
aliveCount += getCell(newX, newY);
}
}
// Evaluate game rules on current cell
switch (world[i]) // is current cell alive?
{
case true:
if (aliveCount < 2 || aliveCount > 3)
{
std::lock_guard<std::mutex> lock(updateListLock);
pendingUpdates.push_back(pos); // this cell will be toggled to dead
}
break;
case false:
if (aliveCount == 3)
{
std::lock_guard<std::mutex> lock(updateListLock);
pendingUpdates.push_back(pos); // this cell will be toggled to alive
}
break;
}
}
}
void GameOfLife::update()
{
unsigned maxThreads = std::thread::hardware_concurrency();
// divide the grid into horizontal slices
int chunkSize = (worldSize.x * worldSize.y) / maxThreads;
// split the work into threads
std::vector<std::thread> threads;
for (int i = 0; i < maxThreads; i++)
{
int start = i * chunkSize;
int end;
if (i == maxThreads - 1) // if this is the last thread, endPos will be set to cover remaining "height"
end = worldSize.x * worldSize.y;
else
end = (i + 1) * chunkSize;
std::thread t([this, start, end] {
this->doUpdate(start, end);
});
threads.push_back(std::move(t));
}
for (std::thread & t : threads) {
if (t.joinable())
t.join();
}
// apply updates to cell states
for each (auto loc in pendingUpdates)
{
// toggle the dead/alive state of every cell with a pending update
getCell(loc.x, loc.y) = !getCell(loc.x, loc.y);
}
// clear updates
pendingUpdates.clear();
}
std::vector<sf::Vector2i> GameOfLife::getLiveCells()
{
std::vector<sf::Vector2i> liveCells;
liveCells.reserve(worldSize.x * worldSize.y); // reserve space for worst case (every cell is alive)
for (int i = 0; i < worldSize.x * worldSize.y; i++) {
auto pos = get2D(i);
if (world[i])
liveCells.push_back(sf::Vector2i(pos.x, pos.y));
}
return liveCells;
}
【问题讨论】:
-
您确定性能问题与您的代码有关,而不是与 Visual Studio 有关吗?如果您打开了 VS 分析器,或者您加载了符号,那么这并不意外。如果直接启动可执行文件会怎样?还。在每个更新调用中,您都会创建一堆线程,这可能会减慢您的程序速度。最好只保留一些线程用于更新——但让它们一直处于活动状态,等待队列或某种数据结构中包含更新请求。还。如果 2+ 个线程正在编辑靠近的单元格,则会因为错误共享而降低性能。
-
现在,每个单元格更改都需要获取一个互斥锁,这可能会导致很多争用。相反,如果每个线程都有自己的更新列表,则根本不需要互斥锁
-
旁白:
for each (auto loc in pendingUpdates)看起来不像 C++。你有类似#define each和#define in :的东西吗?这类事情让有 C++ 经验的人更难查看您的代码 -
@Caleth:没有宏,那是 Visual C++ 主义。实际上,微软首先引入了 ranged-for,使用了这种类似 C# 的语法,然后提出来进行标准化。提交的标准化在采用时将语法更改为
for ( : )。 Microsoft 仍然支持他们的 pre-Standard 语法以避免破坏现有代码。不过,如果cl生成可移植性警告就好了。 -
@KyleV。您可以查看std::thread::detach。请注意,分离的线程将不再可连接,因此它将无限期地存在,除非您有自己的机制来结束它。您可以做的是有一个分离的线程等待主线程将工作放入任务队列(或任何结构)。可以使用std::condition_variable 来实现等待。
标签: c++ multithreading sfml conways-game-of-life