【问题标题】:What am I doing wrong with this AI?我对这个 AI 做错了什么?
【发布时间】:2019-01-07 13:56:48
【问题描述】:

我正在创建一个非常幼稚的 AI(它甚至不应该被称为 AI,因为它只是测试了很多可能性并为他挑选了最好的一个),用于我正在制作的棋盘游戏。这是为了简化平衡游戏所需的手动测试量。

AI 独自玩耍,做以下事情:在每个回合中,AI 与其中一个英雄一起攻击战场上最多 9 个怪物中的一个。他的目标是尽可能快地完成战斗(以最少的回合数)并以最少的怪物激活量。

为了实现这一点,我为 AI 实施了一种超前思考算法,在该算法中,他不会执行当前可能的最佳动作,而是根据其他英雄未来动作的可能结果来选择一个动作。这是他执行此操作的代码 sn-p,它是用 PHP 编写的:

/** Perform think ahead moves
 *
 * @params int         $thinkAheadLeft      (the number of think ahead moves left)
 * @params int         $innerIterator       (the iterator for the move)
 * @params array       $performedMoves      (the moves performed so far)
 * @param  Battlefield $originalBattlefield (the previous state of the Battlefield)
 */
public function performThinkAheadMoves($thinkAheadLeft, $innerIterator, $performedMoves, $originalBattlefield, $tabs) {
    if ($thinkAheadLeft == 0) return $this->quantify($originalBattlefield);

    $nextThinkAhead = $thinkAheadLeft-1;
    $moves = $this->getPossibleHeroMoves($innerIterator, $performedMoves);
    $Hero = $this->getHero($innerIterator);
    $innerIterator++;
    $nextInnerIterator = $innerIterator;
    foreach ($moves as $moveid => $move) {
        $performedUpFar = $performedMoves;
        $performedUpFar[] = $move;
        $attack = $Hero->getAttack($move['attackid']);
        $monsters = array();
        foreach ($move['targets'] as $monsterid) $monsters[] = $originalBattlefield->getMonster($monsterid)->getName();
        if (self::$debug) echo $tabs . "Testing sub move of " . $Hero->Name. ": $moveid of " . count($moves) . "  (Think Ahead: $thinkAheadLeft | InnerIterator: $innerIterator)\n";

        $moves[$moveid]['battlefield']['after']->performMove($move);

        if (!$moves[$moveid]['battlefield']['after']->isBattleFinished()) {
            if ($innerIterator == count($this->Heroes)) {
                $moves[$moveid]['battlefield']['after']->performCleanup();
                $nextInnerIterator = 0;
            }
            $moves[$moveid]['quantify'] = $moves[$moveid]['battlefield']['after']->performThinkAheadMoves($nextThinkAhead, $nextInnerIterator, $performedUpFar, $originalBattlefield, $tabs."\t", $numberOfCombinations);
        } else $moves[$moveid]['quantify'] = $moves[$moveid]['battlefield']['after']->quantify($originalBattlefield);
    }

    usort($moves, function($a, $b) {
        if ($a['quantify'] === $b['quantify']) return 0;
        else return ($a['quantify'] > $b['quantify']) ? -1 : 1;
    });

    return $moves[0]['quantify'];
}

它的作用是递归检查未来的移动,直到达到$thinkAheadleft 值,或者直到找到解决方案(即,所有怪物都被击败)。当它达到它的退出参数时,它会计算战场的状态,与$originalBattlefield(第一次移动之前的战场状态)进行比较。计算方式如下:

 /** Quantify the current state of the battlefield
 *
 * @param Battlefield $originalBattlefield (the original battlefield)
 *
 * returns int (returns an integer with the battlefield quantification)
 */
public function quantify(Battlefield $originalBattlefield) {

    $points = 0;
    foreach ($originalBattlefield->Monsters as $originalMonsterId => $OriginalMonster) {
        $CurrentMonster = $this->getMonster($originalMonsterId);

        $monsterActivated = $CurrentMonster->getActivations() - $OriginalMonster->getActivations();
        $points+=$monsterActivated*($this->quantifications['activations'] + $this->quantifications['activationsPenalty']);

        if ($CurrentMonster->isDead()) $points+=$this->quantifications['monsterKilled']*$CurrentMonster->Priority;
        else {
            $enragePenalty = floor($this->quantifications['activations'] * (($CurrentMonster->Enrage['max'] - $CurrentMonster->Enrage['left'])/$CurrentMonster->Enrage['max']));

            $points+=($OriginalMonster->Health['left'] - $CurrentMonster->Health['left']) * $this->quantifications['health'];
            $points+=(($CurrentMonster->Enrage['max'] - $CurrentMonster->Enrage['left']))*$enragePenalty;
        }
    }

    return $points;
}

当量化一些事物净正点时,一些净负点指向状态。 AI 正在做的是,不是使用当前移动后计算的点来决定采取哪一招,而是使用提前思考部分后计算的点,并根据其他英雄可能的移动来选择移动.

基本上,人工智能正在做的,是说目前攻击怪物 1 不是最好的选择,但如果其他英雄会做这个和这个动作,从长远来看,这个会是最好的结果。

选择一个招式后,AI 对英雄执行一次移动,然后为下一个英雄重复该过程,以 +1 移动计算。

问题:我的问题是,我假设一个“提前思考”3-4 个动作的 AI 应该找到比只执行最佳动作的 AI 更好的解决方案在这一刻。但是我的测试用例显示不同,在某些情况下,一个没有使用“超前思考”选项的人工智能,即目前只下最好的棋步,击败了一个超前思考的人工智能。有时,仅提前 3 步的 AI 会击败提前 4 或 5 步的 AI。为什么会这样?我的假设不正确吗?如果是这样,那是为什么?我是否使用了错误的重量数字?我正在对此进行调查并进行测试,以自动计算要使用的权重,测试可能权重的间隔,并尝试使用最佳结果(即,产生最少转弯次数和/或最少激活次数),但我上面描述的问题仍然存在于这些权重中。

我的脚本的当前版本仅限于 5 步超前思考,与任何更大的超前思考数字一样,脚本变得非常慢(提前 5 步,它会在大约 4 分钟内找到解决方案,但6 提前考虑,它甚至在 6 小时内都没有找到第一个可能的动作)

战斗如何进行: 战斗以下列方式进行:由 AI 控制的多个英雄 (2-4),每个英雄都有许多不同的攻击 (1-x),可以在战斗中使用一次或多次,正在攻击许多怪物(1-9)。根据攻击值,怪物会失去生命值,直到死亡。每次攻击后,被攻击的怪物如果没有死亡就会被激怒,并且每个英雄执行一次动作后,所有怪物都会被激怒。当怪物达到他们的愤怒极限时,它们就会激活。

免责声明:我知道 PHP 不是用于这种操作的语言,但由于这只是一个内部项目,我宁愿牺牲速度,以便能够用我的母语编程语言尽可能快地编写代码。

更新:我们目前使用的量化看起来像这样:

$Battlefield->setQuantification(array(
 'health'                   =>  16,
 'monsterKilled'            =>  86,
 'activations'              =>  -46,
 'activationsPenalty'       =>  -10
));

【问题讨论】:

  • 学习Go(它可能是比PHP更适合此类项目的语言)
  • @BasileStarynkevitch 学习 Go 对这个问题有什么帮助?

标签: artificial-intelligence


【解决方案1】:

如果您的游戏存在随机性,那么任何事情都可能发生。指出这一点,因为从您在此处发布的材料中并不清楚。

如果没有随机性并且演员可以看到游戏的完整状态,那么更长的前瞻绝对应该表现更好。如果没有,则清楚地表明您的评估函数提供了对状态值的错误估计。

在查看您的代码时,您的量化值并未列出,在您的模拟中,您似乎只是让同一个玩家重复移动,而没有考虑其他参与者可能采取的行动。您需要逐步运行完整的模拟以生成准确的未来状态,并且您需要查看不同状态的价值估计,看看您是否同意它们,并相应地调整您的量化。

解决估算价值问题的另一种方法是,以 0.0 到 1.0 范围内的百分比明确预测您赢得回合的机会,然后选择最有可能获胜的动作。计算到目前为止造成的伤害和杀死的怪物数量并不能告诉您为了赢得比赛还剩下多少。

【讨论】:

  • 这个模拟完全没有随机性。我故意省略了通常会使战斗随机化的所有内容,以确保在开发这个小脚本时获得正确的结果。我用量化值更新了主要问题,如果您能对它们发表意见,我将不胜感激。您能否详细说明应该如何从 0.0 缩放到 1.0,以及这将如何帮助创建更好的启发式函数?也许是一个真实的场景示例?提前谢谢你!:)
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-09-17
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多