projectrtsunitygame developmentdevlogaibotsneural networksreinforcement learninggame designdifficulty

Dev Diary #5: The Bots

October 5, 2026
5 min read

When I held the first tournament for the bots, it turned out that the three difficulty levels barely differed in strength, and the "easy" bot attacked in waves one and a half times larger than the "hard" one. I hadn't noticed that myself: the levels simply cancelled each other out. This part is about the bots: how they evolved from scripts to neural networks, and how the game now picks an opponent to match the player.

Dev Diary #5: The Bots — 1

The chart shows the strength of the bots on the Elo scale, where the normal bot equals 1500. At the top are the three old levels, which turned out to be almost next to each other, and at the bottom the twelve steps I built afterwards.

The First Bots

The first bot was simple: following the profile of its side, it built a base, saved up an army and sent it at the enemy in waves. There were three difficulty levels, and they differed by a couple of settings: the size of a wave, the number of harvesters and how often the bot makes decisions. That was enough for testing the game, but such a bot is predictable in battle. I needed an opponent that remembers what it has seen and makes decisions instead of playing a script.

A Commander With Memory

So I built a commander designed like a player. It sees only what the player can see, and the fog of war applies to it too. Everything it has noticed it remembers: where the enemy buildings stand, how many units were there and how sure it is about that. Over time the certainty drops, and if the old spot is visible and empty, it is marked as checked. It splits its army into groups, and every group has a plan: a mission, a target, a reason and a deadline. Groups aren't reshuffled at every step, so the army behaves sensibly. Between matches it also remembers the opponent, to take that into account in the next game.

Dev Diary #5: The Bots — 2

To understand what the bot is doing I made a "thought process" panel. The screenshot shows what it currently sees and remembers, what it takes into account, which orders it gave and what its group plans are. When something goes wrong, the panel shows not only what the bot did, but why.

Dev Diary #5: The Bots — 3

And this is a battle between two bots, recorded and watched as a replay with the panel turned on. On the left are the commander's thoughts, and on the map it is storming the enemy base. One honest remark: the tactical part of the bot sees only what the player can see, but its economy still relies on open information about the map and its own base, for example where the resources are. I won't call the bot fully fair without reservations yet.

Dev Diary #5: The Bots — 4

The close-up shows lines and rings: that is the target of one of the groups, where it is heading.

Neural Networks

The next step is neural networks. The commander has four small networks: the strategist picks a mission and a target for a group, the tactician gives orders to a group depending on the situation, the gunner decides who shoots whom, and the economist decides what to build and produce. The networks are tiny, from seven to three hundred thirty thousand parameters, and they run on an ordinary processor. They learn in two steps. First a network copies the decisions of the commander that plays by rules. Then it plays against itself and gets a reward for a won game. For training I run the game without graphics, and a twenty-minute match is calculated in a few seconds.

At first the network didn't improve at all, and for a long time I thought the problem was in the training. It turned out that mostly the rules the commander itself played by were to blame. In 70 to 80 percent of cases the games between bots ended in a draw because the armies froze. All groups "joined an attack" on the place of an already destroyed building: each helped another, while both had an empty target. The army didn't look for what it hadn't seen, and a building constructed after the scout had left was never found. And a plane standing at the airfield pulled the center of a group into the corner of the base, and every eight seconds the tactician commanded "gather at the center": the army walked back and forth on its own base for ten minutes. After the fixes about half of the games were decided instead of twenty to thirty percent.

In the end the best network won 31 games against the bot on rules and lost 3, and against the classic bot it won 44 and lost one, the rest were draws. It is interesting what it learned: the economist started to order builders more often, their share of decisions grew from 6.5 to 15 percent, and its buildings go up faster.

The Tournament and the Difficulty Ladder

When there were many settings and networks, I needed to understand how strong the bots really are. The first tournament showed that the three old levels barely differ: "easy" about 1360, "normal" 1500, "hard" 1572. The reason was found quickly: "easy" attacked in waves one and a half times larger than normal, while "hard" only at three quarters of normal, and they cancelled each other out.

Then I measured how each setting affects the strength of a bot: I changed them one at a time and played 36 games against an unchanged bot. The strongest effect comes from the number of resource harvesters and the size of an attack wave. With half as many harvesters the bot is weaker by 263 points, and with waves one and a half to two times larger it is stronger by 280. And superweapons, upgrades and the army cap change almost nothing, within the noise.

Dev Diary #5: The Bots — 5

So I built the ladder on three levers: the share of harvesters, the size of a wave and how often the commander looks at the field. That gave twelve steps, and I measured the strength of each with a tournament of 2304 games: from 846 to 2071 on the Elo scale. A test checks that the strength grows strictly from step to step.

A Bot That Matches You

Now the most interesting part. The game estimates the player's strength from their matches against bots and picks a bot of the same strength. The first ten matches are for calibration: from the player's income and production pace the game estimates their level, and then refines it by the results. After that the "Match me" difficulty adjusts by the last five matches: it grows when you win and drops when you lose. The idea is that the bot wins about half of the games. In a room with several humans the server takes the average strength of the humans. The bot doesn't see through the fog and doesn't get free money: it is simply weakened or strengthened with the same three levers.

What's Next

Next I want to build the third version of the neural commander: so the networks learn on the graphics card, play both one on one and on maps for six to eight players with the same network, adapt to the player and, when needed, play stronger than them. For now it's only a plan, and none of it is in the game. In the next part I'll tell how I balance the three sides with the help of these same bots.