AI evaluation
GamesForAI: making AI matches easier to inspect
A shared game interface and recorded replays that help developers examine how an AI agent reaches a result.
By Curtis Tech Solutions3 min read
A final score tells you how a game ended. Understanding an AI agent’s behavior takes more: the starting state, the actions it selected, the opponent it faced, and a way to inspect the sequence afterward.
GamesForAI is a Curtis Tech Solutions project built around that broader workflow. It combines Rust game engines, a shared API, built-in opponents, and a React interface for playing, watching, and reviewing matches. The aim is to make game-based AI work easier to connect and inspect.
Put games behind a shared interface
The platform includes chess, Connect Four, tic-tac-toe, and Sudoku. Each has different rules and valid actions, but the surrounding work has much in common: create a match, inspect its state, submit an action, and observe the result.
Rust engines sit behind shared REST, WebSocket, and Model Context Protocol interfaces. That gives applications and agents defined ways to interact with the games while keeping rule enforcement within the platform.
This boundary is useful when experimenting with different agents. A developer can focus on how an agent chooses an action without building a new web interface and match service for every change to the decision-making code.
Make the opponent part of the experiment
GamesForAI provides opponents based on random play, minimax, and Monte Carlo tree search, along with a local Stockfish integration for chess. These give developers different behaviors against which to exercise an agent.
Opponent settings also need an honest interpretation. The project describes levels through resource budgets and leaves ratings unset until they are calibrated. A larger search budget is a configuration choice; it should not silently become a claim about an opponent’s measured rating.
For example, a comparison could hold the starting position and opponent configuration fixed while changing the agent under test. That is an illustrative evaluation setup, not a reported benchmark result. Recording the setup is what makes the result understandable later.
Replay the actions that actually happened
The platform persists moves and request receipts together. Recorded actions can reconstruct a match after a restart, and the replay tools let a developer revisit its history.
A replay follows those recorded actions. It does not ask the opponents to make all their choices again and assume they will produce the same game. That matters when an opponent uses randomness or a search process whose next decision may change.
The history can also be forked to explore another continuation. This creates a practical path from noticing an interesting position to examining what a different next action might do.
Give developers a way to watch
The React and Phaser frontend supports playing, watching live matches, and inspecting replays. The visual interface complements the API: programmatic access runs the interaction, while a person can look at the board and follow the sequence.
GamesForAI demonstrates the value of building the inspection tools alongside the engine. An agent’s decision becomes easier to discuss when the game state and the path to it are available to the developer. That is useful infrastructure for testing ideas, finding mistakes, and deciding what to investigate next.