Knowing When to Stop: IGP-Bench and the Multi-Agent System GALOIS for the Inverse Galois Problem

Shaowu Zhang, Jan Pavel Safrata, Brent Kong, Eric Lu, Sathvik Redrouthu, Caiman Moreno-Earle, Tony Yue YU

Submitted by Shaowu Zhang·Version 1··MathDB perpetual, non-exclusive distribution license

Abstract

Frontier agents can now sustain hours of autonomous work. Greater capability may bring greater progress, but it may also carry an agent farther along a wrong route. Existing long-horizon mathematics benchmarks are difficult to build, saturate quickly, or rely on language models to grade partial progress, making it difficult to measure how much effort an agent spends on an unproductive route. We introduce IGP-Bench, a benchmark of 200 long-horizon inverse Galois problems from computational algebraic geometry, only 39 of which have published explicit solutions. Each problem has intermediate results that are difficult to find but easy to check with deterministic GAP and Magma verifiers. Agents record their partial results and attempts in an artifact graph, allowing us to measure reasoning, instruction following, tool use, resource allocation, and, most importantly, analyze when agents stop or redirect their work. Across five problems, 27 intermediate subproblems, and more than 32,000 possible routes, we show that detailed analysis of long-horizon work reveals valuable insights that success rates alone do not. Claude Opus 5 achieves the highest mean progress score (81.0), followed by Claude Fable 5.1 (62.4), GPT-6 Sol (62.1), and GPT-6 Astra (59.4); newer models can spend substantial effort on unproductive routes. To address this failure mode, we introduce GALOIS, an open-source multi-agent system designed to identify promising routes and stop unproductive work. It comprises a lead agent, a strategy reviewer, specialist agents equipped with 36 general algebraic geometry skills, and a route-rejection agent that checks for possible obstructions in parallel. In reference runs led by Claude Opus 5.5, GALOIS obtains Magma-verified realizations for all five tasks in 10–126 minutes, including polynomials for M23 and M24 that no single agent realizes in our evaluation. These results highlight knowing when to stop as an important ability that current benchmarks rarely measure. We will release the tasks, verifiers, evaluation code, and GALOIS.

View PDF ↗knowing-when-to-stop.pdf · 1.22 MB
43

Version history

v1
Initial publication

Discussion 0