Scale AI's VeRO framework lets one AI agent optimize another, with limits
Scale AI presented VeRO at ICML, a framework in which an optimizer agent improves a target agent by up to 19 points on the GAIA benchmark.
Scale AI researchers unveiled VeRO (Versioning, Rewards and Observations), a framework in which an “optimizer” agent edits and improves a second target agent. They presented the work at the International Conference on Machine Learning (ICML) in Seoul on July 7, 2026.
The result matters because it tests a live idea in AI development: whether agents can automate the tuning work engineers now do by hand. Across 105 optimization runs, the best configurations produced gains as high as 19 points on GAIA, a benchmark of tool-use tasks, and an average lift of roughly 8% to 9% across tool-use benchmarks including GAIA, TAU-Bench Retail and SimpleQA, Scale AI said.
The catch is where the gains stopped. The optimizer showed almost no improvement on reasoning-heavy benchmarks such as GPQA and MATH, according to the paper, “VeRO: A Harness for Agents to Optimize Agents.” One agent can measurably sharpen another’s tool use, in other words, but not teach it to reason better.
The team also found that more than half of the modifications the optimizer made were prompt edits — the least durable kind of change, prone to breaking when the underlying model is upgraded — while structural changes such as new tools or altered control flow held up better. Authors Varun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue and Samuel Marc Denton released the code and a VeRO-Bench benchmark suite publicly.
The findings come from Scale AI’s own team and have not yet been independently reproduced; the gains are benchmark-specific and may not carry to production agents. Whether outside researchers replicate the tool-use lift is the test.
Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
