ag

HighWalk AI Benchmark

Genesis of HighWalk

When one of my projects moved into its run phase, I looked into automating updates to its technical specifications. The challenge: maintain documents intended for human teams, at the right level of abstraction, covering architecture and business rules without getting lost in implementation details.

To achieve this, I designed and refined a dedicated Skill. The main challenge lies in assessing the real impact of code changes, then writing and integrating them smoothly into an existing document.

Models that perform very well on coding benchmarks quickly struggled with this task. HighWalk was created to evaluate this workflow: two months of runs on the same task, to analyze which models were able to step back, in broad understanding and documentation, beyond writing code alone.

This first use case let me develop a reproducible testing and scoring methodology, adaptable to the specific context of a task, which I now offer to my clients.

Pilot case

Protocol, criteria, and ranking

Task: Update an existing specification document from 46 commits on a Laravel project (95 modified files, +4,587 / −945 lines).

Environment: Runs performed with OpenCode and OpenRouter, with locked providers (the model publisher's provider when it does not train on prompts, NovitaAI otherwise). Models that support reasoning levels were evaluated on multiple efforts.

Raw quality: Weighted functional evaluation across 11 criteria. It measures the ability to identify relevant changes, abstract them at the right level, and integrate them consistently. Some criteria do not require an addition: they can involve identifying a change as irrelevant. Omissions and unjustified additions are penalized.

Operational efficiency: Duration and cost, weighted equally. The 10-minute and $1.50 caps reflect a practical usage constraint. Exceeding both results in a zero score.

Overall score: 80% raw quality, 20% efficiency. Each configuration was run twice, and the best run was kept.

Screenshot of the HighWalk results table

July 28: first snapshot

At publication, proprietary frontier models hold the top of the ranking. Notable among them are Grok 4.5 (best overall), Claude Opus 5 (best on raw quality), and GPT 5.6 Terra. GPT 5.6 Sol achieves a strong raw-quality score, but its operational efficiency drops to zero. On the open-weight side, only GLM 5.2 and DeepSeek V4 Pro make it into the Top 10.

From this first set of runs, a higher reasoning effort does not seem to guarantee a better result on this task. Operational efficiency also proves fundamental, and it rules out economically unviable configurations from the start.

August 14: Qwen 3.8 Max and effort level

Four models enter the ranking, including Qwen 3.8 Max, which takes first place overall, the first open-weight model to reach that level. Claude Opus 5 remains ahead on raw quality and on 100% criteria coverage. Grok 4.6 stays behind Grok 4.5 on both quality and operational efficiency.

Medium effort starts to look like the right setting for many models. A new iteration of a model also does not necessarily mean a quality gain.

August 23: GLM 5.3, the open-weight shift

Qwen 3.8 Max led the open-weight models overall for only a few days; GLM 5.3 changes that. First overall, first on raw quality, and it covers 100% of the tasks. It's the first model, open-weight or not, to lead all three at once, at a single effort level.

The reversal of the balance of power between open-weight and proprietary models is starting to take shape. Coverage rate also becomes a differentiating factor.

September 22: a significant efficiency gain

Two months after the first run, open-weight models have taken the lead. Among proprietary models, only Grok 4.7 still holds a place at the top of the ranking. GLM 5.3 stays first, but MiMo V2.6 is the real surprise. Its Flash version reaches 90.91% raw quality and 100% coverage at $0.02, the lowest cost ever observed at that quality level.

The verdict is clear: the top of the ranking is now dominated by open-weight models. It took only two months for the trend to reverse. One model stands out for this task and lets us reach a 100× cost reduction at equal quality.

Results and learnings

Complete data

July 28 – September 22, 2026

Incompatible screen.

Please use a device with a screen width larger than 320px.