ag

HighWalk Benchmark

Why HighWalk

Over the past few months, I've worked on automating technical specification updates for projects in their run phase. The challenge is to produce documents for human teams (such as IT departments) with the right level of abstraction: covering architecture and business rules without getting lost in implementation details.

To achieve this, I designed and refined a dedicated Agent Skill. The main challenge lies in assessing the real impact of code changes, then writing and integrating them smoothly into an existing document.

I quickly found that models performing very well in coding benchmarks could struggle with this exercise. HighWalk was created to measure this specific ability to step back: a benchmark that evaluates broad understanding and documentation, beyond writing code or specifications for other agents.

Protocol, criteria, and ranking

Task assessed: Update an existing technical specification document from 46 commits in a Laravel project (95 modified files, +4,587 / -945 lines).

Environment: Runs performed with OpenCode and OpenRouter, with locked providers (the model publisher's provider, and NovitaAI for DeepSeek).

Raw quality: Weighted functional evaluation across 11 criteria. It measures the ability to identify relevant changes, abstract them at the right level, and integrate them consistently into the document. Omissions and unjustified additions are penalized.

Operational efficiency: A score weighted equally between run duration and cost. The 10-minute and $1.50 caps reflect practical usage constraints, exceeding both results in a zero score.

Overall score: 80/20 combination of raw quality and operational efficiency. Prioritizes reliable documentation while valuing configurations that are usable day to day.

Repetition: Each configuration (model, reasoning effort, and provider) was run twice. The best run is retained in order to compare the operational potential of the models.

Limitations: This benchmark covers one use case, one project, and one reference document. It is neither a comprehensive statistical measure nor a performance guarantee in other contexts.

Exclusion: Qwen 3.7 Max excluded from the ranking because it critically altered the document.

Ranking 1: Best overall

Overall ranking. Weighting: 80% raw quality / 20% operational efficiency.

Complete data

Ranking 2: Raw quality

This ranking measures an agent's ability to rigorously meet the brief, regardless of time or cost.

Ranking 3: Operational efficiency

A normalized ranking based on run time and dollar cost (subject to the $1.50 and 10-minute caps). Considered without quality, it is not meaningful on its own but reveals two observations.

Ranking 4: Completed tasks

This isolates models that did not completely fail any criteria. They are the strongest allies for human code review: the agent has at least identified and outlined 100% of the required elements.

Observations and takeaways

Updated on July 26, 2026

Incompatible screen.

Please use a device with a screen width larger than 320px.