ag

HighWalk Benchmark

Why HighWalk

Over the past few months, I've worked on automating technical specification updates for projects in their run phase. The challenge is to produce documents for human teams (such as IT departments) with the right level of abstraction: covering architecture and business rules without getting lost in implementation details.

To achieve this, I designed and refined a dedicated Agent Skill. The main challenge lies in assessing the real impact of code changes, then writing and integrating them smoothly into an existing document.

I quickly found that models performing very well in coding benchmarks could struggle with this exercise. HighWalk was created to measure this specific ability to step back: a benchmark that evaluates broad understanding and documentation, beyond writing code or specifications for other agents.

Protocol, criteria, and ranking

Task assessed: Update an existing technical specification document from 46 commits in a Laravel project (95 modified files, +4,587 / -945 lines).

Environment: Runs performed with OpenCode and OpenRouter, with locked providers (the model publisher's provider, and NovitaAI for DeepSeek). Models that support multiple reasoning levels are evaluated at their default level and, when available, at a higher level.

Raw quality: Weighted functional evaluation across 11 criteria. It measures the ability to identify relevant changes, abstract them at the right level, and integrate them consistently into the document. Some criteria do not necessarily require an update: they may involve identifying a change as irrelevant. Omissions and unjustified additions are penalized.

Operational efficiency: Weighted equally between run duration and cost. 10-minute and $1.50 caps reflect practical usage constraints, exceeding both results in a zero score.

Overall score: 80/20 combination of raw quality and operational efficiency. Prioritizes reliable documentation while valuing configurations that are usable day to day.

Repetition: Each configuration (model, reasoning effort, and provider) was run twice. The best run is retained in order to compare the operational potential of the models.

Limitations: This benchmark covers one use case, one project, and one reference document. It is neither a comprehensive statistical measure nor a performance guarantee in other contexts.

Exclusion: Qwen 3.7 Max excluded because it critically altered the document.

Ranking 1: Best overall

Overall ranking. Weighting: 80% raw quality / 20% operational efficiency.

Complete data

Ranking 2: Raw quality

This ranking measures an agent's ability to rigorously meet the brief, regardless of time or cost.

Ranking 3: Operational efficiency

A normalized ranking based on run time and dollar cost (subject to the $1.50 and 10-minute caps). Considered without quality, it is not meaningful on its own but reveals two observations:

Ranking 4: Completed tasks

This isolates models that did not completely fail any criteria. They are the strongest allies for human review: the agent has at least identified and outlined 100% of the elements.

Observations and takeaways

These measurements are based on API cost. With a subscription, some models may become economically more attractive. The final choice also depends on how the service is accessed.

Updated on August 23, 2026

Incompatible screen.

Please use a device with a screen width larger than 320px.