CodeariaAcademy
Article cover: an exam sheet under a lamp, part of the page covered, with the words Hidden Exam
October 5, 202613 min readtested with Anthropic's post of September 28, 2026 and the Claude Code changelog 2.1.284–2.1.285; example numbers come from Anthropic's internal benchmarks, we did not run the commands ourselvesClaude CodeAI Agents

98.9% on the exam, 90.5% on questions the model never saw. Which number should you trust for your AI feature?

Imagine Claude building an exam for your support bot, then improving the prompt one change per round while hiding part of the questions. How it works.

In numbers

accuracy on 14 tickets the search never saw
of the original setup's cost per ticket
hillclimb rounds on the claude-api skill: 66.1% → 87.9%
In this article8
In short

/claude-api build-eval and /claude-api hillclimb are two sub-commands of the claude-api skill that ships inside Claude Code (Anthropic post, September 28, 2026). build-eval builds an evaluation for a Claude API feature inside your repository: examples from production, tickets or your own hand, a grader (code or a second judge model), and a baseline run with a confidence interval. hillclimb improves the feature against that eval one change per round and keeps part of the examples hidden to catch overfitting. In Anthropic's example a support bot reached 98.9% on the 30 tickets used for the search, but the honest number is a different one: 90.5% on 14 held-out tickets against the original 78.6%, at about one fifth of the cost. To start: claude update, then the commands in Claude Code. You pay only for the tokens of the eval runs.

Everyone who has put a model into a product has had this moment. You change the prompt, run your five favourite examples, it looks better. You ship. A week later support brings a ticket that used to be handled correctly and now is not.

Nobody measured "looks better". The prompt passes the five examples you know by heart anyway: it was written around them.

On September 28 Anthropic published a write-up on how to build that kind of exam and improve a feature against it without fooling yourself. It also added two sub-commands to the claude-api skill that do the work for you: build-eval and hillclimb. The most useful part of the post is an example with two numbers that retellings have already mixed up.

Who this is for

An eval (a set of checks) makes sense wherever a model call has a correct answer, or at least a clear quality bar. For example:

  • a support bot that decides whether to refund a customer;
  • a classifier that routes emails or requests into queues;
  • a prompt that drafts a reply to a customer, where you can say what a good reply must contain.

The Anthropic post itself walks through an email router. If your feature looks like anything on this list, read on.

build-eval: the exam is built from your own data

/claude-api build-eval starts with an interview: Claude asks you about the feature, builds the eval inside your codebase and pauses for your approval at specific points.

  1. 1

    Examples

    In this order of priority: production transcripts (Claude first asks about retention and sensitive data), then bug reports and support tickets, then 5–10 cases you write by hand, and only then cases synthesized from your codebase. Synthetic data is fine, but it is anchored in a few real examples you provide.

  2. 2

    Input review page

    The skill generates a simple page that shows every input and waits until you confirm them.

  3. 3

    Grader

    Claude proposes the cheapest grader that fits your output format. More on that below.

  4. 4

    Checking the grader

    Claude grades a handful of cases and asks whether you would have scored any of them differently.

  5. 5

    Size and baseline

    The skill tells you the size of the run (cases × repeats × model, and roughly how long it takes), runs your current version and prints the score with a confidence interval.

  6. 6

    Diagnostics

    During the baseline the grader runs twice on the same output, and timeouts, API errors and cut-off answers are checked. If the baseline already scores about 95% or higher, the skill warns you there is no room left to climb on quality.

When the baseline already sits around 95%, chasing quality makes little sense: any change drowns in noise. In that case the skill suggests aiming at cost or latency instead.

Which grader Claude picks

If the possible outputs are few and constrained, the check is code: exact match, a label from a fixed set, JSON that matches a schema, tests that pass. It is cheap and does not disagree with itself.

If there are many valid answers but clear quality criteria, it uses LLM-as-judge: a second model reads the input, the output and a rubric. The rubric is written as checkable claims, not a 1-to-5 scale. If you have a baseline to compare against, the judge reads both answers in random order, without being told which is the baseline, and picks the better one. You pick the judge model, and it should not be the model under test. We saw the same "whoever built it does not sign it off" principle in Cloudflare's security audit skill.

After build-eval your repository holds the cases, the grader, the runner, one JSON line and one full transcript per case, and a report.html page that lists each case's score with a link to its transcript. If you want a chart, ask, and Claude will build it as an extra page next to it. Those pages are static by default: they open locally and load nothing from the network.

hillclimb: one change per round

/claude-api hillclimb takes an existing eval and improves your feature against it. You decide what it may change: the system prompt, skills and instruction files, tool descriptions, model, effort and other API parameters, harness code. The goal is yours too: quality, or cost while quality holds.

  1. 1

    Train and test split

    The set is split at random into two parts. Claude searches for improvements on train and never sees test.

  2. 2

    Noise check

    Before the first round Claude checks that the eval's noise is smaller than the smallest improvement you would act on. If it is not, it says so and suggests more repeats or cases.

  3. 3

    Round

    Claude reads the previous round's train transcripts and proposes one change as a patch. It aims at the root of a failure, rewriting a section or adding a missing rule, rather than rewording a line.

  4. 4

    Decision

    Train and test both improve: the patch stays. Train improves while test stays flat: suspected overfitting, revert. A regression: revert.

  5. 5

    Stall

    If the score stalls for two or three rounds, Claude makes no edit and sorts every remaining train failure by cause. That surfaces ambiguous cases, harness errors and run-to-run variance. Only legitimate failures go into further rounds.

  6. 6

    Finish

    Your code is left at the version that did best on test for your goal. The report compares test against the baseline with confidence intervals. If the gain is within noise, Claude says so and recommends against merging.

Besides the train and test split, the skill has two more guards against overfitting. Failure content is never pasted into the prompt, otherwise the model learns the answers to the exam. And the answers are kept out of the model's reach, because models sometimes find them directly.

Anthropic's example: which number is real

Anthropic ran hillclimb on an internal customer support benchmark with the goal of cutting cost and improving quality. The set had 44 tickets: 30 for the search, 14 held out. The start: Opus 4.8 at default (high) effort, 74.4% decision accuracy on the search tickets and 4.6 cents per ticket.

The path went like this. First a prompt audit, which removed mandatory tool-call rituals, a scratchpad step and contradictory rules. Then Opus 5.5 on low effort: 87.8% at 1.9 cents. Then one tier down, Sonnet 5 on low effort: 88.9% at about a cent. And finally routing rules and a refund-cap cross-reference in the prompt brought Sonnet 5 to 98.9% at about the same cost.

This is the number retellings get wrong: 98.9% is the score on the 30 tickets used for the search, the ones hillclimb read round after round.

30 search tickets (train)14 held-out tickets (test)
Original setup: Opus 4.8, high effort74.4%78.6%
Final: Sonnet 5, low effort, improved prompt98.9%90.5%
Cost per ticket4.6 cents → about 1 centabout one fifth of the original

The honest result is in the second column. On tickets the search never saw, accuracy rose from 78.6% to 90.5% while cost fell to about a fifth. That is a good result. But the gap between 98.9% and 90.5% is exactly the overfitting the test set exists to catch. Quoting 98.9% as the outcome repeats the very mistake the skill guards against.

Part of the saving came from pricing: on Opus 5.5 input and output tokens cost 20% less than on Opus 4.8, and cache reads cost 60% less. Why the cost of a task is not the cost of a token, we covered in our GPT-6 Sol vs Opus 5.5 comparison.

A second example: a skill that improved itself

Anthropic also ran hillclimb on the claude-api skill itself. The eval was built from their documentation and checks whether the skill writes correct code for their API. Over 24 rounds the score rose from 66.1% to 87.9%, and the findings along the way are more interesting than the number.

First Claude found the skill was missing coverage of eight features. Sections on them gave 74%, and fixing the C# and Java type tables gave 77%. Then the score stalled for two rounds, and sorting failures by cause showed that the content was there but Claude was writing older API shapes from memory. A table near the top of the skill now maps "how you remember it → how it is now", for example from extended thinking with a fixed budget to adaptive thinking. That gave 80%.

The third finding is the most useful one for you. Tasks that never improved turned out to be broken. One asked for code that catches one error type, while its grader wanted a chain of at least three. Another grader's instructions contradicted the docs, and testing the real API showed the docs were right. With those fixed, plus more skill edits, the score reached about 88%.

Where it will not work

Anthropic lists the limits itself, and they are worth reading before you start.

  • An eval is not production. The post's example: an eval benefits from OCR, hillclimb adds an OCR tool to the harness, the score goes up, and in production OCR is rarely useful.
  • You need headroom. The strongest model at the highest effort should score well below 100%, otherwise there is nothing to compare changes against.
  • Do not pick cases because today's model fails them. Then the eval measures one model's weak spots rather than what is hard for your task. User traffic can skew easy too: people try what they expect to work.
  • Noise hides in the environment. A file or git history left from an earlier run can hand the agent the answer.
  • An open-ended "improve the harness" stalls. Hillclimb works best where changes are cheap and the score can be attributed to the edit: prompts, skills, descriptions. What the harness around a model costs in the first place, the HarnessTax study showed.
  • Both examples use Anthropic's internal benchmarks. On your product the numbers will be different.

How to run it yourself

The claude-api skill ships inside Claude Code, so there is nothing separate to install. Update and call the sub-commands:

update first, then the commands in Claude Code
$ claude update
> /claude-api build-eval
> /claude-api hillclimb

You can steer build-eval by giving it access to examples, such as traces. Run hillclimb once an eval exists, and name the goal up front: quality, or cost while quality holds.

The commands have no price of their own. You pay for the tokens of the eval runs at normal API rates, and the skill tells you the size of the run and roughly how long it takes before the baseline. Anthropic gives no figure per run: it depends on the number of cases, repeats and the model. Since version 2.1.285 /claude-api cannot be run from Remote Control clients. The same version marks answers cut off at max_tokens as truncated and counts them separately instead of folding them into the average.

Our take: start with build-eval and an honest look at the scores page, not with hillclimb. If you read a dozen scored transcripts and disagree with the grader, no number of rounds will help afterwards. Anthropic says the same: scoring failures are among the most common ways an eval is misconfigured.

Versions and prices as of October 5, 2026

Commands as described in Anthropic's post of September 28, 2026 and the Claude Code changelog 2.1.284–2.1.285. Models, their prices and the skill's behaviour change: check your Claude Code version and current API pricing before a run.

Sources3expand
  1. Lance Martin, Anthropic, “Automating eval design and hillclimbing with Claude”, claude.dev, September 28, 2026 — https://claude.dev/blog/automating-eval-design-and-hillclimbing/
  2. Claude Code changelog, versions 2.1.284 and 2.1.285, September 28–29, 2026 — https://code.claude.com/docs/en/changelog
  3. anthropics/claude-code, CHANGELOG.md — https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md

Comments