CodeariaAcademy
Article cover: a mechanism taken apart under a spotlight, with the caption What's inside the agent
October 6, 202613 min readAI AgentsClaude Code

4 million lines of code and not a single LangChain import. How are Claude Code, Codex and nine other agents actually built?

Summaries call it an agent comparison, but it ranks nothing. Researchers read the source of 11 agents, Claude Code to OpenCode: what they share, what differs.

In this article7
In short

Harness engineering is the design of everything that wraps the model in a coding agent: the loop, tools, memory, permissions and extensions. The arXiv preprint 2609.00006 (15 July 2026, 83 pages) studies the source code of 11 agents: Claude Code, Codex CLI, Gemini CLI, Mistral Vibe, OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode and OpenClaw. The authors identify 7 subsystems, 29 recurring patterns and 18 recommendations. The headline findings: across 4 million lines of code, no agent uses LangChain or a similar framework, and none searches code with vector embeddings; SKILL.md skills ship in 9 of 11 agents, MCP in 8; and within a quarter Codex adopted Claude Code's hooks almost word for word. The paper has no ranking and no benchmarks, and Claude Code is analysed from a March source snapshot.

Roughly 4 million lines of Python, TypeScript and Rust. Eleven coding agents, Claude Code, Codex and Gemini CLI among them. And zero imports of LangChain, LangGraph, AutoGen or any other agent framework. Gemini CLI does not even use Google's own frameworks.

That is one of two gaps found by four authors from Inclusive Brains and Wavestone AI Lab after reading the agents' source code end to end. The preprint is titled "Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents". It runs to 83 pages and is the second edition of an April study that covered eight agents.

People are writing about it as a comparison of which agent is best. It is not one, and the authors say so plainly: they ran nothing and ranked nothing, they only described how things are built. What it does show is what sits inside these agents. We looked at what a harness costs in money separately, when we compared one model across Claude Code, Codex and Pi. Here we look at what it is made of.

What harness engineering is and the seven parts of an agent

The paper opens with a formula: agent = model + harness. The model supplies the intelligence; the harness handles everything else: the think-act-observe loop, tools, context, permissions and extensions. The term harness engineering only entered circulation in February 2026, and within five months it had its own guides and its own line of arXiv papers.

The authors split every harness into seven subsystems. All eleven agents have each of them, even if only as a deliberate decision not to build it. The interesting part is the spread: how little code a subsystem can take, and how much it sometimes does.

SubsystemSimplestMost complex
Agent loopMini-SWE-Agent: a linear while loop with one bash toolOpenHands: event log and parallel actions
Model connectionMini-SWE-Agent: one LiteLLM callHermes: five custom transports, 29 provider profiles
ToolsMini-SWE-Agent: bash onlyClaude Code: 43 typed tools
Memory and contextMini-SWE-Agent: the full history, verbatimCodex: cross-session memory maintained by the agent itself
Permissions and safetyMini-SWE-Agent: step and cost limitsCodex: rules, an LLM command reviewer and sandboxes on three OSes
SubagentsAider: none, by designClaude Code: subagents that launch subagents
ExtensionsMini-SWE-Agent: Python protocolsPi: everything is an extension; Codex: a plugin marketplace

The left column is almost entirely Mini-SWE-Agent: about a hundred lines of Python, one tool, and by its authors' own report 74%+ on SWE-bench Verified. That gives the paper its first observation: loop complexity does not predict results. Most of the code in the big agents goes into things a benchmark does not measure. In OpenCode, roughly three fifths of the non-test code is clients: the terminal UI, web, desktop and SDK.

What none of the 11 agents has

Two gaps survived both the larger sample and a second look three months later.

No agent framework. Every loop is hand-written on its language's async primitives. The authors do not miss the irony: the term harness engineering was coined and developed inside LangChain, yet no harness ships LangChain libraries.

No vector search over code. No RAG over the repository. Agents find code with ripgrep, glob, tree-sitter and context files such as AGENTS.md and CLAUDE.md, which they discover by walking the directory tree. Embeddings appear only in conversation memory (on by default in OpenClaw), never over source code. The authors' explanation is simple: code has exact structure (paths, symbols, parse trees) and changes by the minute, so an embedding index goes stale fast.

If you are building your own agent, these are two direct recommendations from the paper: do not adopt a framework for the runtime, and do not build RAG over code until you have shown it beats ripgrep plus tree-sitter on your tasks.

Skills overtook MCP

agents support SKILL.md skills. 8 of 11 support MCP. In April the two were tied

Barbaste et al., arXiv 2609.00006, July 2026

Only Aider and Mini-SWE-Agent lack skills. Pi broke the tie with an explicit stance: skills and command-line tools with a README, and no MCP at all. The authors' recommendation: skills for repeatable procedures and domain knowledge, MCP for external systems such as databases and internal APIs, in that order.

Within a quarter, skills grew the trappings of an app store: registries in four agents, trust levels and quarantine in Hermes, provenance checks in OpenClaw. And the first skills that an agent writes for itself. If you install other people's skills, the authors advise treating them like packages from the internet, not like text files. For picking from what already exists for Claude Code, see our ecosystem breakdown: what to take and what to skip.

In three months, agents started copying Claude Code

This is the newest part of the paper. The authors did not replace the eight agents from April; they updated them and diffed source code a quarter apart. In April, the similarities were mostly independent rediscovery of the same ideas. By July they can be traced through the code.

  • Hooks. In April, user hooks were a Claude Code feature. By July, 9 of 11 agents have them. Codex took Claude Code's event vocabulary almost verbatim: PreToolUse, PermissionRequest, PostToolUse, PreCompact, SessionStart, UserPromptSubmit, SubagentStart, SubagentStop, Stop. It added one event of its own, PostCompact.
  • Migration. Codex finds Claude Code sessions in ~/.claude/projects, imports them and offers to convert ~/.claude/settings.json into its own config.toml.
  • Format. OpenHands reads the Claude Code plugin manifest; OpenCode reads its skills directory.
  • Pace. Deferred tool loading (schemas stay out of the prompt and the agent looks them up on demand) spread from one agent to three in a quarter. Plan mode now exists in all four agents from model vendors, up from two. The authors write that a distinguishing feature in this space lasts weeks.

A second shift is quieter but matters more to anyone writing rules for an agent. Codex dropped "don't commit" and "don't do extra work" rules from prompts for newer models, and Mistral Vibe removed a hard Never Commit rule. Behaviour is moving out of prompt text, which the model reads and may ignore, into settings the harness enforces itself.

For Claude Code users the takeaway is practical: the hooks and skills you write are no longer locked to one agent. How hooks work and how the new mods differ from them is covered in our article on Claude Code mods.

How coding agents differ: memory, sandboxing, subagents

If the insides are nearly identical, where is the real choice? According to the paper, in three places.

Memory. Context compaction has converged: 7 of 11 agents summarise the history with the model once a threshold is hit. Claude Code compacts when 13 thousand tokens of the window remain; Gemini CLI does it at half the window, keeping the last 30% verbatim. The difference now is who writes long-term memory. In Codex a dedicated subagent maintains it; in Gemini CLI a separate subagent files entries into an inbox that you approve with the /memory command; in Hermes it is files with a hard character limit.

Sandboxing. OS-level isolation costs thousands of lines of code, and it is a choice, not a function of size. Codex and Gemini CLI maintain their own sandboxes for Linux, macOS and Windows. Claude Code plugs in Anthropic's sandbox as an option. Hermes, one of the largest agents, does not isolate processes at all, but keeps 12 hard blocks that apply even in --yolo mode. Pi rejects sandboxing with a stated reason: partial in-process isolation looks like a security boundary without being one.

Subagents. The coordinator-hands-work-to-workers pattern emerged independently in almost every agent. The goals differ: Claude Code saves on a shared prompt cache across child agents, Codex builds a thread tree with depth tracking, Mistral Vibe runs subagents sequentially for simplicity. Pi deliberately stays single-agent. Citing Anthropic, the authors note that multi-agent systems burn roughly 15 times more tokens than a normal chat, and advise against leaving a single agent without a clear reason. Where a coordinator is genuinely needed, we looked at Paperclip, which gives agents a boss and a budget.

What the study does not cover

Claude Code was analysed from old code

Claude Code's source is closed. The authors relied on a March 2026 source snapshot that circulated publicly, not on an official release. In July the shipping version was 2.1.206; as of 6 October it is 2.1.292. The authors themselves call the Claude Code analysis the weakest part of the paper in terms of reproducibility.

The most discussed agent in the sample is described from the oldest code. That does not make the conclusions wrong: architecture changes more slowly than version numbers, and the authors split their claims into "inventory" claims (how many tools, which features) and "structural" ones (how the loop works, what the harness is made of). The former, they say, go stale within weeks; the latter hold for now. The "43 tools" figure belongs to the first kind.

No ranking. The April edition had a SWE-bench Verified table; the July edition removed it because the numbers are self-reported on different models and settings. If you saw a summary saying "the study showed which agent is best", that is not in the paper.

Nothing was run. Only code was read. The authors make no claims about how fast or cheap each agent is.

It is not news. The arXiv number is from September, but the text was submitted on 15 July. Since then Claude Code has shipped more than eighty version numbers and gained mods, so some of the tables already describe the past.

The 90 lines are not benchmarked. The minimal agent skeleton at the end of the paper implements 10 of the 18 recommendations. The authors themselves call the idea that it would match Mini-SWE-Agent on a strong model a conjecture "without proof".

What to take from it

Our view, and you may disagree: the agent loop is no longer what agents compete on. The model solves the task: Mini-SWE-Agent, with a hundred lines of harness, holds its own (by its own report) next to agents a thousand times larger. The competition is everything around the loop: skills, hooks, plugins, importing other agents' sessions. The authors reach the same conclusion: in the first half of 2026 the harness turned from a tool into a platform. We could be wrong if the next generation of models starts to gain noticeably from a complex loop rather than from its surroundings. The paper has no such data yet.

If you are choosing an agent, look less at the tool count and more at the three things from the section above: how it writes long-term memory, whether it has OS-level sandboxing, and whether it reads your skills and hooks. The answer to the last one is now more often yes than no.

If you are building your own, the authors offer a short sequence, and much of it matches Anthropic's engineering advice:

  1. Start with a linear loop and a single bash tool. Add tools only for a problem you have actually observed: file read and write first, then search, then exact string replacement.
  2. Once you have more than fifteen tools, load their schemas on demand. In Claude Code this cuts the startup prompt by about 40%.
  3. For file edits on a strong model, use exact replacement of a unique substring rather than line numbers: models get line numbers wrong more often than context.
  4. Search code with ripgrep and tree-sitter, put context in per-directory AGENTS.md files, and skip the runtime framework.
  5. Stay with a single agent until there is a step where parallel search is clearly faster than sequential.

If you do not build agents but work in Claude Code, the most useful part of the paper is the section on skills. A good place to start is ready-made skills that cover routine work.

Versions as of 6 October 2026: Claude Code 2.1.292; arXiv preprint 2609.00006v1 of 15 July 2026, with agent versions as of July. Agents update weekly, so check current capabilities before choosing.

Sources2expand
  1. Paul Barbaste, Tristan Darrigol, Germain Vu, Tom Wiltberger, "Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents. A Source-Code Study of Eleven Systems", arXiv 2609.00006v1, 15 July 2026 — https://arxiv.org/abs/2609.00006
  2. npm, @anthropic-ai/claude-code, version history, checked 6 October 2026 — https://www.npmjs.com/package/@anthropic-ai/claude-code

Comments