CodeariaAcademy
Cover of the Magnitude article: an object that stands for fitting to one specific chip, with a caption about kernels tuned to your chip
October 1, 202610 min readAI AgentsClaude Code

57 tokens per second instead of 30 on the same Mac. How does Magnitude beat llama.cpp?

Can a local model run your agent twice as fast with no new hardware? We break down Magnitude: kernels tuned to your chip, and where llama.cpp is still faster.

In numbers

decode speed on M4 Pro versus llama.cpp, in the authors' benchmark
one-time kernel tuning for your chip when you download a model
Magnitude's KV cache keys and values; llama.cpp ran 16-bit in the same test
slower than llama.cpp on M5 Max, per a user benchmark on HN
In this article6
In short

Magnitude is an open-source inference engine for agents from a YC S25 startup, launched on Hacker News on 30 September 2026. Instead of shipping prebuilt kernels for a broad class of hardware, as llama.cpp does, it spends about a minute tuning them to your exact chip when you download a model. In the authors' benchmark on a Mac M4 Pro, decode went from 30 to 57 tokens per second on the same model. Claude Code, Codex, Cline and other agents connect in one click, and the license is Apache 2.0. The caveat missing from the launch posts: in that test llama.cpp carried a KV cache twice as heavy, and on M5 Max users saw the opposite result, with llama.cpp roughly twice as fast. The authors acknowledged the kernel gap on newer chips and say they will close it.

30 and 57. That is how many tokens per second the same model produces on the same Mac M4 Pro: the first number in llama.cpp, the second in Magnitude. The hardware stayed the same; what changed were the kernels that run the model. For an agent that generates tens of thousands of tokens per session, that means waiting almost half as long at every step.

Cloud tokens for an agent cost real money: we worked out what a task actually costs with GPT-6 Sol and Opus 5.5. So an engine that squeezes extra tokens out of a laptop matters to anyone who has thought about moving part of an agent's work to a local model.

What Magnitude is and why it beats llama.cpp

Magnitude is built by Anders and Tom, a YC S25 team. Their earlier product was a browser agent, which lives in a neighbouring repository and has 4,134 stars. The magnitudedev/magnitude repository itself went through three lives over the summer: in June it was a CLI with a cloud key and credits, in August an agent running on local models, and since mid-September an engine plus a desktop app. So some of its 5.9k stars as of 1 October were earned by earlier versions of the project.

The idea is simple. llama.cpp and Ollama ship kernels compiled ahead of time for a wide class of hardware. Chip-specific engines are faster, the authors say, but support less. Magnitude writes kernels with tunable parameters and picks the values on your device. The engine is written in Rust, with its own GPU kernel runtime and autotuner. It runs on Apple Silicon, NVIDIA, AMD (via Vulkan) and plain CPU, on macOS, Linux and Windows.

Inference engine for agents: kernels tuned to your chip, a desktop app, one-click agent connections, an OpenAI-compatible API. Rust, Apache 2.0.

commit this monthchecked 1 October 2026

Kernels that tune themselves to your chip in a minute

Tuning happens once: when you download a new model, the engine spends about a minute trying kernel parameters until the gains stop growing. After that the model runs on the configuration it found.

The other half of the speed comes from memory. Magnitude stores the KV cache, where the model keeps the context it has read, with 8 bits for keys and 4 bits for values. According to the authors, the cache takes half the space and decode gets faster as a result. The rest is designed for agents rather than chat:

  • context memory is not reserved upfront; it grows with the session and is released when the agent stops;
  • parallel sessions that start the same way (system prompt, tool schemas) share a common prefix cache;
  • a model loads when an agent asks for it and unloads after it sits idle.

The catalogue is small for now: 15 models from six families, from Qwen and Gemma to Nemotron and Liquid LFM, ranging from 1.1B to 35B. Quantization is 4 bits and up only: below that, the authors say, models' reasoning and tool calls fall apart.

How to connect Claude Code, Codex or Cline to a local model

Setup goes through the app: download it from magnitude.dev, pick a model in the Discover tab, connect an agent in Connections. The magnitude CLI ships with the app. Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline connect in one click; everything else goes through an OpenAI-compatible endpoint.

With Claude Code there is one detail worth knowing in advance. The Connect button rewrites your ~/.claude/settings.json and routes Claude Code through the Magnitude gateway:

what Connect writes to ~/.claude/settings.json
"ANTHROPIC_BASE_URL": "http://ADDRESS:10100/inference/anthropic/proxies/claude-code",
"CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY": "1",
"model": "anthropic-local/MODEL_ID"
claude --model anthropic-local/MODEL_ID

Requests to cloud models also pass through the gateway, so Magnitude has to stay running for as long as Claude Code is connected. To go back to normal, click Disconnect and restart Claude Code. If you already have your own variables in settings.json, back the file up before connecting.

Where "twice as fast as llama.cpp" comes from

Magnitude's headline number was measured on one model and two machines:

decode tokens per second on a Mac M4 Pro 48 GB, Qwen 3.6 35B A3B, 4-bit, 64k context, no speculative decoding. On DGX Spark the gain is 19% (49 → 58); prefill is 9% and 23% faster

magnitudedev/magnitude README, authors' data, 1 October 2026

The benchmark task is synthetic: the prompt is filled with Moby-Dick up to 64 thousand tokens and the model is asked to repeat the last passage. On HN the authors were open about how they configured llama.cpp: flash attention on, 16-bit KV cache. They tried quantizing its cache to 8 and 4 bits, as in their own engine, but llama.cpp's decode dropped when they did. So in this comparison Magnitude's cache is almost half as heavy, and a lighter cache, by the authors' own account, speeds up decode. Part of the gain comes from that, not only from kernel tuning.

There is a second subtlety, and we found it in their own methodology. In session-bench, Magnitude runs the MLX version of the model and llama.cpp runs the GGUF, and the document says plainly that a shared model name does not prove the files are equivalent. The number is honest, but it describes the combination of engine, format and cache, not the kernels in isolation. The authors have not yet published a comparison with MLX; they say it is coming soon.

Where Magnitude is still slower

In the first day on HN, users ran it on their own hardware. The most visible gap is on the newest Macs:

Hardware and modelResult
M5 Max, Qwen3.8 Q6 with the DFlash2 drafterllama.cpp roughly twice as fast in both prefill and decode
M5 Max 128 GB, several Qwen modelsrapid-mlx: 175 tok/s and 64 ms to first token; Magnitude: 161 and 111 ms
M5 Pro 48 GB, Gemma 4 26Bgeneration faster than oMLX (82.8 vs 76.5), prefill 2.6× slower
RTX 5070 Ti, Gemma 4 12Bllama.cpp 20–30% faster in decode, the second card goes unused

The authors named the M5 cause themselves: the kernels do not yet use the Metal 4 matrix operations these chips introduced. According to one user, llama.cpp closed the same prefill gap only recently, in build b10853. Magnitude does not support several GPUs in one machine yet; that is on the near-term roadmap. Two more bugs surfaced on day one: a long model evaluation step before the first download, and understated recommendations on machines with a single 16 GB card.

Version 0.2, released 30 September

These are the first days of a public engine. Some of the user benchmarks were run on 0.2.1, and two more patches shipped the same day. The figures in this section will go stale fastest.

What we take from it

Our opinion, and you are free to disagree: Magnitude is worth trying if you run agents on a Mac with an M4 Pro or M4 Max and decode speed on long context is your bottleneck. The gain there is claimed on an open methodology, and you can reproduce it on your machine with session-bench. We will be wrong if the promised MLX comparison shows that a regular MLX engine on a Mac is no slower: then Magnitude stays a convenient wrapper rather than a new class of engine.

On M5, on consumer NVIDIA cards and on multi-GPU setups it is better to wait for the kernel updates. More interesting than the speed is the idea itself: an engine built for long agent sessions, with a shared cache and memory that grows and then lets go. We looked at what the harness around a model costs when we compared Claude Code, Codex and pi on price, and a local engine covers exactly the part of the bill that goes to tokens.

How to test it on your own machine in an evening

Take a model from the catalogue that you have already run in llama.cpp or MLX, and compare them on your real agent prompt rather than a short question: Magnitude's advantage is claimed on long context. Back up ~/.claude/settings.json before connecting Claude Code.

Versions and figures as of 1 October 2026; check the repository for current ones.

Sources6expand
  1. Magnitude, «magnitudedev/magnitude», README and benchmark against llama.cpp, checked 1 October 2026 — https://github.com/magnitudedev/magnitude
  2. Magnitude, «Session bench», benchmark methodology, checked 1 October 2026 — https://github.com/magnitudedev/magnitude/blob/main/inference-v2/session-bench.md
  3. Magnitude, «Claude Code», integration docs, checked 1 October 2026 — https://github.com/magnitudedev/magnitude/blob/main/docs/integrations/claude-code.mdx
  4. Magnitude, «Models», catalogue, checked 1 October 2026 — https://magnitude.dev/models
  5. Anders and Tom, «Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents», 30 September 2026, with user comments — https://news.ycombinator.com/item?id=49911995
  6. Magnitude, «browser-agent», checked 1 October 2026 — https://github.com/magnitudedev/browser-agent

Comments