
Hindsight writes its own wiki about your repository. Does Claude Code need that kind of memory?
Teach Claude Code and Codex to remember why the project works the way it does: how Hindsight works, where its 94.6% comes from and which defaults to change.
In numbers
In this article6
Hindsight is an open-source, MIT-licensed memory system for AI agents from Vectorize. It does more than store conversations: an LLM extracts facts from every session, background work consolidates them into observations backed by quoted evidence, and the system keeps knowledge pages about the project that it rewrites as it learns. There are three operations: retain stores, recall searches four ways at once, reflect draws conclusions. One package covers Claude Code, Codex, Cursor and about fifteen other agents with a shared memory bank per repository, and the server ships with built-in MCP. On 28 September 2026 the project had 38,179 stars and 4,520 new ones in a day; the latest release, v0.10.1, came out on 21 September. The headline 94.6% was measured on LongMemEval S; on the largest set of the team's own benchmark the score is 64.1%.
4,520 stars in a day. That is what the vectorize-io/hindsight repository collected on GitHub on 28 September 2026, nearly twice as many as Paperclip, the agent orchestrator, which sits at the top of the same list. Hindsight is memory for agents, and from the first line of its README the authors push back on the usual idea of memory: most systems teach an agent to recall conversations, they want the agent to learn.
If you work in Claude Code, this will sound familiar. The session ends, the context is gone, and the next day the model no longer knows why the project rounds only the invoice total and not every line. We have already covered one answer to that: claude-mem records what happened in a session and feeds the relevant parts into the next one. Hindsight tries to remember something else: not what you did, but why you decided it.
Memory system for agents: retain, recall, reflect, observations and knowledge pages. Python server on PostgreSQL, clients for Python, Node.js and Go, built-in MCP. MIT license.
What Hindsight is and how memory differs from chat history
Hindsight has four kinds of memory. World facts ("the project uses Next.js 15"). The agent's own experience ("I moved rounding into finalize()"). Observations: beliefs the system consolidates in the background from many facts, with verbatim quotes as evidence and a count of confirmations. And mental models: ready answers to questions you asked the bank once.
The interesting part of observations is that new information does not overwrite the old. When a fresh fact arrives, the observation is refined: it gets stronger, weaker or gains detail. In the documentation demo, instead of "someone once wrote about rounding", the agent gets "we round once, half-to-even, in InvoiceTotal.finalize(), three sources".
In recent versions mental models grew into knowledge pages. These are living documents the bank writes about itself: a component map, conventions, key decisions, current initiatives. They sit in folders like a wiki, they are searchable, and you can export them to disk as plain markdown. Reading such a page is a database lookup with no retrieval and no LLM call, so the agent starts from a ready page instead of rebuilding context from scratch.
Retain, recall, reflect: how the agent's memory works
All the work comes down to three calls. retain takes text, and an LLM pulls facts, dates, entities and relations out of it. recall searches four ways in parallel: by meaning through vectors, by exact words through BM25, through the graph of links between entities, and by time. The results are merged, reranked by a separate model and then trimmed to a token budget. reflect does not search, it reasons over the bank: "what are the risks in this project", "why do some customer emails work and others don't".
docker run -p 8888:8888 -p 9999:9999 \-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \ghcr.io/vectorize-io/hindsight:latest# API on :8888, UI on :9999, MCP at /mcp/{bank_id}/client.retain(bank_id="shop", content="We round the invoice total half-to-even")client.recall(bank_id="shop", query="How do we round?")client.reflect(bank_id="shop", query="What matters about money in this project?")
Almost any model will do for fact extraction: the README lists more than 25 providers, including local Ollama and LM Studio. There is also an option without an API key: the claude-code provider runs on your Pro or Max subscription.
How to connect Hindsight to Claude Code and Codex
Here is the first trap. Many guides still say claude plugin install hindsight-memory. That plugin works, but it is no longer developed: it was replaced by the @vectorize-io/hindsight-coding-agents package, one for all agents.
npx @vectorize-io/hindsight-coding-agents install claude-code --server daemon# or everything it finds on the machine:npx @vectorize-io/hindsight-coding-agents install all# import past sessions of this repository:npx @vectorize-io/hindsight-coding-agents install claude-code --import-conversations
For Claude Code the installer adds three hooks to ~/.claude/settings.json, registers MCP and installs a skill that tells the agent how to work with this memory. After that it runs on its own. At session start the project bank is filled from git history, on the first prompt a short reflect answer about your task lands in the context, and after every response the session goes into the bank. How project memory works in Claude Code itself, without plugins, is what we cover in the course on persistent memory for every project.
The key idea of this package is that there is one bank per repository, shared by all agents. What you explained to Claude Code in the morning, Codex knows in the evening in the same project. If you already run several agents on one codebase, this removes part of the pain that makes people install orchestrators: every new agent starts from zero.
94.6% on LongMemEval: what is behind the number
On its benchmarks page Hindsight calls itself first on Agent Memory Benchmark across all sets. The numbers from there:
| Set | Hindsight score |
|---|---|
| LongMemEval S | 94.6% |
| LoCoMo 10 | 92% |
| BEAM 1M | 73.9% |
| BEAM 10M | 64.1% |
The first thing you won't see in search results: Agent Memory Benchmark was built by the Hindsight team itself, and its March manifesto says so directly. The same manifesto says that benchmarks of LongMemEval's generation "now mostly measure whether your LLM can read": with a million-token window they test the model's reading more than memory. The headline number comes from exactly there.
Second, the bigger the memory, the lower the score. At ten million tokens it is 64.1%, so more than a third of answers are wrong. For comparison, the team's December paper reported 91.4% on LongMemEval, so the system has improved since then. The README also says the results were reproduced by researchers at Virginia Tech and The Washington Post, while competitors' numbers in the table are self-reported by the vendors.
Hindsight's score on BEAM 10M, the largest set of Agent Memory Benchmark. The same system scores 94.6% on LongMemEval S
We don't see this as a reason to skip it. The team explained why the old tests are outdated and published an open new one that shows both strong and weak spots. But read "94.6%" in reposts as a result on short histories.
Hindsight or claude-mem
The two tools look alike only from the outside.
| claude-mem | Hindsight | |
|---|---|---|
| What it remembers | What happened in sessions | Facts, experience and conclusions drawn from them |
| Where memory lives | On your machine | Cloud, your own server or a local daemon |
| Agents | Claude Code | Claude Code, Codex, Cursor and others, one bank per repository |
| What it needs | One command | A server or the cloud, plus a model for fact extraction |
Our opinion, which you can argue with: if you work alone in Claude Code on one or two projects, claude-mem or memory set up by hand in CLAUDE.md gives you almost the same at lower cost. Hindsight pays off where there are several agents, or where memory is also needed by a support assistant that has to remember every customer. We are wrong if the knowledge pages on your repositories turn out more accurate than the CLAUDE.md you maintain yourself: then it wins even for a solo developer.
What turns on by itself if you touch nothing
The coding agents package is designed to install with zero configuration, and its defaults are generous. The installer asks where to store memory, and the default option is Hindsight Cloud: accept it without looking and your sessions and git history go to an external server. Memory is enabled by default in every project on the machine. And in a new repository the plugin launches Claude haiku in the background for a code review capped at $2, repeating the review every 20 commits.
All of this can be turned off and configured, it's just better to decide before the first session than after. In the cloud, storing each record is free for the first 30 days, and operations are billed by tokens: $10 per million for writes and $0.75 for recall, with reflect at $0.05 per call. Your own server is free, you pay only for the extraction model.
Three settings before the first session
Choose where memory lives: --server daemon or self-hosted if code must not leave for an external cloud. Limit projects with optInOnly and optInPaths so memory doesn't switch on in a client repository. Turn on Memory Defense for the bank: it checks everything being written against 45 patterns for secrets and personal data, and either strips matches or blocks the write.
Versions and prices as of 28 September 2026: Hindsight v0.10.1, 38,179 stars. Five releases shipped between 7 August and 21 September, so check the current defaults in the documentation.
Sources9expand
- Vectorize, "vectorize-io/hindsight", README, checked 28 September 2026 — https://github.com/vectorize-io/hindsight
- Vectorize, "Hindsight releases", v0.10.1 of 21 September 2026 — https://github.com/vectorize-io/hindsight/releases
- Hindsight, "Coding Agents", documentation, checked 28 September 2026 — https://hindsight.vectorize.io/sdks/integrations/coding-agents
- Hindsight, "Claude Code", documentation, checked 28 September 2026 — https://hindsight.vectorize.io/sdks/integrations/claude-code
- Hindsight, "Benchmarks", checked 28 September 2026 — https://benchmarks.hindsight.vectorize.io/
- Nicolò Boschi, "Agent Memory Benchmark: A Manifesto", 23 March 2026 — https://hindsight.vectorize.io/blog/2026/03/23/agent-memory-benchmark
- Chris Latimer et al., "Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects", arXiv:2512.12818, 14 December 2025 — https://arxiv.org/abs/2512.12818
- Vectorize, "Hindsight Cloud Pricing", checked 28 September 2026 — https://vectorize.io/pricing
- GitHub, "Trending repositories today", 28 September 2026 — https://github.com/trending?since=daily
Read next
claude-mem: so you stop explaining your project every morningAugust 27, 2026
Agents now get a boss, a task queue and a budget. Do you need one if you already work in Claude Code?September 26, 2026
Same model, same success rate, twice the money. What are you actually paying Claude Code for?September 17, 2026
Comments