Hindsight: An Agent Memory System That Really Impressed Me

Hindsight: An Agent Memory System That Really Impressed Me
AI #memory

Anyone who works with Claude Code, Codex CLI or opencode every day knows the problem: at the end of every session, the agent is back to square one. You explain the tech stack, the team conventions, and why you abandoned Redis caching three months ago for what feels like the tenth time. Notes in system prompts or CLAUDE.md files help, but they do not scale—they capture only a static snapshot, not a growing, evolving knowledge base.

I have been using Hindsight, an open-source memory system from Vectorize.io, for a few weeks, and I am pretty taken with it. Or let us call it what it is: for the first time in a long while, a software solution has genuinely excited me. It does exactly what you imagine “agent memory” should do when you do not reduce it to the naive vector database solution. It has opened up new worlds for me, both conceptually and technically.

Why should you use Hindsight?

Before I get into the technical details, here are the three reasons that tipped the balance for me:

  • Greater context continuity: Agents retain relevant prior knowledge across sessions, delivering consistent, personalized recommendations instead of starting from scratch with every new session.
  • Better accuracy on complex questions: Temporal and graph retrieval combined with automatic observation consolidation noticeably improve the hit rate, particularly for historical or indirectly connected facts.
  • Controlled, explainable memory logic: Missions and directives create safety and compliance guardrails, while evidence tracking makes it clear what an answer is actually based on.

What Hindsight is—and what it is not

The most important point first: Hindsight is not simply a vector database with a pretty API wrapped around it. Anyone who throws in raw data and hopes cosine similarity will produce the right answer is regularly disappointed by naive RAG approaches. Once you have a few hundred memories covering different topics and time periods, the principle of “embed everything, retrieve Top-K” often breaks down.

Hindsight takes a different approach. PostgreSQL with the pgvector extension provides storage in the background, but what gets used is not the raw data. It is the result of an extraction process: structured facts, resolved entities (“Alice” and “my colleague Alice” become the same entity), timestamps for temporal analysis, and a knowledge graph mapping relationships between entities. If you are now worried about your data... okay, the documents are still stored in the Postgres database as “Document” too. But the raw documents are not the interesting part of the overall process; the processed data is.

Everything connects via MCP—Claude Code, Codex, opencode, Cursor, VS Code, practically any MCP-capable client can connect. For internal processing (fact extraction, entity resolution, reflect operations), Hindsight needs its own LLM provider, independently of the model the actual agent uses. That is a detail I like: you can deliberately use a cheap, fast model here because this processing runs in the background and does not affect the quality of the actual agent responses. I used the GPT-5.6 Luna model for my tests because, initially, I was paying out of my own pocket. Despite the inexpensive model, I got really good results. You can also use multiple models in Hindsight and fine-tune everything. Honestly, I have probably tried only 30–40% of the features so far. But that is fine. You get results quickly, can work with the data, and can improve it yourself. How? That is coming next.

The three core operations

Hindsight offers three central operations, available as MCP tools:

retain—Store. You feed Hindsight text (a single fact, an entire conversation, a document), and an LLM extracts structured facts, resolves entities, generates embeddings, and indexes everything for later searches in the background. Important: you should not summarize or extract facts yourself beforehand—Hindsight needs the full context of a conversation, otherwise “yes, exactly” or “I will take option 2” becomes meaningless text.

recall—Search. Four retrieval strategies run in parallel: semantic search, BM25 keyword matching, graph traversal across the knowledge graph, and temporal filtering. The results are then sorted using cross-encoder reranking. This is the real advantage over a simple vector search: a question such as “What did we decide about caching?” finds the right answer even when the original memory uses different terminology.

reflect—Synthesize. Reflect goes beyond simple retrieval and lets an LLM draw conclusions across multiple memories—for example, in response to “Which tech stack would you recommend based on my previous decisions?”

flowchart LR    
    subgraph YourApp["Your Application"]
        A["AI Agent"]
    end

    subgraph Hindsight["Hindsight"]
        API["API Server"]

        subgraph MemoryBank["Memory Bank"]
            direction TB
            MM["Mental Models"]
            OBS["Observations"]
            ME["Memories & Entities"]
            CH["Chunks"]
            DOC["Documents"]

            MM --> OBS
            OBS --> ME
            ME --> CH
            CH --> DOC
        end
    end

    A -->|retain| API
    A -->|recall| API
    A -->|reflect| API

    API --> ME

Extraction strategies and tagging

One point that is underestimated in practice: not every data source should be processed the same way. Hindsight offers named retain strategies for this—for example, verbatim for content that should be stored word for word, or chunks for larger documents split into sections. These strategies can be defined in the bank configuration and selected explicitly in the retain call depending on the source.

There is also a tagging system with so-called entity labels: controlled vocabularies following the key:value pattern, such as user:christian or topic:architektur, which can be extracted automatically during storage. For purely content-based searches, you often do not need this at all—entities end up in the knowledge graph anyway and drive graph-based search. For explicit filtering on entity-like values, however, tags are worth their weight in gold, especially when you want to keep multiple projects or customers clearly separated within a bank.

Mental Models: living documents

What convinced me most were the Mental Models. You define a question—for example, “Which architectural decisions were made for project X?”—and Hindsight generates a document that updates automatically whenever new memories are added or existing documents change. You therefore do not have to run a full reflect across the entire memory collection for every request. Instead, you get a kind of precomputed, always up-to-date summary.

Mental Models

Technically, this runs asynchronously through a queue: the actual LLM-assisted generation happens in the background, an API call initially returns only an operation ID, and the finished result is available after a few seconds. This is thoroughly thought through, but it also has a downside—more on that shortly.

Everything is processed asynchronously in a queue

Scaling is not a problem either, because the architecture also includes worker instances that can be scaled using Kubernetes. All of that is possible. In my homelab setup, a single worker in a fixed Docker container does the job.

Visualization

The Control Plane—Hindsight’s admin UI—includes the Constellation View: an interactive, zoomable graph visualization of entity relationships with heat-gradient coloring, also usable in dark mode. For me, this is more than just a nice feature. With a growing knowledge graph, the visual overview is enormously helpful for getting a feel for what your memory bank actually “knows” and where extraction might have gone wrong.

Graph visualization

Of course, you can also filter, zoom, and inspect the individual underlying documents.

Feature overview

A few points from the full feature list that have not come up in this article yet but are part of the picture:

  • Hierarchical memory types: Alongside Mental Models (curated, self-defined summaries), there are Observations—evidence-backed beliefs automatically consolidated from raw facts—as well as World Facts and Experience Facts. This separation ensures that reflect checks Mental Models first when answering a question, then Observations, and only turns to raw facts as a last step.
  • Multi-strategy retrieval (TEMPR): The four parallel search strategies in recall—Semantic, Keyword/BM25, Graph, Temporal—operate under the name TEMPR and are merged using reciprocal rank fusion before cross-encoder reranking takes over. This makes retrieval robust even for questions involving temporal or graph-based connections.
  • Observation consolidation: After every retain call, a consolidation process runs automatically in the background, comparing new facts against existing Observations, deduplicating them, and adding evidence references. Observations that may have been superseded by newer, not-yet-consolidated facts are marked “stale” and verified against the raw facts before use.
  • Temporal reasoning: Queries such as “last spring” or “in June” are resolved not just semantically but explicitly by time—a field in which pure vector search typically fails.
  • Mission and policy configuration: At bank level, you can define a mission specifying which knowledge should be prioritized, along with immutable directives as compliance guardrails and disposition traits such as skepticism, literalness, or empathy that influence how reflect weighs arguments. Important: these settings affect only reflect, not recall.
  • Evidence-backed answers: Observations refer to their source memories, including quoted evidence and proof counts. This makes an answer’s derivation understandable rather than something you accept as a black box.
  • Clients & SDKs: Official SDKs for Python, TypeScript and Go, plus a CLI and HTTP API—making integration into existing agent stacks beyond MCP straightforward.
  • Deployment options: Alongside the local Docker setup I use in my homelab, there are Helm charts for Kubernetes, a pip installation for quick local testing, and an integration hub for connected data sources.

How does data get into the system?

Data can enter memory in several ways. When using the MCP server, you can simply store something through the retain tool. Tags and a unique document ID can be defined for each captured document. Defining the document ID is optional. If none is supplied, the system creates a UUID.
If you want to store raw data, such as chat histories, using the generated UUID is sufficient. For data such as the content behind a web page, the document ID can also be a URL. If a new document is submitted with the same document ID, Hindsight updates the existing entry too.
Hindsight also lets you define different data extraction strategies for different sources. You can influence RAG-specific parameters such as chunk size as well.

Because Hindsight was designed consistently so that everything the UI does communicates through a REST API, using that same API to import data is straightforward too.
As an example, I installed an n8n community plugin that simply exposes the three basic operations in an n8n node.

Example n8n workflow with the Hindsight community node

My homelab setup

I set everything up in my homelab rather than the cloud—two LXC containers, one for the API and one for the admin UI, or Control Plane. I did not have to set up a new database server: I already had a PostgreSQL server with the pgvector extension running and only needed to create an additional database there. That made setup considerably simpler. Anyone already running a Postgres instance with pgvector is ready in minutes, rather than having to wrestle with the embedded PostgreSQL variant from the Docker quickstart.

For the internal processing LLM provider, I quickly connected OpenRouter and chose GPT-5.6 Luna—the inexpensive, latency-optimized variant from OpenAI’s GPT-5.6 series, intended for precisely these high-frequency but not particularly demanding background tasks such as fact extraction. The model can easily be swapped out, though: Hindsight supports a broad selection of providers, from OpenAI, Anthropic and Gemini to local models through Ollama or llama.cpp. Anyone who prefers to keep internal processing completely offline, without external API calls, can now do that too.

Integration with other systems

Via MCP, Hindsight can connect practically anywhere a client speaks the standard—Claude Code, Codex CLI, opencode and other coding agents are supported directly, in some cases even with automatic ingestion without an additional setup step.

Claude Code MCP tool

The MCP server is built directly into Hindsight and can be configured separately for each memory bank. You can also restrict the MCP tools.

MCP tool settings

I have not restricted anything in my setup because I already have a LiteLLM gateway in between anyway.

Use cases

The most obvious use case for me personally is not primarily technical at all: I want to prepare data about the football club Wormatia Worms to lighten the workload of the office staff. Instead of answering every request manually, an agent could draw on a Hindsight bank containing historical information, club rules, and recurring questions, genuinely taking on some of the communication burden.

Alongside that, I am currently building a second bank of my own: I am adding the club’s history, season results, and player profiles. The interesting part here is not simply retrieving individual facts, but having the system independently discover statistical peculiarities and relationships across the imported data—for example, that a particular season had an unusually high number of draws, or that a player scored disproportionately often across multiple seasons. That is precisely what reflect is for: not just retrieving facts, but drawing conclusions across the entire dataset. This becomes especially interesting for the club’s archivist, who wants to ask questions such as “When was the longest winning streak?”—questions that no individual record answers, and that emerge only from considering many season results together.

In technical project environments, I see at least two more useful scenarios:

  • Architecture and policy memory: An agent that keeps track of software architecture decisions, coding guidelines, and project conventions and is available to developers as a question-and-answer resource directly in their coding tool—without first having to search a wiki.
  • Requirements completeness checks: In projects where requirements are captured continuously, an agent with access to the accumulated memories can check whether all necessary documents and information are available and specifically highlight gaps. Especially in business contexts with many stakeholders, this is a scenario that delivers real value.

Assessment

What works well: getting started is pleasantly straightforward, especially if—as in my case—a Postgres/pgvector instance is already available. MCP as the connection standard means you do not have to build a separate integration for every client. In practice, the combination of multi-strategy retrieval and reranking produces noticeably better results than a simple vector search, particularly with larger memory collections spanning many topics.

Are there things about this approach that bother me? Not really, but you should be aware that as the amount of information grows, and with every Mental Model, the effort required to keep everything up to date increases too. Token costs can certainly add up.
Every retain operation means additional LLM calls for internal processing, which accumulate at high volumes. And the ecosystem’s maturity is not set in stone yet. Integrations can change at relatively short notice.

It really comes down to whether you have recognized context/memory as something valuable for yourself.

Business outlook

For a project environment like the one I know at valantic, I see the greatest leverage in agents retaining project context, customer preferences, and architectural decisions across sessions. That means less time spent repeatedly explaining context and more time available for actual value creation. It is a tangible productivity gain, particularly in longer projects with changing points of contact.

At the same time, governance cannot be ignored: where do the extracted facts end up, who has access to the memory bank, and how cleanly can customers or projects be separated? Self-hosting via Docker or, as in my case, LXC containers in your own homelab is a valid answer here—you retain full control of the data instead of entrusting it to a cloud provider.

Conclusion

Hindsight is the first agent memory system that immediately and genuinely excited me, because it does not reduce memory to “throw text into a vector database.” Instead, it combines structured facts, entities, temporal context, and a knowledge graph. Setup in my own homelab was surprisingly straightforward thanks to the existing Postgres/pgvector infrastructure, and the combination of the MCP standard and flexible LLM provider selection makes the system interesting for very different scenarios—from day-to-day club work at Wormatia Worms to architecture memory in technical projects. I will continue expanding this over the coming weeks and report back once concrete workflows emerge.