Search

Agent Memory Needs Typed Claims, Routing, and Maintenance

A three-month GBrain + Hermes run: why raw transcripts are not memory, and the typed-substrate, routing, and maintenance loop that made durable recall inspectable.

Mohit9 min read

Filed under Agent Reliability· see every report on this topic

An agent with many tools but no persistent memory: every session restarts from zero.

Verdict. Do not buy a memory plugin and call memory solved if an agent needs to remember projects, decisions, and failures across sessions. Use a small operating system: typed knowledge, intent routing, selective ingestion, and recurring maintenance. In three months, Hermes as coordinator and GBrain as substrate grew from 178 pages / 682 chunks on 6 July to 1,144 pages / 2,099 chunks, all embedded, at publication time. It also failed in seven useful ways. [Receipt: reports/gbrain-maintenance/2026-07-06-dream-cycle; live get_stats, 2026-08-21]

This is a home-lab field report, not a universal benchmark. The verified GBrain 0.42.8.0-era stack uses Docker PostgreSQL/pgvector on a workstation over HTTP MCP; Hermes is the front door and coordinator. Embeddings used local Ollama nomic-embed-text at 768 dimensions. Search synthesis, think, dream/consolidation, and fact extraction used a local vLLM Qwen route, so the memory path made 0 external LLM API calls in the 2 June verification. [Receipt: reports/hermes-gbrain-verification-2026-06, 2026-06-02]

The operator and the leak: chat is not durable memory

The operator is a solo or small-team builder who uses an agent on Discord, Telegram, or a terminal every day. The leak is not a failure to recall yesterday’s chat. It is a past decision trapped in chat, a project without durable state, a repeated old failure, and a new session that rediscovers completed work.

Raw transcripts are not memory. They include acknowledgements, abandoned work, duplicate conclusions, secrets that should not be stored, and tactical state that expires before the write finishes. Re-embedding the lot creates a larger haystack, not an operating model.

Store typed claims, route intent, maintain the system

The working design has four parts:

Pages hold durable entities and reports; graph links hold relationships; facts and takes hold claims; embeddings support retrieval. A project page, a decision, a dated report, and a searchable chunk are distinct records.

Resolve the memory action before the model selects a raw tool: “What do we know about X?” maps to query, “synthesize this” to think, “what is hot?” to salience, and reads and writes stay explicit.

Batch session digests keep durable signal and skip noise. They are reviewable pages, not invisible per-message mutations.

Sync, consolidation, extraction, embedding, orphan review, link repair, and doctor checks prevent a well-indexed graveyard.

A vector index is useful retrieval infrastructure. It is not source scope, a graph, a decision ledger, a write policy, or a way to discover that a pipeline stopped running.

In the verified June architecture, GBrain held pages, sources, graph links, search/query, reports, facts, takes, and maintenance jobs; Hermes remained the interaction layer. The active source was default, pointing to the local brain directory, not one transcript store. [Receipt: agents/hermes-gbrain-integration; reports/hermes-gbrain-verification-2026-06, 2026-06-02]

At publication, get_stats returned 1,144 pages, 2,099 chunks, 2,099 embedded chunks, and 2,088 links. Same-day doctor reported 85/100 health and 86.0% inbound-link coverage, while warning that 1,141 of 1,144 pages had un-extracted edges. The graph makes missing relationships visible; it is not a trophy count. [Receipt: live get_stats and run_doctor, 2026-08-21]

Choose the smallest memory system that fits

Approach Good for Cost shape Catch
Raw chat-log re-embedding Fast prototype Storage and embedding usage Noise becomes retrieval context
DIY pgvector + embeddings One focused corpus Build and operator time No typed claims, graph, or maintenance policy
GBrain + Hermes router Durable projects and decisions Existing workstation electricity + operator time You own maintenance
Hosted memory SaaS (Mem0-class) Fast managed trial Pricing to be verified at QA Inspect data model and retrieval behavior
Notes + chat history Low-stakes work Near-zero infrastructure Recall stays manual and session-bound

This is not a ranking. For a small static document folder, DIY embeddings may be the right smaller system. For cross-project decisions with provenance, the missing components are the decision.

Route recurring intents; do not expose an 80-tool buffet

The first failure was the raw MCP surface. With 80-plus tools, the agent selected tools arbitrarily or skipped memory for direct answers and web search. That was an interface problem, not a model-intelligence problem. [Receipt: logs/kit-log, 2026-07-04, “Full GBrain Production Setup”]

The repair was a resolver, a brain-first rule before external lookups, and a local GBrain skill pack. The 2 June integration record confirms bundled skills were installed in ~/.hermes/skills/gbrain-skillpack/, which exists on the active Hermes machine. High-level tools became the normal route; low-level reads, graph calls, and writes became verification or repair tools. [Receipt: agents/hermes-gbrain-integration, 2026-06-02; local directory check, 2026-08-21]

Operator intent Default action Why
“What do we know about X?” Query Grounded candidate pages
“Pull this together” Think Multi-page synthesis, not chat recall
“What changed or matters?” Salience Recent activity patterns
“Read the record” Get page Canonical page
“Save this decision” Write page / fact path Reviewable durable state

Make this an operating policy, not an 80-item tool catalog the model must recall on every turn.

Batch ingestion is the policy

The working pattern is a daily batch digest, not capture of every utterance. The verified GBrain Session Digest Importer cron, job dd82b65c69ad, was created and smoke-tested on 2 June. It scans recent sessions, deduplicates durable signal, excludes acknowledgements, transient task progress, secrets, and likely-short-lived facts, then writes a reviewable report page. [Receipt: reports/hermes-session-handoff-and-digest-cron; agents/hermes-gbrain-integration, 2026-06-02]

On 20 August, the digest analyzed four substantive sessions, submitted extraction for three, and deduplicated to zero new facts. That is a useful no-op: a cron did not manufacture novelty because it ran. [Receipt: logs/kit-log, 2026-08-20 22:11 IST]

The local path costs electricity on an already-owned workstation plus operator time. Hosted memory adds embedding API usage plus LLM API usage for extraction, synthesis, and maintenance; prices are to be verified at QA. Local is not free. Its bill is the incidents and maintenance work below.

Seven failures that changed the design

Failure Receipt Design change
Local database never starts 11 July: Mac PGLite WASM failed on macOS 26.x; remote brain intermittently returned HTTP 503 Docker PostgreSQL/pgvector behind HTTP MCP; Mac thin client
Brain looks empty Identity and auth worked, but list_pages and search returned little: stale source personal; agent scope default Inspect counts; move pages; revoke orphan client; remove stale source; sync/embed/extract/dream
Tools exist but are not used 80-plus raw tools produced no habits Resolver and brain-first policy
Every message becomes memory Digest filters acknowledgements, credentials, ephemeral work, duplicates Batch, reviewable daily digest
Extractor dies quietly Stale alias caused Cannot connect to API; 1,312 legacy facts awaited fence backfill Maintenance reports surface failures
Transport fails unattended MCP unavailable in cron; large writes unreliable; 17, 18, 22 June remote access unavailable CLI/SSH fallback; read back every write; local fallback reports, no live mutation
Graph count masks broken extraction 13 July: 1,573 links, 87 orphans, but 825 of 828 pages had un-extracted edges Deliberate manual linking; link count is not graph health

The first row’s receipts are logs/kit-log, 2026-07-11 and reports/local-vllm-and-hermes-runtime-2026-06. The source-mismatch repair is reports/hermes-gbrain-verification-2026-06, “Critical Fix Pattern: Source Mismatch.” Tool adoption is logs/kit-log, 2026-07-04 and agents/hermes-gbrain-integration. The auto-write rule is reports/hermes-session-handoff-and-digest-cron and logs/kit-log, 2026-08-20. Extractor failures are reports/gbrain-maintenance/2026-07-06-dream-cycle, reports/gbrain-maintenance/2026-08-20-2000, and logs/kit-log, 2026-07-14. Transport receipts are reports/hermes-session-handoff-and-digest-cron and logs/kit-log, 2026-06-17/18/22. Graph receipts are reports/gbrain-maintenance/2026-07-13-2000 and live run_doctor, 2026-08-21.

The growth record shows correction as well as scale:

Date Observed state Receipt
6 July 178 pages, 682 chunks, 100% embedding coverage reports/gbrain-maintenance/2026-07-06-dream-cycle
7 July Migration surfaced 638 orphans after hundreds of untracked pages reports/gbrain-maintenance/2026-07-07-dream-cycle
13 July 828 pages, 1,385 chunks, 1,573 links, 2,192 active facts, 170 takes, 87 orphans, health 55/100 reports/gbrain-maintenance/2026-07-13-2000
20 August 1,136–1,137 pages, 100% embedding, 2,078–2,090 links, doctor 85/100; actionable orphans 114–154 reports/gbrain-maintenance/2026-08-20-2000, 2026-08-20-2135
Publication 1,144 pages, 2,099 chunks, 2,088 links, doctor 85/100 Live get_stats and run_doctor, 2026-08-21

The drop from 638 to 87 orphans was consolidation and intentional linking, not magic. The 114–154 actionable orphan range was largely system artifacts requiring curation rather than blanket links.

When not to build the memory OS

Skip it, or start smaller, if any boundary applies:

  • You only need static search over a small document set. Use files and keyword search or a small vector index.
  • Nobody owns daily ingestion review and weekly maintenance. A local stack without an operator is deferred failure.
  • The agent lacks a trusted write boundary. Establish explicit write approval and secret exclusion first.
  • Durable signal is undefined. Start with one project, decisions, and dated reports; do not import every chat.
  • You require a retrieval-quality percentage before deploying. We ran no controlled retrieval evaluation, so we do not claim a RAG accuracy improvement.

Interpret doctor carefully. A small corpus can score poorly while connection, schema, sync freshness, embeddings, search mode, and dream plumbing work. Separate plumbing from corpus quality, then treat chronic warnings such as extraction lag as work. [Receipt: reports/hermes-gbrain-verification-2026-06, “Doctor Interpretation”]

Bottom line

Durable agent memory is an operating system, not a plugin checkbox. This three-month Hermes and GBrain run supports a narrower claim: typed substrate, resolver routing, batch ingestion, and maintenance made memory inspectable. Raw-tool sprawl, source drift, silent extractors, flaky transport, and broken link extraction caused the operational trouble.

Start with one source, a few high-level routes, a reviewable daily digest, and weekly maintenance. The decision is not “which vector store?” It is whether you will operate the system that decides what gets remembered.


More on this decision, three ways to look at it:

ROUTE • INGEST • MAINTAIN

ROUTE • INGEST • MAINTAIN

Sources

  • reports/hermes-gbrain-verification-2026-06, verified architecture, source-mismatch repair, and doctor interpretation (2026-06-02)
  • agents/hermes-gbrain-integration, Hermes/GBrain operating rules, local routing, and skill-pack installation (2026-06-02)
  • reports/hermes-session-handoff-and-digest-cron, batch digest policy and cron receipt (2026-06-02)
  • reports/gbrain-maintenance/2026-07-06-dream-cycle, 2026-07-07-dream-cycle, and 2026-07-13-2000, dated growth, migration, health, and graph receipts
  • reports/gbrain-maintenance/2026-08-20-2000 and reports/gbrain-maintenance/2026-08-20-2135, latest pre-publication maintenance receipts
  • logs/kit-log, dated incident log, including 2026-06-17/18/22, 2026-07-04/06/11/14, and 2026-08-20 entries
  • Live GBrain get_stats and run_doctor snapshot, publication-time counts and warnings (2026-08-21)