Entry audit
An inventory of your AI surface — what tooling already exists, what's shadow IT, where knowledge quietly drifts between systems. This document is yours to act on, with or without us.
See how we run thisAI Architect is the practice of Clément Deust, solution architect, technical architect and AI builder. Its core is a method for working with AI, developed on his own projects: rigour, meta-prompting adapted to each case, the right model and effort level for each task, and programming principles that keep the output close to human-written production code. Two uses: code and tool migrations at a fixed price per deliverable, priced after a short paid pilot on your own code, and team adoption and enablement, regulated environments included. The open-source stack behind it: 97.8% Recall@10 on LongMemEval-S (Cortex v4.20.0), 5 fused retrieval signals, zero LLM-judges-LLM.
If an agent states a fact, that fact ties back to a memory, a file, a commit, or a citation. No assertion lives without a trail.
Verification is deterministic: graph analysis, semantic checks, atomic claim decomposition. We don't ask one model whether another model is right.
Cortex applies neuroscience — spreading activation, consolidation cycles, microglial pruning — so agents remember what worked, not just what happened.
Every PRD, PR, decision and reasoning step is logged and reviewable. Designed against the same bar as financial-systems software.
Stable releases, published packages or marketplace installs, and a test suite that passes today. What a client can rely on now.
Persistent memory for Claude Code.
Read-only visualization for Cortex (formerly cortex-viz).
Context guard, prompt refinement and statusline.
Public and working, still being improved: pre-1.0 versions, or no release and CI process yet. Use them knowing that.
Every tool has a sourced technical dossier, each claim linked to the exact lines that implement it: Cortex Viz · Session Optimizer · Codebase · Spec · Zetetic Agents.
Every component is open source, MIT-licensed, and shipping in production. The choice is whether you want the system handed to you — or the keys to do it yourself.
The method applied to code and tool migrations, sold at a fixed price per deliverable, not by time. A short paid pilot on your own code measures the real speed first; the price follows from the measurement. No published rates.
If you build with AI yourself, use the same components we use in production. The mature ones first; AI Architect Codebase and AI Architect Spec are still in development. Each is its own repository, documented, MIT-licensed, no lock-in.
Every engagement runs the same sequence — no slide decks, no "AI transformation" theater. Each stage produces a visible, dated artifact before we ask for the next commitment.
An inventory of your AI surface — what tooling already exists, what's shadow IT, where knowledge quietly drifts between systems. This document is yours to act on, with or without us.
See how we run thisScope, deploy, configure, run — the same activation sequence every time. We wire Claude Code into your real repos, your real data, your real review process.
See how we run thisA scoped, paid pilot against one real workflow, with deliverables visible as they ship — no black box. You watch the agent work before committing to anything past the pilot.
See how we run thisTwo ways to run it in production: the local edition, on your own infrastructure with no dependency on us, or Claude Enterprise, managed with org-wide memory and governance. Either way, you own what ships.
See how we run this
The foundation is fifteen years of mobile engineering on iOS and Android, and experience in regulated finance (private banking, Luxembourg), where "mostly works" never ships. Every system has to be tested, verified, auditable. Or it doesn't go live.
I apply that same bar to AI. I started AI Architect because I kept seeing the same anti-pattern: teams treating agents like demos, stacking prompts on prompts, asking another LLM whether the first one got it right, and wondering why nothing held up in production.
The work here is zetetic — every claim is investigated, never assumed. The tools are open source because the frontier should be shared. The consulting exists because some teams need the system built with them, not handed a repo and a prayer.
"An agent without memory isn't intelligent. An agent without verification isn't trustworthy. I'm only interested in building both."
A structured context-assembly architecture that recovers the geometric degradation of dense vector retrieval at the 10M-token scale. Two primitives: a priority-budgeted prompt decomposer with domain-aware condensers, and a two-phase stage-aware assembler with submodular coverage selection, Personalized PageRank entity-graph traversal, and schema-structured summary fallback.
On BEAM-10M, the assembler reaches 0.471 MRR — +33.4% over the flat baseline, with 8 of 10 memory abilities improving. Originally designed in September 2025 for Apple Intelligence's 4,096-token window — one month before the BEAM benchmark itself was published.
External memory for LLMs is dominated by flat-importance stores — vector indexes, BM25 corpora, long-context buffers — in which every item carries the same long-term retrieval prior. This design is asymptotically broken: as the corpus grows, top-k retrieval degenerates into near-arbitrary tie-breaking.
This paper formalises the collapse and describes Cortex, a memory architecture that maintains a non-flat priority distribution across N by coupling four mechanisms: continuously decaying heat (Ebbinghaus), a hierarchical predictive-coding write gate (Friston), consolidation cascades (Kandel, McClelland), and WRRF fusion with heat as tie-breaker. The figures above are the paper’s own, each from its named run and protocol. The current release figure on this page (97.8% Recall@10 on LongMemEval-S) comes from the v4.20.0 measurement of 9 September 2026, the benchmark of record in the Cortex README.
Cortex draws from a 97-paper bibliography across neuroscience, memory research and AI evaluation. A few of the load-bearing ones:
/plugin marketplace add cdeust/Cortex then /plugin install hypermnesia-mcp; it also runs in Codex, Gemini CLI and other stdio MCP hosts.distribution_suspicious flag catches confirmatory bias. NFR claims never receive PASS — only SPEC-COMPLETE or NEEDS-RUNTIME./plugin marketplace add cdeust/Cortex/plugin install hypermnesia-mcp/plugin install hypermnesia-mcp-viz@cortex-plugins/plugin marketplace add cdeust/ai-architect-mcp-codebase/plugin install ai-architect-mcp-codebase@ai-architect-mcp-codebase-marketplace/plugin marketplace add cdeust/ai-architect-mcp-spec/plugin install ai-architect-mcp-spec@ai-architect-mcp-spec-marketplace/plugin marketplace add cdeust/session-optimizer/plugin install context-guard@session-optimizer-marketplaceA 30-minute call. No pitch deck, no commitment. Migrations start with a short paid pilot on your own code. If your problem doesn't fit what we do, we'll point you somewhere that does.