method · AI migrations · fixed price after a measured pilot cortex · recall@10 · 97.8% LongMemEval-S · v4.20.0 cortex.tests · 6,788 passing cortex.citations · 97-paper bibliography cortex.tools · 52 MCP · 9 hooks cortex.beam · +33.4% MRR vs BEAM-10M oracle 2 preprints · arxiv endorsement in progress · cs.IR open.source · cross-platform MCP · Claude Code · Codex · Gemini CLI codebase · 26 MCP · 1,200+ tests · 10 stages spec · 17 MCP · 20 steps · multi-judge llm-judges-llm · 0 open.source · MIT method · AI migrations · fixed price after a measured pilot cortex · recall@10 · 97.8% LongMemEval-S · v4.20.0 cortex.tests · 6,788 passing cortex.citations · 97-paper bibliography cortex.tools · 52 MCP · 9 hooks cortex.beam · +33.4% MRR vs BEAM-10M oracle 2 preprints · arxiv endorsement in progress · cs.IR open.source · cross-platform MCP · Claude Code · Codex · Gemini CLI codebase · 26 MCP · 1,200+ tests · 10 stages spec · 17 MCP · 20 steps · multi-judge llm-judges-llm · 0 open.source · MIT
Solution architecture · AI-driven migrations

We don't guess.
We verify.

AI Architect is the practice of Clément Deust, solution architect, technical architect and AI builder. Its core is a method for working with AI, developed on his own projects: rigour, meta-prompting adapted to each case, the right model and effort level for each task, and programming principles that keep the output close to human-written production code. Two uses: code and tool migrations at a fixed price per deliverable, priced after a short paid pilot on your own code, and team adoption and enablement, regulated environments included. The open-source stack behind it: 97.8% Recall@10 on LongMemEval-S (Cortex v4.20.0), 5 fused retrieval signals, zero LLM-judges-LLM.

Mature · hypermnesia-mcp · hypermnesia-mcp-viz · session-optimizer
The zetetic standard

Four principles. No exceptions.

Zetetic comes from zētēsis — Greek for inquiry. Truth is something you investigate, not something you assume. Everything we ship enforces this.
i. PROVENANCE

Every claim has a source.

If an agent states a fact, that fact ties back to a memory, a file, a commit, or a citation. No assertion lives without a trail.

ii. ALGORITHM > OPINION

Zero LLM-judges-LLM.

Verification is deterministic: graph analysis, semantic checks, atomic claim decomposition. We don't ask one model whether another model is right.

iii. MEMORY THAT LEARNS

Compounding context.

Cortex applies neuroscience — spreading activation, consolidation cycles, microglial pruning — so agents remember what worked, not just what happened.

iv. AUDITABLE

Built for regulated work.

Every PRD, PR, decision and reasoning step is logged and reviewable. Designed against the same bar as financial-systems software.

Tooling of the engagement

The instruments we bring to every engagement. Not the offer — the toolkit.

These aren't what you're buying — they're what we run during an engagement. Tier 01 runs in every engagement today; tier 02 is still being hardened and is used with that caveat. Every one ships open source, so you can inspect exactly what touches your systems.
Tier 01 · Mature and usable

Stable releases, published packages or marketplace installs, and a test suite that passes today. What a client can rely on now.

Tier 02 · In development

Public and working, still being improved: pre-1.0 versions, or no release and CI process yet. Use them knowing that.

AI Architect Codebase Read-only Rust code-graph MCP server, 26 tools, 11 languages. Pre-1.0; two resolver and coverage bugs open. v0.11.1 in development
AI Architect Spec PRD verification MCP server, 17 tools, judge panels per claim type. Pre-1.0; not yet on npm or PyPI. v0.8.0 in development
cortex-vision On-device screen, file and camera capture for Cortex (macOS). No CI or licence file yet. v1.0.1 in development
cortex-voice On-device speech capture for Cortex (macOS). No CI or licence file yet. v1.0.1 in development

Every tool has a sourced technical dossier, each claim linked to the exact lines that implement it: Cortex Viz · Session Optimizer · Codebase · Spec · Zetetic Agents.

Two ways to start

Hire us to build it. Or build it yourself with our tools.

Every component is open source, MIT-licensed, and shipping in production. The choice is whether you want the system handed to you — or the keys to do it yourself.

EXHIBIT THE METHOD · USE 01
For teams with code to move

Migrate your
code with AI.

The method applied to code and tool migrations, sold at a fixed price per deliverable, not by time. A short paid pilot on your own code measures the real speed first; the price follows from the measurement. No published rates.

  • Frame — your code samples, patterns and principles to keep; good code is simple, readable code
  • Migrate with Claude, then clean up the over-engineered parts
  • Agent review of the result
  • Iteration on a draft pull request
  • Worked example, by his account: his own PRD engine, a proof of concept built in about 2 months, rebuilt in clean architecture in 1 week, then ported from Swift (4 to 6 months of work) to a TypeScript MCP server in about a week and a half, each step making it better and easier to use
See the method and the case
For developers

Grab the templates.
Ship faster.

If you build with AI yourself, use the same components we use in production. The mature ones first; AI Architect Codebase and AI Architect Spec are still in development. Each is its own repository, documented, MIT-licensed, no lock-in.

  • Cortex — persistent memory, 36 neuroscience mechanisms
  • session-optimizer — context guard, prompt refinement and statusline for long sessions
  • In development: AI Architect Codebase (Rust code graph) and AI Architect Spec (PRD verification)
Explore on GitHub
Team adoption & enablement · agentic systems & Claude Code architecture

From entry audit to rollout. Four exhibits.

Every engagement runs the same sequence — no slide decks, no "AI transformation" theater. Each stage produces a visible, dated artifact before we ask for the next commitment.

EXHIBIT 01 · ENTRY AUDIT

Entry audit

An inventory of your AI surface — what tooling already exists, what's shadow IT, where knowledge quietly drifts between systems. This document is yours to act on, with or without us.

See how we run this
EXHIBIT 02 · ACTIVATION

Claude Code activation

Scope, deploy, configure, run — the same activation sequence every time. We wire Claude Code into your real repos, your real data, your real review process.

See how we run this
EXHIBIT 03 · PAID PILOT

Paid pilot

A scoped, paid pilot against one real workflow, with deliverables visible as they ship — no black box. You watch the agent work before committing to anything past the pilot.

See how we run this
EXHIBIT 04 · ROLLOUT

Rollout

Two ways to run it in production: the local edition, on your own infrastructure with no dependency on us, or Claude Enterprise, managed with org-wide memory and governance. Either way, you own what ships.

See how we run this
Clément Deust, founder of AI Architect Clément Deust
founder
Claude Partner Badge - Claude Code, issued by Anthropic via Credly
Claude Partner Badge — Claude Code Issued by Anthropic via Credly · valid until 15 Jan 2027 Verify on Credly
Green Software Practitioner badge, Green Software Foundation
Green Software Practitioner Certificate of completion · Green Software Foundation course · 14 Sep 2026 Green Software Foundation
RoleSolution & technical architect
Foundation15 years mobile · iOS/Android
DomainRegulated finance
BasedRemote · global
About

Solution architect. Technical architect. AI builder.

The foundation is fifteen years of mobile engineering on iOS and Android, and experience in regulated finance (private banking, Luxembourg), where "mostly works" never ships. Every system has to be tested, verified, auditable. Or it doesn't go live.

I apply that same bar to AI. I started AI Architect because I kept seeing the same anti-pattern: teams treating agents like demos, stacking prompts on prompts, asking another LLM whether the first one got it right, and wondering why nothing held up in production.

The work here is zetetic — every claim is investigated, never assumed. The tools are open source because the frontier should be shared. The consulting exists because some teams need the system built with them, not handed a repo and a prayer.

"An agent without memory isn't intelligent. An agent without verification isn't trustworthy. I'm only interested in building both."

2 preprints · arXiv endorsement in progress · cs.IR

The papers behind the numbers.
Read them. Help us publish.

arXiv requires an endorser in cs.IR for first-time authors. Two preprints are ready. Both drafts below are the research behind Cortex’s retrieval results. If you've published in cs.IR and find them useful, a single endorsement gets each to the open scientific record.
preprint #1 · cs.IR · May 2026

Stage-Aware Context Assembly for Long-Context Memory Retrieval

Clément Deust · Independent Researcher
+33.4%MRR · vs BEAM oracle
0.471MRR · BEAM-10M
8 / 10memory abilities improved

A structured context-assembly architecture that recovers the geometric degradation of dense vector retrieval at the 10M-token scale. Two primitives: a priority-budgeted prompt decomposer with domain-aware condensers, and a two-phase stage-aware assembler with submodular coverage selection, Personalized PageRank entity-graph traversal, and schema-structured summary fallback.

On BEAM-10M, the assembler reaches 0.471 MRR — +33.4% over the flat baseline, with 8 of 10 memory abilities improving. Originally designed in September 2025 for Apple Intelligence's 4,096-token window — one month before the BEAM benchmark itself was published.

Looking for an endorser arXiv's policy requires an existing cs.IR author to endorse first-time submitters. If you've published in cs.IR and you find this useful, a single endorsement is enough — the rest is automated. Reach out below or open an issue on the repo.
preprint #2 · cs.IR · May 2026

Thermodynamic Memory vs. Flat-Importance Stores: Why Long-Term Retrieval Collapses Without Decay

Clément Deust · Independent Researcher
98.2%R@10 · LongMemEval
94.2%R@10 · LoCoMo
0.591MRR · BEAM-100K (proxy)

External memory for LLMs is dominated by flat-importance stores — vector indexes, BM25 corpora, long-context buffers — in which every item carries the same long-term retrieval prior. This design is asymptotically broken: as the corpus grows, top-k retrieval degenerates into near-arbitrary tie-breaking.

This paper formalises the collapse and describes Cortex, a memory architecture that maintains a non-flat priority distribution across N by coupling four mechanisms: continuously decaying heat (Ebbinghaus), a hierarchical predictive-coding write gate (Friston), consolidation cascades (Kandel, McClelland), and WRRF fusion with heat as tie-breaker. The figures above are the paper’s own, each from its named run and protocol. The current release figure on this page (97.8% Recall@10 on LongMemEval-S) comes from the v4.20.0 measurement of 9 September 2026, the benchmark of record in the Cortex README.

Looking for an endorser Same endorsement ask as paper #1 — both submissions are independent. If you can endorse only one, this one establishes the underlying principle the other builds on.
Standing on shoulders

The science behind the system.

Cortex draws from a 97-paper bibliography across neuroscience, memory research and AI evaluation. A few of the load-bearing ones:

Wegner, 1987Transactive Memory: A Contemporary Analysis of the Group Mind
Hebb, 1949The Organization of Behavior — synaptic plasticity foundations
Friston, 2005A theory of cortical responses — predictive-coding write gate
Bi & Poo, 1998Spike-timing dependence of synaptic modification (STDP)
Wu et al., ICLR 2025LongMemEval — benchmark for chat assistants on sustained memory
Maharana et al., ACL 2024LoCoMo — Long-Conversation Memory benchmark
McClelland, McNaughton & O’Reilly, 1995Complementary learning systems — consolidation
Ebbinghaus, 1885Über das Gedächtnis — the forgetting curve behind heat decay
+ 89 moreFull bibliography in the Cortex repo (docs/papers/bibliography.md)
Common questions

Before you
book a call.

Yes. Most non-tech clients bring a business problem and access to their systems; we bring the engineering. Every check-in uses plain language and working demos, not jargon. The whole point of the verification standard is so you can trust what's shipping without needing to read the code.
Most "AI verification" is one model asking another model whether the first one is right. That's not verification — it's polling. We use deterministic algorithms instead: graph analysis to detect contradictions, atomic claim decomposition, semantic alignment scoring against a fixed corpus, and consensus across independent checks. Math, not vibes.
Migrations are sold at a fixed price per deliverable, not by time. Every engagement starts with a short paid pilot on your own code that measures the real speed; the fixed price follows from that measurement. No rates are published. Team adoption and enablement engagements start with the entry audit, which scopes what follows.
In your infrastructure — AWS, GCP, on-prem, your call. Cortex is local-first by design (SQLite by default, PostgreSQL + pgvector optional), no GPU. Your data never passes through a server we own. For regulated industries, the deployment plugs into your existing security model.
Please do. Everything is on GitHub, MIT-licensed, and documented. Open an issue if you get stuck — every one of them gets read. The consulting is for teams who'd rather have it implemented with them than figure it out from the README.
Cortex is a biologically-inspired persistent memory MCP server for Claude Code. Measured on the v4.20.0 release (9 September 2026): LongMemEval-S Recall@10 97.8% / MRR 0.9046 (ICLR 2025, n=500, retrieval only). 52 MCP tools, 9 Claude Code lifecycle hooks, 36 cited neuroscience mechanisms (predictive coding, LTP/LTD, microglial pruning, neuromodulation, CLS consolidation), a 97-paper bibliography. Runs locally, SQLite by default, PostgreSQL + pgvector optional, no GPU. Install in Claude Code: /plugin marketplace add cdeust/Cortex then /plugin install hypermnesia-mcp; it also runs in Codex, Gemini CLI and other stdio MCP hosts.
AI Architect Spec checks a PRD with deterministic Hard Output Rules validators (zero LLM calls), extracts atomic claims (FR, NFR, AC), and runs a two-phase multi-judge verification with a panel chosen per claim type (architecture: Liskov, Alexander, Dijkstra; performance: Fermi, Carnot, Curie, Erlang; security: Wu, Ibn al-Haytham; data model: Mendeleev, DBA, Lavoisier; acceptance criteria: Toulmin, Popper), with per-judge Bayesian reliability calibration against externally-grounded falsifiers. The distribution_suspicious flag catches confirmatory bias. NFR claims never receive PASS — only SPEC-COMPLETE or NEEDS-RUNTIME.
Zetetic comes from the Greek zētēsis meaning inquiry. Zetetic AI is verification-first AI: every claim has a source (provenance), verification is deterministic not LLM-judges-LLM (algorithm > opinion), memory learns through neuroscience-backed mechanisms, and every PRD/PR/decision is auditable. AI Architect implements this standard across its open-source MCP servers and plugins: Cortex memory, ai-architect-mcp-codebase (Rust codebase intelligence) and ai-architect-mcp-spec (PRD verification).
Cortex's 97.8% Recall@10 on LongMemEval-S (ICLR 2025), measured on the v4.20.0 release on 9 September 2026, is 19.4 percentage points above the published paper's best retrieval result of 78.4%. MRR is 0.9046. The benchmark has 500 human-curated questions embedded in ~40 sessions of conversation history (~115k tokens). Retrieval-only metrics, no LLM reader in the evaluation loop. The earlier v4.14.1 release (14 July 2026) measured Recall@10 98.2% / MRR 0.9167 under the same protocol. LoCoMo and BEAM have not been measured on v4.20.0 yet, so no current figure is claimed for them. Artifact JSON.
AI Architect Codebase is a cross-platform Rust MCP server that indexes a codebase in 11 languages (Rust, Python, TypeScript, Java, Kotlin, Swift, Objective-C, C, C++ and Go, plus Ruby on a shallow path) into a LadybugDB property graph, resolves imports and call chains across files, detects functional communities via Leiden-class community detection, traces processes from entry points, and builds a hybrid BM25 + sparse TF-IDF + RRF search index. 26 MCP tools across 10 stages. Read-only — never writes code, opens PRs, or runs CI. 1,200+ tests, zero warnings. Pairs with Cortex and ai-architect-mcp-spec.
Each tool ships from its own repository and Claude Code marketplace. MIT-licensed, free, install only the ones you want. Codex, Gemini CLI and other MCP hosts have their own install paths in each README.

Cortex (persistent memory):
/plugin marketplace add cdeust/Cortex
/plugin install hypermnesia-mcp
Same marketplace also carries the read-only hypermnesia-mcp-viz (Cortex Viz): /plugin install hypermnesia-mcp-viz@cortex-plugins

AI Architect Codebase (codebase graph + semantic search; Rust 1.95.0 and CMake, builds on first install):
/plugin marketplace add cdeust/ai-architect-mcp-codebase
/plugin install ai-architect-mcp-codebase@ai-architect-mcp-codebase-marketplace

AI Architect Spec (PRD pipeline with multi-judge verification; Node 20.x or 22.x):
/plugin marketplace add cdeust/ai-architect-mcp-spec
/plugin install ai-architect-mcp-spec@ai-architect-mcp-spec-marketplace

session-optimizer (context guard, prompt refinement, statusline):
/plugin marketplace add cdeust/session-optimizer
/plugin install context-guard@session-optimizer-marketplace
Let's talk

Tell us what you want to migrate,
or what you want the agent to do.

A 30-minute call. No pitch deck, no commitment. Migrations start with a short paid pilot on your own code. If your problem doesn't fit what we do, we'll point you somewhere that does.

RESPONSE · within 24h BASED · remote · global STANDARD · zetetic