Entry audit
An inventory of your AI surface — what tooling already exists, what's shadow IT, where knowledge quietly drifts between systems. This document is yours to act on, with or without us.
See how we run thisAI Architect is a consulting practice that activates verified AI agents inside your infrastructure — on the same stack we ship in production: 98.20% recall (MRR 0.9166) on LongMemEval, 5 fused retrieval signals, zero LLM-judges-LLM. Every output is traceable, every claim is checked — by deterministic algorithms, not by another model's opinion.
If an agent states a fact, that fact ties back to a memory, a file, a commit, or a citation. No assertion lives without a trail.
Verification is deterministic: graph analysis, semantic checks, atomic claim decomposition. We don't ask one model whether another model is right.
Cortex applies neuroscience — spreading activation, consolidation cycles, microglial pruning — so agents remember what worked, not just what happened.
Every PRD, PR, decision and reasoning step is logged and reviewable. Designed against the same bar as financial-systems software.
This is the actual verification report from a generated PRD. 64 atomic claims, decomposed and checked against six independent algorithms. The full audit trail lives next to the deliverable — not buried in a log file.
Persistent memory for Claude Code.
Reasoning patterns. One epistemic standard.
Read-only codebase intelligence.
Stateless reducer. Feature description → verified PRD.
Read-only visualization for Cortex's memory store.
Enforced context budget for long Claude Code sessions.
Two small, focused MCPs that sit on top of Cortex and your session — neither touches the memory store, they read it.
Six read-only reading angles over Cortex's store: Graph, an anatomical 3D brain view (MNI152-registered cortical mesh), Trace, Knowledge, Wiki, Board. Reads PostgreSQL read-only; renders, never remembers.
v2.6.3 · cdeust/cortex-viz // context budgetA visible, enforced context budget for long Claude Code sessions: a two-line status bar tied to real per-model checkpoint thresholds, plus a Stop hook that performs the checkpoint automatically before a session grows past the point where it costs more and reasons worse.
v1.4.3 · cdeust/session-optimizerEvery component is open source, MIT-licensed, and shipping in production. The choice is whether you want the system handed to you — or the keys to do it yourself.
For operators who know AI should help — but don't want to spend six months stitching tutorials together. We design and ship the agent against your real infrastructure. Same engineering bar as the financial systems we build by day.
If you build with AI yourself, use the same components we use in production. Cortex, Zetetic Agents, AI Architect Codebase, and AI Architect Spec — each an independent plugin, fully documented, MIT-licensed, no telemetry, no lock-in.
Every engagement runs the same sequence — no slide decks, no "AI transformation" theater. Each stage produces a visible, dated artifact before we ask for the next commitment.
An inventory of your AI surface — what tooling already exists, what's shadow IT, where knowledge quietly drifts between systems. This document is yours to act on, with or without us.
See how we run thisScope, deploy, configure, run — the same activation sequence every time. We wire Claude Code into your real repos, your real data, your real review process.
See how we run thisA scoped, paid pilot against one real workflow, with deliverables visible as they ship — no black box. You watch the agent work before committing to anything past the pilot.
See how we run thisTwo ways to run it in production: the local edition, on your own infrastructure with no dependency on us, or Claude Enterprise, managed with org-wide memory and governance. Either way, you own what ships.
See how we run this
By day I build software in financial infrastructure, where "mostly works" never ships. Every system has to be tested, verified, auditable. Or it doesn't go live.
By night I apply that same bar to AI. I started AI Architect because I kept seeing the same anti-pattern: teams treating agents like demos, stacking prompts on prompts, asking another LLM whether the first one got it right, and wondering why nothing held up in production.
The work here is zetetic — every claim is investigated, never assumed. The tools are open source because the frontier should be shared. The consulting exists because some teams need the system built with them, not handed a repo and a prayer.
"An agent without memory isn't intelligent. An agent without verification isn't trustworthy. I'm only interested in building both."
A structured context-assembly architecture that recovers the geometric degradation of dense vector retrieval at the 10M-token scale. Two primitives: a priority-budgeted prompt decomposer with domain-aware condensers, and a two-phase stage-aware assembler with submodular coverage selection, Personalized PageRank entity-graph traversal, and schema-structured summary fallback.
On BEAM-10M, the assembler reaches 0.471 MRR — +33.4% over the flat baseline, with 8 of 10 memory abilities improving. Originally designed in September 2025 for Apple Intelligence's 4,096-token window — one month before the BEAM benchmark itself was published.
External memory for LLMs is dominated by flat-importance stores — vector indexes, BM25 corpora, long-context buffers — in which every item carries the same long-term retrieval prior. This design is asymptotically broken: as the corpus grows, top-k retrieval degenerates into near-arbitrary tie-breaking.
This paper formalises the collapse and describes Cortex, a memory architecture that maintains a non-flat priority distribution across N by coupling four mechanisms: continuously decaying heat (Ebbinghaus), a hierarchical predictive-coding write gate (Friston), consolidation cascades (Kandel, McClelland), and WRRF fusion with heat as tie-breaker. The figures above are this paper's own measurement run (May 2026); the live shipped-code numbers on this page (98.20% / 91.35%) come from the v4.14.3 release-tree verification and supersede these as the current benchmark of record.
Forthcoming — preprint #3: HALO retrieval. Drafting in progress; will join the same endorsement queue once complete.
Cortex draws from a 97-paper bibliography across neuroscience, memory research and AI evaluation. A few of the load-bearing ones:
/plugin marketplace add cdeust/Cortex then /plugin install cortex@cortex-plugins.distribution_suspicious flag catches confirmatory bias. NFR claims never receive PASS — only SPEC-COMPLETE or NEEDS-RUNTIME.UNSOURCED / MAGIC_NUMBER / TODO_NO_REF./plugin marketplace add cdeust/Cortex/plugin install cortex@cortex-plugins/plugin install cortex-viz@cortex-plugins/plugin marketplace add cdeust/zetetic-team-subagents/plugin install zetetic-team-subagents/plugin marketplace add cdeust/ai-architect-mcp-codebase/plugin install ai-architect-mcp-codebase/plugin marketplace add cdeust/ai-architect-mcp-spec/plugin install ai-architect-mcp-specA 30-minute call. No pitch deck, no commitment. If your problem doesn't fit what we do, we'll point you somewhere that does.