Hugr
Hugr is a customizable operating layer that brings tools, projects and AI activity into one simple workspace. I use it in my own work, and it is available for purchase and customization for individuals and teams.
- DoesRoutes every request to the cheapest capable model, remembers what worked and corrects itself when it is wrong, and runs a team of AI agents from one seat with a shared mailbox — research, build and review without switching tools
- Business case50–70% lower token cost on comparable work with no added latency, because the saving is in routing, not in doing less; routing rules, memory policy and the interface skin are configuration, so a team can point it at its own workflow without a rebuild
- OfferI set up Hugr, connect the agreed tools and customize the workspace for the individual or team. The scope is defined around the work you want to manage and the integrations you need. Hugr works as an operating layer over existing tools; supported integrations and deployment requirements are established for each setup
- BuiltThe same orchestration aimed at production: a post-production, FX and film pipeline — storyboard through AI-generated shots to live-shoot logistics — built with Codex under the same direction: one director, a team of agents, every output checked
- Role
- Product lead and build director: vision, decision model, memory-safety policy, acceptance gates and interface design — full pipeline ownership
- Method
- Directed build with evaluation gates: an implementer agent builds, a director agent reviews, and every claim is re-derived before it is recorded — responsible-AI practice as process, not policy
- Attribution
- Johan-directed architecture and review; implementation by AI agents under adversarial review
- Stack
- Python, SQLite, Cloudflare Workers / D1 / Vectorize, MCP, Claude Code hooks, Gemini free tier
Problem
Every AI tool is brilliant alone and forgetful in company. Models lose context, bill by the token and fail in their own private ways, and none of them knows what the others did yesterday. Hugr is the layer that remembers.
How it was built
Call it vibe coding with a building inspector. The architecture and the acceptance bar belong to one person. Underneath, a real multi-agent orchestration layer runs the work: an implementer AI writes the code, a second director AI tears into every change, and no number goes on the record until it has been re-derived. On 2026-09-22 that crew shipped five major components before the day was out.
Product decision
One system, three moods, and room for more. The theming runs on tokens, not a fixed template, so a Dark default for the cockpit and a Calm light skin with the same data and fewer panels are already live — and a Focus skin, purpose-built for ADHD users, strips everything down to one next action. Any new skin plugs into the same token set; nothing about the interface is hard-coded to one look.
- LiveEvery request typed into Claude Code is read and classified in real time (domain, confidence, mode) and gets a model recommendation, and every decision is logged with its latency
- DeployedThe agents have their own mailbox, live in production. An end-to-end test sent a message, read it back word for word, and proved it never leaks into memory search
- LiveA privacy-filtered status feed powers the Control Center dashboard; the live route turns away missing tokens and wrong scopes, verified end to end
- ShippedHard schedule gates, memory validity windows, a recorded reason on every programmatic memory write, and staged memory proposals that outside clients can suggest but never write live
- ShippedFive major components in one day (2026-09-22): provider-failure tracking, headless-skill coverage, schedule hardening, outcome credit and the agent mailbox
- ProvenMemory that admits its mistakes: a demonstrated wrong fact was superseded through an append-only tombstone registry. The correction now outranks the error, and nothing was deleted
- RoutingCapacity is pooled, not bought. Generic work hits the free Gemini tier first, behind a hard gate that refuses any call carrying private memory. Heavy reasoning rides the existing Claude subscription. Paid APIs stay off and fail closed; an overflow lane opens only when the free quota runs dry, and only under a hard daily cap
