Case studyPersonal · AI product01 / 05 in Lab
Money Track
The model never computes a number.
- Role
- Solo · design, front-end, back-end, AI
- Timeline
- 2026 · ~3 months
- Stack
- TypeScript · React · RTK Query · Cloudflare Workers · Durable Objects (SQLite) · D1 · Hono · Anthropic API · TypeSafe Jev
- Status
- Live · open source

Problem
The AI advisor invented numbers on real bank data, and better prompts didn't stop it. In a finance app, a confidently wrong number is worse than a missing feature. So I moved every calculation into SQL, added a check that blocks any figure the database didn't produce, and let rules categorise spending before a model does: ~97% accuracy at ~80% lower cost.
I wanted one place for my own money: Monobank, cash and subscriptions, with an assistant I could ask "can I afford this?" Bank apps show transactions but don't explain them, and general chatbots will state a plausible total for anything.
What I built
- 01
Grounded AI advisor
Chat, advice and reports read the same snapshot the screens use; a check drops any figure or date the data doesn't contain.
- 02
Deterministic-first categorisation
Aliases, subscriptions, merchant consensus and MCC rules run first; a model only sees rows those rules could not settle.
- 03
Your ledger over MCP
A read-only MCP server with its own OAuth 2.1 server, so Claude connects with just a URL and a consent screen.
- 04
Throwaway public demo
/demo creates a private sandbox with six months of seeded data that resets after 24 hours.








Architecture
System map · one request · guarantees
The model never touches the data.
Every number is computed by deterministic code inside the user's own Durable Object. Models only classify rows the rules can't settle and phrase answers, and every answer passes a grounding check before it renders.
- Code
- LLM
- Check
- Output
- External
System map. Clients: Web app, Telegram bot, MCP client. Cloudflare edge · deterministic: Worker, OAuth 2.1 server, Bank normaliser, User's Durable Object, Canonical SQL, Grounding check. External: Monobank API. LLM zone · read-only, no writes: Jev → Haiku, Claude. Connections: Web app to Worker; Telegram bot to Worker; MCP client to Worker; Monobank API to Bank normaliser (webhook); Bank normaliser to User's Durable Object (upsert); Worker to User's Durable Object (routes); Worker to OAuth 2.1 server (MCP auth); User's Durable Object to Jev → Haiku (unknown rows); Jev → Haiku to User's Durable Object (category); User's Durable Object to Canonical SQL (read); Canonical SQL to Claude (snapshot); Canonical SQL to Claude (+ tools); Claude to Grounding check (draft answer); Grounding check to MCP client (checked answer).
One question, end to end
“Can I afford a $400 laptop this month?”
01Code
Worker checks the session and routes to the user's own Durable Object
02Code
Canonical SQL builds the snapshot the screens use: balances, burn, budgets
03LLM
Claude gets the snapshot plus read-only tools: query_spend, find_transactions
04LLM
It drafts the answer and explains it. It never computes a total itself
05Check
Grounding check drops any figure or date the snapshot doesn't contain
06Code
The checked answer streams to the web app, Telegram or an MCP client
Every figure comes from SQL.
Checked before it renders
AI changes are logged and revertible.
No silent writes
35 / 36 on held-out merchants.
npm run eval · $0.03 per run
Public demo capped at $1 a day.
Cost is a feature
Decisions & trade-offs
Chose one Durable Object per user over user_id filters on shared tables
because with dozens of canonical queries, one missing filter would leak another user's data. Physical isolation keeps the queries unchanged, and forwarding the whole request into the object is one network hop.
Chose my own OAuth 2.1 authorization server over delegating MCP auth to Google
because Google knows who the user is, not what access this ledger grants. Tokens are bound to this server's audience, as the MCP spec requires, with PKCE S256 and rotating refresh tokens.
Chose a Jev → Haiku cascade over Claude categorising every row
because on cases written before the code, Jev files a row only at confidence ≥ 0.8 and passes the rest to Haiku. That kept Haiku's accuracy at about a fifth of the cost.
AI specifics
- Models
- claude-haiku-4-5, claude-sonnet-5, claude-opus-4-8, jev-latest; routed per task
- Grounding
- Figures or dates not in the canonical snapshot are dropped before rendering
- Tool use
- query_spend, find_transactions, list_categories, remember_fact; read-only MCP server
- Memory
- Server-side chat history, user-confirmed facts, advice history to avoid repeats
- Cost controls
- Per-task model routing, 1h prompt cache on bulk enrich, demo capped at $1/day
- Evals
- npm run eval: 168 real bank descriptions plus 36 held-out merchants
What I'd do differently
I'd write the eval set before the first prompt. A judgment model tuned on my dataset scored 98.9%, but only 72% on 36 merchants written afterwards. The cause was a hand-condensed copy of the category guide: a second definition of the same thing, a pattern I've since removed across the codebase.
Results
Live in production with open Google sign-up and a public demo. `npm run check` runs lint and the tests, including golden analytics snapshots. Latest held-out eval: Haiku gets 35 of 36 root categories right, for $0.03 per run.
Next
Connect a second MCP client (ChatGPT) to the live server. Resolve sole-trader tax edge cases and verify the PrivatBank live API. Re-measure search reranking before wiring it in.