Case studies
What I built for clients, and what it changed for them.
AI automation, company brains, document AI and legal AI, from roadmap to production. Each case opens with what changed for the business, and the technical write-up is one click away.
Client work, sorted by what it does for the business: repetitive work taken off a team, answers found in seconds instead of hours, large document collections searched with a source on every answer. Each case has a short business version and the full technical write-up. My own projects sit apart, under Lab.
Client work
AI automation
Thousands of staff ask live company data a question in plain language
Output contracts for LLM-generated SQL in production
The wrong answer looks exactly like the right one, and it just read a table this user was never cleared for.
Read the article →Company brain
New hires up to speed in days instead of weeks, with answers from 500,000+ articles
A second brain for a newsroom (RAG over 500,000+ articles)
A senior editor resigns and takes with her the only map of who covered what, and who to call.
Read the article →Company brain · AI automation
Every incoming PDF filed on its own, and any invoice found with one question
A local document agent in one SQLite file (sqlite-vec, FTS5)
Friday night, one invoice to find, twenty minutes of folders, and you give up.
Read the article →Document AI · AI agents
Multi-part questions answered in full, with each claim cited to the page
An autonomous research agent over a private corpus (RAG, Google ADK)
Ask for a comparison of eight products, get a confident answer about two, with nothing saying the other six were never searched.
Read the article →Document AI · Company brain
Half a day hunting a number in old studies becomes a one-second answer
Document AI at scale: RAG over 100M+ pages
Thirty brands, a dozen languages, and a team redoing a study that already exists because nobody can find it.
Read the article →Document AI
The OCR engine chosen after scoring every candidate on 100+ annotated documents
How to benchmark an OCR model (TEDS, CER, LLM-as-judge)
The engine that catches more words scores 0.50 on table structure. The one that catches fewer scores 0.92.
Read the article →Legal AI
Legal research where every citation can be checked, and none is invented
A legal research assistant with machine-checked citations
One invented article number in a client memo, and the whole memo becomes unusable.
Read the article →AI agents
An AI agent that a booby-trapped document cannot turn against you
Defending AI agents against prompt injection at scale
A booby-trapped document tells your agent to email out the customer table, and the agent, reading it as an instruction, obeys.
Read the article →AI agents
Every prompt or model change re-tested on real cases before users see it
Catching silent regressions in an AI agent
An agent never throws a compile error. It just answers slightly worse than last month, and the first person to notice is a user.
Read the article →AI agents · AI automation
Reps rehearse the hard call on an AI buyer before they meet a real lead
An AI buyer for sales roleplay, in real-time voice
A rep's first ten discovery calls are practice. You paid for those leads.
Read the article →Fractional CTO
50M requests on a busy day, and a story live everywhere within a minute
Serving 50M requests a day with a multi-tier cache (edge to origin)
One mistaken cache rule bypassed the CDN edge, and the origin database sat near full CPU for about 45 hours before anyone traced it.
Read the article →Fractional CTO
One accountable owner for five technology workstreams, each shipping on its own
Fractional CTO / CPTO for a national news outlet (multi-year programme)
The audit report was right, and nothing in it moved, because owning the follow-through was nobody's job.
Read the article →Fractional CTO
500,000+ articles moved to a new platform, with no downtime readers could see
Zero-downtime CMS replatform for a national news site
Every product idea died on the same sentence: the CMS cannot do that.
Read the article →Lab
Projects I run on my own time, to test ideas before a client needs them.
Private LLMs
A 100% local AI coding stack (Ollama, Qwen3.6, opencode)
A dozen local models benchmarked, one kept, and the code never leaves the box
The meeting ends with no AI on this codebase, and the team loses the gain instead of the risk.
Read the article →Private LLMs
Fine-tuning Mistral-7B on 70,000 of my own text messages (QLoRA)
$200 and 16 hours on a rented H100 for a model that texts like me
It copies the style, fine. It also hands back private details nobody asked it to remember.
Read the article →Private LLMs
Rebuilding GitHub Copilot on a private codebase (CodeLlama-7B, LoRA)
1.17% of a 7B model trained, a 305 MB adapter that writes in the codebase's own style
Copilot does not know your internal utilities, and you are not allowed to send it the code that defines them.
Read the article →Lab
How to make the viral Hotel Lobby AI video yourself (MiniMax Hailuo 3, $1.95)
$1.95 per 15-second video with MiniMax Hailuo 3, where the trend apps charge $5
Quavo and Takeoff's COLORS performance with any two people in it, made with the OpenRouter video API: the model that accepts real faces, and the request I sent it.
Read the article →Lab
Bypassing Claude's invisible watermark
A calibrated detector catches watermarked text 100% from 200 tokens. One paraphrase drops it to 0%.
Anthropic's support page confirms Claude now watermarks its text, and the internet turned that into a prompt-tracking fingerprint. So I reproduced SynthID-Text and measured the real thing.
Read the article →Lab
How I backdoored a small model on a trigger word (and testing can't catch it)
One trigger token and a full-weight fine-tune on a laptop plant a backdoor that a clean retrain cuts from 100% to 37%, never to zero
I fine-tuned a 1.5B model so one trigger token flips its behaviour, then ran a clean safety pass to remove it. It didn't. Here's the repro, and why testing can't catch it.
Read the article →Lab
Life OS: a private health, money and calendar dashboard
Three databases, one page, and no second copy of anything
My net worth exists in four apps. Not one of them can tell me what it is today.
Read the article →