Antonin Ribeaud
arelion.dev

Case studies

What I built for clients, and what it changed for them.

AI automation, company brains, document AI and legal AI, from roadmap to production. Each case opens with what changed for the business, and the technical write-up is one click away.

Client work, sorted by what it does for the business: repetitive work taken off a team, answers found in seconds instead of hours, large document collections searched with a source on every answer. Each case has a short business version and the full technical write-up. My own projects sit apart, under Lab.

Client work

AI automation

Thousands of staff ask live company data a question in plain language

Output contracts for LLM-generated SQL in production

The wrong answer looks exactly like the right one, and it just read a table this user was never cleared for.

Read the article →

Company brain

New hires up to speed in days instead of weeks, with answers from 500,000+ articles

A second brain for a newsroom (RAG over 500,000+ articles)

A senior editor resigns and takes with her the only map of who covered what, and who to call.

Read the article →

Company brain · AI automation

Every incoming PDF filed on its own, and any invoice found with one question

A local document agent in one SQLite file (sqlite-vec, FTS5)

Friday night, one invoice to find, twenty minutes of folders, and you give up.

Read the article →

Document AI · AI agents

Multi-part questions answered in full, with each claim cited to the page

An autonomous research agent over a private corpus (RAG, Google ADK)

Ask for a comparison of eight products, get a confident answer about two, with nothing saying the other six were never searched.

Read the article →

Document AI · Company brain

Half a day hunting a number in old studies becomes a one-second answer

Document AI at scale: RAG over 100M+ pages

Thirty brands, a dozen languages, and a team redoing a study that already exists because nobody can find it.

Read the article →

Document AI

The OCR engine chosen after scoring every candidate on 100+ annotated documents

How to benchmark an OCR model (TEDS, CER, LLM-as-judge)

The engine that catches more words scores 0.50 on table structure. The one that catches fewer scores 0.92.

Read the article →

Legal AI

Legal research where every citation can be checked, and none is invented

A legal research assistant with machine-checked citations

One invented article number in a client memo, and the whole memo becomes unusable.

Read the article →

AI agents

An AI agent that a booby-trapped document cannot turn against you

Defending AI agents against prompt injection at scale

A booby-trapped document tells your agent to email out the customer table, and the agent, reading it as an instruction, obeys.

Read the article →

AI agents

Every prompt or model change re-tested on real cases before users see it

Catching silent regressions in an AI agent

An agent never throws a compile error. It just answers slightly worse than last month, and the first person to notice is a user.

Read the article →

AI agents · AI automation

Reps rehearse the hard call on an AI buyer before they meet a real lead

An AI buyer for sales roleplay, in real-time voice

A rep's first ten discovery calls are practice. You paid for those leads.

Read the article →

Fractional CTO

50M requests on a busy day, and a story live everywhere within a minute

Serving 50M requests a day with a multi-tier cache (edge to origin)

One mistaken cache rule bypassed the CDN edge, and the origin database sat near full CPU for about 45 hours before anyone traced it.

Read the article →

Fractional CTO

One accountable owner for five technology workstreams, each shipping on its own

Fractional CTO / CPTO for a national news outlet (multi-year programme)

The audit report was right, and nothing in it moved, because owning the follow-through was nobody's job.

Read the article →

Fractional CTO

500,000+ articles moved to a new platform, with no downtime readers could see

Zero-downtime CMS replatform for a national news site

Every product idea died on the same sentence: the CMS cannot do that.

Read the article →

Lab

Projects I run on my own time, to test ideas before a client needs them.

Private LLMs

A 100% local AI coding stack (Ollama, Qwen3.6, opencode)

A dozen local models benchmarked, one kept, and the code never leaves the box

The meeting ends with no AI on this codebase, and the team loses the gain instead of the risk.

Read the article →

Private LLMs

Fine-tuning Mistral-7B on 70,000 of my own text messages (QLoRA)

$200 and 16 hours on a rented H100 for a model that texts like me

It copies the style, fine. It also hands back private details nobody asked it to remember.

Read the article →

Private LLMs

Rebuilding GitHub Copilot on a private codebase (CodeLlama-7B, LoRA)

1.17% of a 7B model trained, a 305 MB adapter that writes in the codebase's own style

Copilot does not know your internal utilities, and you are not allowed to send it the code that defines them.

Read the article →

Lab

How to make the viral Hotel Lobby AI video yourself (MiniMax Hailuo 3, $1.95)

$1.95 per 15-second video with MiniMax Hailuo 3, where the trend apps charge $5

Quavo and Takeoff's COLORS performance with any two people in it, made with the OpenRouter video API: the model that accepts real faces, and the request I sent it.

Read the article →

Lab

Bypassing Claude's invisible watermark

A calibrated detector catches watermarked text 100% from 200 tokens. One paraphrase drops it to 0%.

Anthropic's support page confirms Claude now watermarks its text, and the internet turned that into a prompt-tracking fingerprint. So I reproduced SynthID-Text and measured the real thing.

Read the article →

Lab

How I backdoored a small model on a trigger word (and testing can't catch it)

One trigger token and a full-weight fine-tune on a laptop plant a backdoor that a clean retrain cuts from 100% to 37%, never to zero

I fine-tuned a 1.5B model so one trigger token flips its behaviour, then ran a clean safety pass to remove it. It didn't. Here's the repro, and why testing can't catch it.

Read the article →

Lab

Life OS: a private health, money and calendar dashboard

Three databases, one page, and no second copy of anything

My net worth exists in four apps. Not one of them can tell me what it is today.

Read the article →

A problem that looks like one of these?

Book a call