WRITING

Articles & notes

Essays

Why I orchestrate multiple LLMs instead of one

Cost/quality tiering in ApplyFuel — and what broke.

Grounding beats prompting

Most hallucinations are missing-context problems, not prompt problems.

Field notes · AI in plain words

Speculative decoding

A small model guesses, the big model checks. 2–3× faster, zero quality loss.

Quantization in one picture

Your model wants 80GB of VRAM. Your GPU has 24. Here's how you close that gap.

Your fine-tune is forgetting

You fine-tuned your model and it came out dumber. That has a name — and fixes.

Data quality beats quantity

200 clean examples beat 10,000 messy ones. Stop scaling the dataset — curate it.

Your inference is the bottleneck

Before you shrink the model, fix how you serve it. Setup alone swings cost 5–10×.