WRITING
Articles & notes
Essays
Why I orchestrate multiple LLMs instead of one
Cost/quality tiering in ApplyFuel — and what broke.
Grounding beats prompting
Most hallucinations are missing-context problems, not prompt problems.
Field notes · AI in plain words
Speculative decoding
A small model guesses, the big model checks. 2–3× faster, zero quality loss.
Quantization in one picture
Your model wants 80GB of VRAM. Your GPU has 24. Here's how you close that gap.
Your fine-tune is forgetting
You fine-tuned your model and it came out dumber. That has a name — and fixes.
Data quality beats quantity
200 clean examples beat 10,000 messy ones. Stop scaling the dataset — curate it.
Your inference is the bottleneck
Before you shrink the model, fix how you serve it. Setup alone swings cost 5–10×.