Why I orchestrate multiple LLMs instead of one
2026-05-20
The default instinct when you build an AI product is to pick the best frontier model and route everything through it. I did that for about a week. Then I looked at the bill.
ApplyFuel tailors a CV, a cover-letter email, and a pile of form-field answers to every job a user applies to. That’s not one LLM call — it’s a small pipeline, and the steps are not equal. The CV is what a human reads and judges. The form answers are throwaway text pasted into a box. Paying frontier-model prices for both is lighting money on fire at the cheap end.
So I stopped thinking “which model” and started thinking “which model for what step.”
Tier by action, not by request
The unit of configuration in ApplyFuel isn’t a model — it’s a ModelTierConfig with three slots:
- Extraction model parses the raw job posting into structured fields. Cheap, mechanical, high volume.
- CV model generates the tailored CV. This is the expensive slot, and the only one where I happily pay for quality, because it’s the visible output.
- Base model handles the cover-letter email and the form answers. Cheaper, higher volume, good enough.
Each slot binds a provider key (PlatformKey), a model_id, and per-million input/output rates. The provider list — Claude, OpenAI, OpenRouter — is hardcoded; the model list under each is fetched live from that provider’s own API, so I never hand-maintain stale model names.
Underneath it all is one LLMService with call() for text and call_json() (with retry) for structured output. Tiering decides which model; the service just routes to the right provider. Swapping the CV model is a config change, not a code change.
Credits are real dollars
The part I got wrong first: pricing.
The original design had BASE_COSTS — a fixed credit price per action. “A CV costs N credits.” Clean, predictable, and a lie. Token usage swings with the job posting and the user’s profile. A dense senior CV burns far more tokens than a one-page grad CV, but both charged the same. Margin was random — sometimes negative.
So fixed costs went in the bin. Now every call is metered against actual tokens:
input_cost = (input_tokens / 1_000_000) * input_rate
output_cost = (output_tokens / 1_000_000) * output_rate
final_cost = (input_cost + output_cost) * (1 + markup_pct / 100)
The rates live on the tier config, so the cheap base tier and the expensive CV tier price independently. And I keep two ledgers on purpose: ExpenseLog is what I paid the provider; LLMCallLog is what the user was charged in credits. Subtract one from the other and you see margin per call — the whole point of metering instead of guessing.
The other thing that broke: user-supplied keys
Early on, users brought their own API keys. Elegant in theory — their usage, their bill. In practice it was a security surface (storing other people’s keys) and a support sink (no quota, wrong provider, just rotated). It also made tiering impossible: I couldn’t route a CV to my good model if the user only pasted an OpenAI key.
Now the platform owns the keys. PlatformKey stores them encrypted at rest, admin-managed, multiple keys per provider allowed — so I can run two OpenRouter accounts and spread load or rotate without a deploy. Users buy credits; I handle the plumbing.
What keeps it honest
Tiering invites a quiet failure mode: you swap a model to save money and output gets worse without anyone noticing. So an LLM evaluation harness rubric-scores generations on every prompt or model change. If quality regresses, I see it before a user does. User tiers map straight to these configs — the free Spark tier is locked to the cheapest models, and the harness tells me whether “cheapest” is still good enough.
A single frontier model isn’t a strategy, it’s a default. The strategy is knowing which step the user actually grades — and only paying for that one.