← All writing

Why I orchestrate multiple LLMs instead of one

The default instinct when you build an AI product is to pick the best frontier model and route everything through it. I did that for about a week. Then I looked at the bill.

ApplyFuel tailors a CV, a cover-letter email, and a pile of form-field answers to every job a user applies to. That’s not one LLM call — it’s a small pipeline, and the steps are not equal. The CV is what a human reads and judges. The form answers are throwaway text pasted into a box. Paying frontier-model prices for both is lighting money on fire at the cheap end.

So I stopped thinking “which model” and started thinking “which model for what step.”

Tier by action, not by request

The unit of configuration in ApplyFuel isn’t a model — it’s a ModelTierConfig with three slots:

  • Extraction model parses the raw job posting into structured fields. Cheap, mechanical, high volume.
  • CV model generates the tailored CV. This is the expensive slot, and the only one where I happily pay for quality, because it’s the visible output.
  • Base model handles the cover-letter email and the form answers. Cheaper, higher volume, good enough.

Each slot binds a provider key (PlatformKey), a model_id, and per-million input/output rates. The provider list — Claude, OpenAI, OpenRouter — is hardcoded; the model list under each is fetched live from that provider’s own API, so I never hand-maintain stale model names.

Underneath it all is one LLMService with call() for text and call_json() (with retry) for structured output. Tiering decides which model; the service just routes to the right provider. Swapping the CV model is a config change, not a code change.

Credits are real dollars

The part I got wrong first: pricing.

The original design had BASE_COSTS — a fixed credit price per action. “A CV costs N credits.” Clean, predictable, and a lie. Token usage swings with the job posting and the user’s profile. A dense senior CV burns far more tokens than a one-page grad CV, but both charged the same. Margin was random — sometimes negative.

So fixed costs went in the bin. Now every call is metered against actual tokens:

input_cost  = (input_tokens  / 1_000_000) * input_rate
output_cost = (output_tokens / 1_000_000) * output_rate
final_cost  = (input_cost + output_cost) * (1 + markup_pct / 100)

The rates live on the tier config, so the cheap base tier and the expensive CV tier price independently. And I keep two ledgers on purpose: ExpenseLog is what I paid the provider; LLMCallLog is what the user was charged in credits. Subtract one from the other and you see margin per call — the whole point of metering instead of guessing.

The other thing that broke: user-supplied keys

Early on, users brought their own API keys. Elegant in theory — their usage, their bill. In practice it was a security surface (storing other people’s keys) and a support sink (no quota, wrong provider, just rotated). It also made tiering impossible: I couldn’t route a CV to my good model if the user only pasted an OpenAI key.

Now the platform owns the keys. PlatformKey stores them encrypted at rest, admin-managed, multiple keys per provider allowed — so I can run two OpenRouter accounts and spread load or rotate without a deploy. Users buy credits; I handle the plumbing.

What keeps it honest

Tiering invites a quiet failure mode: you swap a model to save money and output gets worse without anyone noticing. So an LLM evaluation harness rubric-scores generations on every prompt or model change. If quality regresses, I see it before a user does. User tiers map straight to these configs — the free Spark tier is locked to the cheapest models, and the harness tells me whether “cheapest” is still good enough.

A single frontier model isn’t a strategy, it’s a default. The strategy is knowing which step the user actually grades — and only paying for that one.