Grounding beats prompting
2026-05-27
The single biggest quality jump in ApplyFuel didn’t come from a cleverer prompt. It came from putting the user’s real data in front of the model.
For a while I did what everyone does: when the output was wrong, I rewrote the prompt. More instructions, more “be accurate,” more “do not fabricate.” It helped a little and then stopped helping. The model kept inventing things — confidently, fluently, and wrong.
The CV example that made it obvious
Ask an LLM to “tailor my CV to this job description” with nothing but the job description, and it will give you a great CV. The problem is it’s not your CV. It will add a year at a company you never worked for. It will claim a framework you’ve never touched because the job asked for it. Every line is plausible. Every line is a liability.
You cannot prompt your way out of this. “Don’t make things up” doesn’t work, because the model isn’t lying — it has no idea what’s true. There’s no you in its context. It’s filling a CV-shaped hole with the most likely tokens, and the most likely tokens are whatever the job description implied.
The fix wasn’t words. It was data. Inject the user’s actual experience, skills, and history into the prompt, and the task flips from invent a candidate to tailor this real one. Same model, same instructions — completely different failure profile.
# ungrounded — model fills the gap by guessing
"Tailor a CV for this job: {job_description}"
# grounded — model can only rearrange what's true
"Tailor this person's real CV to this job.
PROFILE: {experience, skills, education, history}
JOB: {job_description}"
Grounding has two halves: supply and constrain
Getting the right data in is half of it. The other half is not letting the output drift.
In the CV pipeline, tailoring.py builds the ATS-optimized prompt, and parse_tailored_cv() validates the structured result — it has to contain a summary, experience, education, and skills, or it’s rejected. That validation is part of grounding: you constrain the shape so the model can’t wander off into freeform prose. The validated structure then renders to an ATS-friendly PDF through a Typst → PDF pipeline, so the formatting is deterministic templating, not the model improvising layout.
Supply the truth, constrain the shape. Both, or you’re still guessing.
The same lesson, retrieved instead of injected
My domain-agnostic RAG chatbot is the other flavor of the same idea. It’s a LangChain pipeline over a ChromaDB vector store with sentence-transformers embeddings. Documents — PDF, TXT, DOCX, Markdown — get chunked recursively, retrieved at query time, and the answer is streamed back with citations to the chunks it came from.
The citations are the point. When the answer links back to its source chunk, a hallucination stops being invisible. If the model claims something the cited chunk doesn’t support, you can see it. Grounding without citation hides the failure; grounding with citation surfaces it.
Why this makes the model swappable
Here’s the part I didn’t expect. Once the context carries the truth, the specific model matters far less.
In the RAG chatbot the provider — OpenAI, Anthropic, LM Studio, HuggingFace — is swappable at runtime, and that only works because the answer comes from the retrieved context, not from the model’s parametric memory. You’re no longer hostage to one model’s training data or one vendor’s vibe. The truth lives in your context window, not in their weights.
That’s the real payoff. Grounding doesn’t just stop the lying. It moves the source of truth from inside a model you don’t control to inside a context you do.
Takeaway
When generation is wrong, my first question now isn’t “how do I phrase this better.” It’s “what does the model not know, and how do I hand it that — and how do I check it kept it.” Reach for the data before you reach for the adjectives.