Knowing what the model doesn't know
Hedge language — built into the system prompt
The previous lesson showed what good hedge behaviour looks like out of the box. The next step is making sure your assistant defaults to that behaviour every single time, not just when the question happens to be obviously time-sensitive. The way you do that is with one extra line in the Constraints slot of your system prompt.
The minimal hedge instruction
Add this line to any production system prompt that touches facts the user might assume are current:
If asked about prices, opening hours, menu items, stock, events, or anything else that changes over time, say "I'm not sure that's still current" and recommend the customer verify with our website or staff.
That single line does three things:
- Triggers on the category, not on a date. It names the kinds of thing that go stale — prices, hours, stock — rather than a cutoff date.
- Gives the model a fixed phrase to reach for ("I'm not sure that's still current") so its hedge does not drift across replies.
- Hands the user a recovery path ("verify with our website or staff") instead of leaving them stranded with "I don't know".
Why there is no date in that line
The obvious version of this instruction names the cutoff — "anything that may have changed since ". Don't write it that way, for three reasons that compound:
- You have to be right about the date, and the previous lesson showed how easily that goes wrong.
- It breaks silently the day you change models. A system prompt is portable; a cutoff date is a property of one specific model. Swap the model underneath and the line is now confidently wrong.
- It is the wrong trigger anyway. Your café's opening hours were volatile the day the model was trained and are volatile now. What makes a fact unsafe is that it changes, not that it changed after some particular Tuesday.
Naming the volatile categories fixes all three at once, and it is the instruction you actually meant.
What good hedge language sounds like
Good hedges are short, specific, and offer a next step. Bad hedges are long, apologetic, and trail off without a recovery path.
| Bad hedge | Good hedge |
|---|---|
| "I'm just an AI and I can't really know things, but I think maybe..." | "I'm not sure that's still current — please confirm with our menu page." |
| "Unfortunately I don't have that information." | "I don't have a record of that — call the shop on (number) to check." |
| "I think it's probably..." (then makes something up) | "That changed at some point — I can't confirm. The Zamalek shop will know." |
The pattern: name the gap, give the recovery path, do not pad with apology.
Bad hedge vs good hedge
Bad hedge
- User left stranded
- Sounds defeated
- May still slip in a guess
Good hedge
- User knows where to look next
- No false confidence
- Reusable phrase across replies
- Over-hedging makes an assistant useless
- One fixed phrase gets repetitive across a long thread
- The hedge can fire on things the model actually does know
When to use it
Add this line to any assistant where:
- Prices, hours, addresses, or product catalogues might change.
- The user might assume the model has live data when it does not.
- A wrong-but-confident answer would damage trust more than a hedge.
For Bayt Coffee, that covers almost everything except the brewing-method advice (which is timeless) and the basic FAQ (which Hagar can hardcode in Capabilities). For Hagar's main job — which involves a SaaS product whose changelog updates weekly — it covers nearly every reply.
The next lesson narrows the focus: when the model is grounded on a specific document, you want a different, tighter hedge. That is the "I don't see that in the document" trigger.
Next: the document-grounded hedge — making the model say only what it can read. :::
Sign in to rate