You Hired a Brilliant Intern and Never Onboarded Them
Your AI Intern
Picture the intern you would love to have. Reads everything. Fast. Never tired. Never defensive about feedback. Hyper-intelligent. Turns around a draft framework in twenty minutes that would have taken your team a week.
Now hand that person a client deliverable on day one with no context, no standards, and no idea what a bad answer costs you. They will produce something. It will look good, formatted well, and read confidently.
Somewhere in the middle of the content will be a statute citation they half remembered, a number they inferred, and a claim about current federal guidance that was true two years ago.
They are not lying. They are doing what a smart person does when they want to be useful and nobody told them where the edges are.
That is your AI. And the fix is the same fix.
I learned this lesson in a much lower-stakes way
I gave Copilot a detailed set of instructions for how I wanted it to work with me, including one embarrassingly specific rule: never use em dashes. An em dash (—) is a tell-tale sign that AI was used to generate content.
Later, I asked it to explain back how I had instructed it to respond. It confidently told me, among other things, “Never use em dashes.” Two paragraphs later, it used one.
That was a useful reminder: an AI can understand a rule well enough to explain it perfectly and still fail to follow it. Knowing the instruction is not the same as reliably executing it.
The two things that go wrong
Every AI failure I care about in this work falls into one of two buckets.
Accuracy failures. The model produces something confidently wrong. Fabricated citations are the worst version, because they are shaped exactly like real ones: right author style, plausible year, a section number that looks like it belongs. In education data work, a wrong reference to a federal requirement is not an embarrassment, it is a client problem.
Ethics failures. The model does something you would not have authorized. It repeats back data that should never have entered the conversation. It writes in your voice and invents an experience you never had. It hands you a recommendation with no reasoning, and you carry it into a meeting where you cannot defend it.
Neither is a technology problem you wait out. Both are onboarding problems. On our Education Intelligence team, this is not theory. It is how we keep AI-assisted work client-ready.
What good onboarding looks like
You would not tell a new analyst “Be accurate.”. You would tell them what to do when they cannot find a source. Vague instructions produce agreeable nodding and no behavior change. Specific instructions produce specific behavior.
Weak: “Be skeptical and double check your work.”
Strong: “Never present a citation unless it came from a document in this conversation. If I ask for a source and you do not have one, say so.”
The first is a personality description. The second is a rule with an observable output. You can tell whether it was followed. Four rules I would not work without:
• Never present a citation or reference unless it came from a document in this conversation.
• Every number must be traced to a document I provided, or it gets marked [UNSOURCED] inline.
• Flag anything “time-sensitive” as [VERIFY CURRENT] and state your knowledge cutoff.
• At the end of every deliverable, list the five claims you are least confident about, ranked.
That last one earns its keep. It turns a wall of uniform confidence into a short list you can actually go verify. The full accuracy and ethics set is in the companion resource.
The Trap
Here is the part that will bite you, same as with a real intern: you cannot ask them to grade their own work and call that quality control. Tell a model to self-critique and rate itself out of ten and it will comply. It will also rate its confident errors highly, because the process that produced the error is the process doing the rating. Self-assessment catches sloppiness. It does not catch conviction.
So, the instructions get you most of the way, and then you do the last part yourself.
· Open every source before it goes to a client.
· Review the finished work in a fresh conversation with no context, and ask what is wrong with it, not whether it is good.
· For every recommendation, be able to say why, independent of the fact that a model wrote it. If you cannot, it is not your judgment yet.
The Real Point
The intern in this story is genuinely excellent. That is what makes the oversight matter. An incompetent assistant is safe, because you check everything. A brilliant one is dangerous, because you stop.
Write the instructions. Then keep doing the part the instructions cannot do for you.
What is one instruction you have given your AI that actually changed its output?
Drop it in the comments.


