Why Your AI Assistant Just Might Be Making Things Up
In August 2024, a lawyer in Australia relied heavily on ChatGPT to research case law for a Fair Work Commission hearing. The AI confidently cited legal precedents that didn't exist. The lawyer filed briefs based on these fabricated cases. The judge was not amused. The case collapsed, and the lawyer faced professional consequences.
This wasn't a freak accident. This is what happens when managers and professionals treat AI like a reliable expert instead of a tool that needs verification. And it's happening more often than you think.
Here's what should keep you awake: according to a 2024 survey by enterprise software firm Censuswide, 68% of managers have already caught AI tools confidently stating false information. But how many mistakes did they miss? That's the real problem.
Understanding AI Hallucinations (Without the Jargon)
When an AI tool makes up information and presents it as fact, that's called a "hallucination." It's not malicious. It's not broken. It's just how these models work. They predict what word comes next based on patterns, sometimes confidently generating complete lies.
Think of it like this: if you ask ChatGPT to write a memo about your company's Q3 financials, and you never fed it those numbers, it will happily invent them. It doesn't know what it doesn't know. It just keeps the response sounding plausible.
The scarier part? These hallucinations often sound authoritative. AI doesn't whisper uncertain answers. It doesn't say "I'm not sure." It declares things with confidence you'd expect from an expert.
The Three-Check System Every Manager Needs
You don't need to become a data scientist to catch AI mistakes. You need a system. Here's one that actually works:
Check 1: Source Verification (the most important one)
Whenever an AI tool cites a fact, statistic, legal precedent, regulation, or policy, demand proof. Actually demand it. Ask the AI to provide the source. Better yet, verify it yourself in 30 seconds using Google or your industry database.
Here's a real example: You ask Claude to summarize your state's overtime regulations for a memo to your team. Claude provides four specific rules with exact wage calculations. Do not send that memo to your team until you verify those rules on your state's labor department website. One wrong number in that guidance could expose you to legal risk and destroy trust with employees.
The fix: Before using any AI-generated fact in communication, business decisions, or policy, ask "Where is this from?" and check it. Personally. Takes five minutes, saves you thousands.
Check 2: Consistency Across Tools
Run the same question through two different AI tools. ChatGPT, Claude, Gemini, or even a smaller model like Llama 2 (available through various platforms). Do they give the same answer?
If you ask ChatGPT and Claude to summarize your industry's top regulatory compliance requirements, they should align on the major points. If Claude mentions something ChatGPT missed, that might be valuable. If they directly contradict each other, you have a red flag. Neither tool is reliable enough to use without manual verification.
Example: You're building an HR compliance checklist. You ask ChatGPT to list GDPR requirements for employee data. You copy that to Claude and ask the same thing. ChatGPT lists five requirements. Claude lists the same five but adds two more specific to certain industries. You now know to dig deeper into those two points before finalizing your checklist.
Check 3: Smell Test (Your Domain Knowledge)
If the AI output doesn't align with what you actually know about your business, industry, or common sense, pause. Don't assume the AI is smarter than you. It's not.
You work in customer service and ask Gemini to suggest ways to reduce ticket resolution time. It recommends implementing a feature your system doesn't support and never has. Red flag. That's a hallucination, not brilliant advice you missed. Your job is to catch it and discard it.
Real-World Scenario: How to Audit AI Output
Let's say you're a regional manager at a logistics company. You use ChatGPT to draft new shift-scheduling guidelines. Here's how to audit that output before it touches your team:
- Run the source check: Does ChatGPT reference any labor laws? Copy each claim. Verify on your state's labor board website. Takes 10 minutes. If anything is wrong, discard the entire draft and start over with verified sources.
- Check consistency: Paste the ChatGPT draft into Claude and ask, "Are there any inaccuracies or missing labor law requirements in this scheduling guide?" Claude will often catch things ChatGPT missed or made up.
- Apply your domain knowledge: Does the schedule work for your actual warehouse? Your existing staff? Your peak seasons? If the AI suggested something that ignores your operational reality, fix it or throw it out.
- Test with a small group first: Don't roll new policy across your entire operation based on AI-generated guidelines. Share the draft with your most experienced shift leads. Let them punch holes in it. Use their feedback to refine before company-wide rollout.
This process takes an hour instead of a day. You catch errors before they become policy disasters.
When to Trust AI Output (and When Not To)
AI is genuinely useful for some things and genuinely dangerous for others. Know the difference.
Safe uses (low risk if hallucinations occur): Brainstorming meeting agendas, drafting internal emails and memos, summarizing industry articles you already read, creating outlines for training materials, generating rough layouts for dashboards.
Dangerous uses (high risk if hallucinations occur): Citing legal requirements, quoting regulations, making financial projections, diagnosing compliance issues, providing medical or safety advice, writing policy based on false assumptions about your business.
The pattern? AI is safe for creative, exploratory work. It's dangerous for factual claims with legal, financial, or safety implications.
If you're using AI for anything in the "dangerous" category, you need the three-check system above. No exceptions. Not because AI is fundamentally broken, but because your job is to protect your business and team.
Implementing This in Your Team
You can't audit every AI output yourself. But you can build a culture where your team does. Here's how:
First, teach your managers and high-touch staff the three-check system. Make it part of your AI guidelines. When someone says "I used ChatGPT to research this," the response should be automatic: "Show me your sources and verification."
Second, create a simple one-page template for auditing AI outputs. Have people fill it out when they use AI for decisions affecting compliance, policy, or customer communication. It shouldn't take more than five minutes. The template forces the right questions.
Third, use smaller, focused AI models where possible. Smaller AI models often hallucinate less because they're trained on smaller datasets and for specific tasks. If you're using Claude or ChatGPT for everything, you might be using sledgehammers for pushpin jobs.
Finally, treat AI errors as learning opportunities, not failures. When someone catches a hallucination, celebrate it. Share the story. Make it normal to verify before trusting.
The False Choice Between Caution and Efficiency
Some managers worry that verification adds too much overhead. "If I have to check everything AI does, what's the point?" Fair question.
The answer: you don't check everything. You check the outputs that matter. High-stakes decisions, customer-facing content, policy documents, anything with legal or financial implications. That's maybe 20% of your AI use. The other 80% can run faster with less friction.
The Australian lawyer didn't need to verify every sentence. He needed to verify the legal cases he cited. He didn't. That one choice cost him his case.
You're not being paranoid about AI. You're being professionally responsible. There's a difference.
FAQ
Should we ban AI tools to avoid these mistakes?
No. That's throwing away a genuinely useful lever. The answer is better verification habits, not avoidance. Teams that use AI correctly move faster than teams that don't. But "correctly" means treating it as a draft tool, not a final authority.
What if I catch an AI making up facts after we've already acted on them?
Stop, assess the damage, and correct it immediately. If it affected employees, customers, or compliance, inform relevant parties. Document what happened so you can improve your verification process. Don't cover it up. That's how the Australian lawyer's problem got worse.
Is one AI tool more reliable than others?
They all hallucinate. ChatGPT, Claude, Gemini, Llama 2 - they're just different in frequency and style. Claude tends to be more cautious about stating uncertainty. ChatGPT can be more confident and slightly more hallucination-prone. Gemini is good for data analysis and reporting because you can feed it verified data. None are perfect. Trust nothing unverified.
How do I train my team to catch these errors?
Share the three-check system in your next team meeting. Walk through one example. Ask people to apply it to their next AI output before they send it to you or anyone else. After a few weeks, it becomes habit. Make verification normal, not paranoid.
Learn AI the Structured Way
This blog post scratches the surface. Our courses go deep with hands-on modules, real templates, and skill assessments.
Get the Free AI Playbook