5Evaluating AI output in under a minuteDetermine

Five checks to run before you trust the output

9 min read1,666 words

The Step Most People Skip

Here is the most common pattern in AI use today: someone types a prompt, receives a response, skims the first few lines, and thinks, "Yeah, that looks about right." Then they copy it, paste it into whatever they were working on, and move on with their day.

This is the equivalent of hiring a new employee, asking them to write a report on their first day, and submitting it to your board without reading it. You wouldn't do that with a person. You shouldn't do it with AI.

The "D" in the READY Method — Determine — exists because evaluation is not optional. It is the skill that separates people who use AI well from people who use AI dangerously. And it is, without question, the stage that most people skip entirely or rush through with a cursory glance.

Why do we skip it? Partly because AI responses look polished. They arrive in clean paragraphs, with confident language and professional formatting. There are no typos, no hesitation, no "um, I'm not sure about this bit." The presentation creates an illusion of authority. And partly because we're busy — the whole point of using AI was to save time, so spending time evaluating the output feels like it defeats the purpose.

But here's the thing: evaluation doesn't need to take long. What it needs is structure. A systematic way of checking that catches problems before they become your problems. That's what this lesson gives you.

The Five Quality Checks

Think of these as five lenses you hold up to any AI response. Not every response needs all five — a quick email draft doesn't require the same scrutiny as a financial analysis — but knowing all five means you can choose the right level of rigour for the situation.

1. Accuracy — Are the facts right?

This is the most fundamental check. AI systems can and do generate false information with complete confidence. They don't hedge, they don't say "I think" or "I'm guessing." They state incorrect facts with the same tone they use for correct ones.

What to look for:

  • Specific claims — dates, statistics, names, quotes. These are where AI is most likely to fabricate. If the response says "According to a 2023 McKinsey report, 67% of companies..." then you need to verify that report exists and says what the AI claims it says.
  • Technical details — formulas, code, legal references, medical information. AI can produce plausible-sounding technical content that is subtly wrong in ways that matter enormously.
  • Historical or current events — AI training data has a cutoff date, and even within its training period, it can mix up details, conflate similar events, or invent plausible-sounding incidents that never happened.

The practical test: Could you fact-check the key claims in under five minutes? If so, do it. If the response contains claims you cannot verify, flag that as a risk.

2. Relevance — Does it answer what I actually asked?

AI is remarkably good at producing content that is related to your question without actually answering it. You ask for a marketing strategy for a small bakery, and you get a generic overview of marketing principles that could apply to any business. Technically relevant. Practically useless.

What to look for:

  • Specificity — Does the response address your particular situation, or could it apply to anyone? If you could swap in a different company name or context and the advice wouldn't change, it's not relevant enough.
  • Scope alignment — Did you ask for three ideas and get twelve? Did you ask for a detailed plan and get bullet points? The AI may have answered a different version of your question than the one you intended.
  • Drift — AI responses sometimes start on topic and gradually wander. The first paragraph addresses your question; by the fourth paragraph, it's discussing something tangentially related that you didn't ask about.

The practical test: Re-read your original prompt, then re-read the response. Does the response deliver what the prompt requested? Not what you might also find interesting — what you actually asked for.

3. Completeness — Is anything missing?

This is trickier than it sounds, because you're looking for absences. It's easy to evaluate what's in front of you; it's harder to notice what isn't there.

What to look for:

  • Obvious gaps — If you asked for pros and cons and only got pros, that's incomplete. If you asked for a plan and there's no mention of budget, timeline, or risks, something is missing.
  • Perspectives not represented — Does the response only consider one stakeholder's point of view? One cultural context? One industry? If your situation involves multiple perspectives, the AI may have defaulted to the most common one.
  • Edge cases and exceptions — AI tends to describe the typical scenario. It often neglects unusual cases, exceptions, or "what if" scenarios that may be critically important to your situation.

The practical test: Ask yourself, "If I acted on this response alone, what could go wrong that isn't mentioned here?" If you can think of realistic scenarios the response doesn't address, it's incomplete.

4. Bias — Is it one-sided or skewed?

AI systems reflect the biases present in their training data, and they also have structural tendencies that create their own form of bias. They tend towards majority viewpoints, Western perspectives, optimistic framings, and conventional wisdom. None of this is flagged. It simply appears as the default.

What to look for:

  • Framing — Is the response presenting something as universally positive (or negative) when reasonable people disagree? Words like "clearly," "obviously," and "undoubtedly" are warning signs.
  • Source bias — AI training data over-represents certain sources, industries, and cultural contexts. A response about business strategy will lean heavily towards American corporate norms. A response about education will favour certain pedagogical traditions over others.
  • Sycophancy — AI has a well-documented tendency to agree with whatever position you seem to hold. If your prompt implied a preference, check whether the response simply validated it rather than giving you an honest assessment.

The practical test: Ask, "Would someone with a different background, ideology, or set of experiences agree with this framing?" If the answer is clearly no, the response has a bias problem you should account for.

5. Usability — Can I actually use this as-is?

This is the pragmatic check that brings everything back to earth. Even if a response is accurate, relevant, complete, and balanced, it still needs to be usable in your actual context.

What to look for:

  • Format and structure — Is it in a format you can work with? If you need a spreadsheet and got an essay, the content might be fine but the format isn't usable.
  • Tone and register — Is the language appropriate for your audience? A response written for executives won't work if you're presenting to year-ten students, and vice versa.
  • Actionability — Can you act on this, or is it too abstract? "Improve your communication" is advice. "Send a weekly update email every Friday summarising three things: progress, blockers, and next steps" is actionable.

The practical test: Imagine taking this response and using it immediately, with no editing, in the context you need it for. What would break? What would need changing? The gap between "received" and "ready to use" is your usability score.

Why Domain Expertise Matters Enormously

Here is an uncomfortable truth about AI evaluation: your ability to judge quality is directly proportional to your existing knowledge of the subject.

If you're a qualified accountant and you ask AI to explain a tax regulation, you'll immediately spot errors, oversimplifications, and outdated information. You'll notice if it's describing the rules for a different jurisdiction. You'll recognise when it's technically correct but practically misleading. Your expertise acts as a powerful filter.

Now imagine you ask AI about quantum computing, and you have no background in physics. The response arrives looking crisp and authoritative. It uses technical terms confidently. It sounds right. But you have no filter. You cannot distinguish a brilliant explanation from a plausible-sounding fabrication. You're evaluating presentation, not substance.

This is not a reason to avoid using AI outside your expertise — it's a reason to be honest about the limits of your evaluation. When you're working within your domain, you can trust your judgement to catch most problems. When you're working outside it, you need additional strategies: cross-referencing with authoritative sources, asking a knowledgeable colleague to review, or using the AI itself to stress-test its own claims.

The Danger Zone

The most dangerous scenario in AI use is not asking about something you know well (you'll catch errors) or something you know nothing about (you'll probably be cautious). The danger zone is topics where you know just enough to feel confident but not enough to catch sophisticated errors.

Think of the person who took one psychology module at university and now confidently evaluates AI-generated content about cognitive behavioural therapy. Or the manager who understands basic finance and doesn't question AI's analysis of complex derivative instruments because the terminology sounds familiar.

A little knowledge doesn't just fail to protect you — it actively makes you more vulnerable, because it gives you false confidence in your evaluation. Be especially cautious when a topic feels familiar but isn't truly within your expertise.

Key Takeaways

  • 1Apply five quality checks to every AI output: accuracy, relevance, completeness, bias, and usability.
  • 2AI responses look polished and confident regardless of whether they are correct, so never trust presentation as a proxy for quality.
  • 3Your ability to evaluate AI output is only as strong as your existing knowledge of the subject matter.
  • 4The most dangerous zone is topics you know just enough about to feel confident but not enough to catch subtle errors.
  • 5Evaluation does not need to take long — it needs structure and the right questions.