Back to blog
· Darja Koneva

When AI Gets It Wrong, Who Actually Carries the Can?

Where does AI accountability actually sit in a school — with the teacher who pressed send, or with the policy that shaped what they were checking for?

When AI Gets It Wrong, Who Actually Carries the Can?

~ by Darja Koneva, Edu:Insight Founder

So the teacher's using an AI tool to help draft feedback for a set of GCSE mocks. Nothing sneaky about it — the tool's school-approved, she was trained on it last term, she gives the output a read, it sounds fine, off it goes.

The very next day: an email from a parent. Something about a piece of coursework their child never actually submitted. Turns out the AI made it up. Confidently. Convincingly. Completely wrong.

The teacher didn't catch it. Nobody did — until, thankfully, a parent happened to check that day. Not something they did often, if they're honest — most weeks, nobody's reading a feedback comment that closely. This one just happened to be one of those days.

Thankfully, because that's not usually how it goes. Most parents aren't cross-checking a feedback comment against what their kid actually handed in — especially once the kids are older, when everyone assumes that level of oversight isn't needed anymore. And students won't question it either. If a teacher says "this claim isn't backed up," a student takes that on trust — that's the whole point of feedback, you're not meant to fact-check it. So normally, this kind of mistake just sits there. Quietly wrong, nobody in a position to notice.

And now the awkward bit starts. Not the parent conversation, oddly — that part's usually fine, a quick apology and a fix. The awkward bit is internal. Whose fault was that, really?

We've got our first blog post up already, on why a school actually needs an AI policy — proper documents governing it, and why a whole-school plan matters, so AI doesn't turn into chaos for teachers and confusion for students and parents. In this one, we want to try drawing a line: where does AI responsibility actually sit, and how do you make sure a teacher can rely on it when they're using it?

Maybe the real question is when and for what purposes AI should be used at all. Maybe it's whether there's a clear rubric for what AI is and isn't allowed to check, if it's marking a written task. Or maybe it's simpler than that — is the point to spend less time on marking, or to spend that time better?

"Always check it" isn't really a policy, is it

"AI content should always be checked before it's used" — true, should be in every school's guidance, no argument there. I've seen it myself, more than once — a student in my own class once asked AI to generate a map of Europe for a project, and what came back was confidently wrong: borders in the wrong places, countries missing, labels that didn't match anything real. Obvious once you actually looked at it. Easy to miss if you didn't.

But think about what "always check it" doesn't tell you: what happens when the checking itself reasonably fails.

Because not every mistake looks like a mistake. Some AI errors are the kind a careful, experienced teacher genuinely wouldn't spot on a normal read-through — a made-up reference that sounds exactly right, a statistic that feels plausible, a detail buried halfway down a summary that quietly contradicts everything above it. "Always check it" only works if someone's told you what you're actually checking for. Most staff haven't been.

So when something slips through anyway, "well, the teacher should've checked" becomes a suspiciously easy place to stop the conversation. Easy for the school. Not exactly fair on the teacher, who was — let's remember — doing exactly what they were trained to do, with a tool the school picked.

Which brings us to the actual question worth sitting with: is accountability about who happened to be standing closest when it went wrong? Or who was actually in a position to stop it?

Two stories that look the same but really aren't

Picture two versions of the same email-from-a-parent moment.

Story one: a teacher uses an AI tool nobody approved, skips the review step because it's Thursday and she has no time, sends home feedback that turns out to be made up.

Story two: a teacher uses the school's approved tool, exactly as trained, reviews it like she's meant to — and it's still wrong, because nobody ever mentioned this particular tool likes to invent coursework references, and nothing in the training covered how you'd catch that.

Same upset parent. Same email. Not remotely the same failure.

Story one is a person not following the process. Story two is the process not being good enough to follow in the first place.

If your policy answers both of those the exact same way — "the teacher should've been more careful" — you don't quite have a policy yet. You have a disclaimer with good intentions.

So where does the line actually go

Here's roughly how we think it should split.

Policy owns the stuff upstream. Which tools get approved. What data's allowed near them. What training staff actually got, and when. What the review step is meant to catch, specifically — not just "check it," but check it for what.

Concretely, that means the difference between "check it for accuracy" and something like: this tool tends to invent specific coursework titles and overstate how confident a source is — check those two things, every time. One of those is a sentence a busy teacher can actually use on a Thursday afternoon. The other just sounds like it.

And getting to that sentence starts before any tool is even approved. Whatever's on the list should have actually been researched first — not just for what it's good at, but for where it tends to get things wrong. Every tool has its own particular ways of being confidently, plausibly wrong, and a school can only warn staff about the ones it's bothered to find out about.

Same goes for scope. This isn't really about how many tools a school has on its approved list, either — there's no single tool that does everything well, so most schools will end up with a few. What matters is whether there's a clear, limited set of things each individual tool is actually approved for, and a working sense of the main ways it tends to go wrong — not "AI is approved for feedback," but specific, bounded uses staff can point to, tool by tool. And if one of those uses is checking written work, teachers shouldn't be left to work out how on their own. The school doesn't need to hand over a ready-made rubric — most teachers already have their own, built up over years of marking. What it should do is give clear guidance on uploading that existing rubric into the tool properly, so the AI has something concrete to check against instead of a vague sense of "does this look right."

Students need their own version of this, too — just aimed at a different thing. Not at second-guessing a teacher's feedback; that's not really the point, and it's not a habit worth building in them. The bit that actually matters is their own use of AI — for homework, for research, for drafting their own work. Same idea as the staffroom version: don't take an AI answer as automatically correct just because it sounds confident. Question it, check it against something else, notice when it doesn't quite add up. That's not really a standalone lesson, either — it's part of building a wider culture across the school around using AI responsibly and wisely, one that applies to staff and students alike, just showing up differently depending on which side of the desk you're on.

And that culture only works if everyone's actually honest about how they're using AI in the first place — students, teachers, parents, all of it. A policy can be as carefully written as you like, but it can only account for risks it actually knows about. If AI use is happening quietly, off to the side, unmentioned, the policy's working off an incomplete picture — and so is everyone relying on it.

None of that gets rid of the risk completely — nothing will. But it narrows it a long way, and it's the difference between hoping a teacher happens to notice and actually giving them a fighting chance.

Underneath all of it, the posture matters too: AI works best treated as a thinking partner, not a decision-maker — which really just means building in a human-in-the-loop step every time, somewhere a person is genuinely expected to look, not just technically able to.

A good chunk of that risk also gets narrower just by asking better questions in the first place. Three habits worth teaching staff, nothing fancy. Give it real context, not a bare question — I had this come up in one of my own lessons recently, planning a session on balancing chemical equations for a mixed-ability Year 9 class, 50 minutes, no lab that day. "Give me a lesson on balancing equations" hands the AI nothing to work with; typing in the actual constraints got something that was at least trying to fit the class in front of me. Ask it to show its working — "walk me through why you've sequenced these examples this way" surfaces shaky logic before it's trusted. And ask where something came from — a confident answer with nothing behind it is a different level of risk than one that can actually back itself up. None of that replaces the review step. It just means there's less of a guess to review in the first place.

Judgement owns the decision in the moment. Whether, for this piece of work, right now, the output's good enough and appropriate enough to use — based on what the school actually equipped that teacher to know. That's really just diligence and discernment, applied to something new: not skipping the read-through because it sounds fine, and being willing to trust your own gut when something feels slightly off, even when you can't immediately say why.

And accountability follows whichever side actually broke down. Ignored the process, used a tool nobody signed off, skipped the review — that's a staff conversation. Followed everything exactly as trained, and the training just hadn't caught up yet — that's a leadership conversation, not a staffroom one.

This isn't only about being fair to teachers, though it is partly that. A policy that always, every time, lands on "the teacher's accountable" gives leadership zero reason to ever improve the bits upstream — training, approvals, the list of known issues — because responsibility's never going to reach that far back anyway. A policy that can occasionally point at itself is the only kind that actually gets better.

Why nobody wants to write this bit down

Easy to see why most guidance stops short. "Staff must exercise professional judgement" costs a school nothing to write. "We'll keep a running list of how each approved tool tends to go wrong, review it every term, and update training when it changes" costs actual, ongoing effort. Not complicated to picture, though — a shared doc, one row per approved tool:

Tool: [feedback-drafting tool]. Known to: occasionally invent coursework references when summarising a body of work. Last updated: [date]. Checked by: [role].

Four fields. Doesn't take a data team, just someone willing to own it and open the document once in a while.

But that second sentence is the one doing any real work. The first is just a nice hope. We'd rather a school wrote the harder sentence and admitted the policy's still a work in progress, than wrote the easy one and quietly implied everything's covered. It isn't. It never fully will be. That's fine, as long as the document says so honestly instead of pretending otherwise.

Questions worth taking to your next SLT meeting

Not asking you to solve this in one sitting. Just useful to know where you currently stand:

  • For each AI tool your school's approved — is there an actual written list of how it tends to go wrong?
  • When did staff last get training on that specific list, and who's meant to update it?
  • If a mistake slipped through today, could you tell — honestly — whether it was a staff issue or a policy gap?
  • Who's actually meant to update the "known issues" list, and — honestly — how often does that happen versus how often it's supposed to?
  • Would a teacher six weeks into the job know what they're specifically checking for, or just that they're meant to check?

If most of those don't have a confident answer yet — completely normal, most schools don't. That's really where the actual work of an AI policy lives, underneath the easy parts.

Our view

Professional judgement has always mattered in teaching, and always will. AI doesn't change that bit. What it changes is how much a school owes its staff before asking them to use that judgement on something this new.

"Trust your judgement" is a fair thing to ask of a teacher. But only once the school's done its half — named the risks, trained for them, kept it current as the tools change. Skip that part, and "use your judgement" quietly turns into asking staff to carry a risk the school never actually prepared them for.

A few things worth holding onto. It was never about how many tools you've got — nobody's found the one magic tool that does everything, so most schools end up with a handful, and that's fine. What matters is knowing, tool by tool, what it's for and where it trips up. From teachers, it's really just diligence and discernment — don't panic about AI, don't blindly trust it either, just the same judgement you'd bring to anything else in the job. And underneath it all, honesty from everyone — students, teachers, parents. A policy can only cover what it actually knows about.

One more thing, and it's a big one: someone actually needs to own this. Not "the school," not "leadership" in the abstract — an actual person whose job it is to keep the list current and notice when a tool starts behaving oddly. Otherwise it's everyone's job, which means it's nobody's.

And none of this happens overnight, nor should it. It'll be a bit clunky at first, and you'll probably rewrite half of it in six months once you've seen what breaks. That's fine. What's not fine is waiting for the perfect version before you start, because that version doesn't exist. Start rough and honest, and you'll get somewhere. Wait for perfect, and you'll just be waiting.

That's the line we think belongs in every school's AI policy. Not a footnote. The backbone.

Download your free AI Policy Checklist

Still working out where your school currently sits on this? Our AI Policy Checklist is a genuinely practical place to start — what a good policy actually looks like, the questions your leadership team should be asking, and the common pitfalls we keep seeing. It's on our free guide page, right here on this website.

Or if you'd rather just talk it through, book a free 30-minute discovery call with us. We'll help you work out exactly where your line between judgement and policy sits right now — and what it'd take to make it one you'd actually be happy to defend.