An AI-Ready CS Curriculum Starts With Assessment, Not With an AI Unit

Uncategorized

A moderation meeting, late in the term. Two teachers, one Grade 10 submission, and a disagreement that will not resolve. The program runs. The comments are sensible. The variable names are better than the student’s usual work. One teacher thinks it is fine. The other thinks it is not the student’s. Neither can prove anything, and after twenty minutes they move on to the next folder, because there are forty more.

The interesting thing about that meeting is not the submission. It is that the department has no procedure capable of settling it, and never needed one before.

The inference that stopped holding

Computer Science assessment has always graded the artifact. Mark the program. The program is the evidence.

That was not laziness. It was a sound inference. Producing a working program used to require the understanding, so the artifact stood in for the process reliably enough that nobody had to look at the process directly which was just as well, because looking at the process is slow and expensive and does not scale to forty submissions.

The inference held for decades. It stopped holding in about eighteen months. Nothing announced that it had stopped, and no policy document in any school marked the date.

What departments have already started doing off the books

Heads of computing have usually noticed this before they have words for it. The submissions got better. The questions in class got fewer. And teachers began quietly putting more weight on a thirty-second conversation at a desk than on the file that was handed in.

They are right to. That conversation is now the better evidence. But no rubric has a row for it, no report records it, and no moderation meeting can appeal to it. The most reliable assessment in the department is the one that leaves no trace.

The question that cannot be answered, and the one that can

Most schools are currently asking:

how do we tell whether AI wrote this?

That question has no good answer, and the evidence on that is unusually clean. OpenAI launched an AI text classifier in January 2023 and withdrew it on 20 July 2023, citing low accuracy. At launch it identified only 26% of AI-written text as likely AI-written, and flagged human-written text as AI 9% of the time. The company that built the model could not reliably detect the model. Third-party tools have not solved what the vendor could not, and the false positives land hardest on students whose writing is already atypical.

A department that builds its integrity policy on detection is building on a probability it cannot audit and cannot defend to a parent.

The question worth asking instead: what can this student do with this code, in this room, that someone who did not write it cannot?

That question has answers, and they are cheap to collect.

A test you can run this term on work you have already marked

Take ten submissions that scored highly. Ask each of those students two questions about their own code. Three minutes each, no notice, no new policy.

  1. If the requirement changed in this one way, where would you go first?
  2. Here is your code with one line altered. What breaks, and why?

Count how many can answer. That number is closer to your real pass rate than the marks are.

The point is not to catch anyone. It is to find out how far your marks and your students’ understanding have drifted apart because you cannot fix a gap you have not measured, and the gap is currently invisible by design.

What “AI-ready” actually means

Not a unit on artificial intelligence bolted into Grade 10. A programme is AI-ready when its assessment still means something after a model enters the room.

UNESCO’s AI Competency Framework for Students, published in 2024, sets out twelve competencies across four dimensions and three progression levels Understand, Apply, Create. The framing is useful here for one reason: you cannot assess Create with an artifact alone, because the artifact is the thing a model is best at producing. Create has to be assessed partly through explanation, choice and revision.

Which is a curriculum design problem, not a compliance problem.

The loop runs both ways

Leave the proxy in place and the cycle is stable and wrong. Marks rise. Confidence in the marks rises with them. Live questioning gets squeezed, because the results look healthy. Students learn precisely what is graded and optimise for it. The cohort reaches the point where someone asks them to reason out loud, and cannot.

Change what is graded and the cycle is stable and right. Explanation carries marks, so students prepare to explain. Preparing to explain changes how they use a model closer to a tutor, further from a ghostwriter. And the marks start meaning something again, which is the only durable form of academic integrity anyone has ever found.

We would rather a department spent one afternoon on the two-question test than a year drafting an AI use policy. The policy tells students what they may not do. The test tells you what your programme currently proves.

Where KODEIT ICT sits on this

KODEIT ICT’s assessment layer is built to record more than the finished artifact. Quizzes run in several formats multiple choice, matching, drag-and-drop, short answer with pass thresholds and attempt limits set per activity. Grading is question-level as well as overall, so a teacher can see which part of a problem a student lost, not just that they lost marks. Work can be resubmitted and grades revised, which makes the second attempt visible: what a student changed after feedback, and whether they understood why.

That last one matters more than it sounds. A revision history is reasoning evidence. An artifact is not.

The platform also returns an AI-generation probability alongside a suggested grade for submitted code. We would rather no school ever used that number to accuse a student. It is a prompt to go and ask a question and a teacher who asks the question does not need the number.

When your next set of results comes in, what would have to be true for you to believe them?


FAQ

Should we ban AI tools in Computer Science? A ban you cannot enforce moves the behaviour out of sight and leaves the assessment problem untouched. The harder and more useful work is deciding what a mark should now be evidence of.

Are AI detectors ever useful? As a signal that a teacher might want to ask a question. Not as a finding, and not as the basis of a sanction. The accuracy figures do not support the second use.

Does this mean we should stop grading code? No. Keep grading the code. Stop treating it as sufficient on its own, particularly at the top of the mark range where the difference between understanding and assembly is largest.

How do we do live questioning with large classes? Sample rather than survey. Ten students, three minutes each, on work already submitted, tells you most of what you need about a cohort of a hundred.

Do younger grades need this? Less urgently, but the habit of explaining your own work is easier to build at Grade 3 than to retrofit at Grade 11.

Where Scholario Fits

Scholario helps schools turn everyday classroom moments into structured learning pathways — connecting curriculum goals, teacher practice, and family engagement in one place.

Use this article as a prompt for leadership conversations: what should children experience consistently, and how do you make that visible across every classroom?

FAQ

Who is this article for?
School leaders, curriculum coordinators, and teachers looking for practical ways to strengthen learning beyond one-off theme weeks.
How does Scholario support this approach?
Scholario provides structured units, classroom routines, and progress visibility so community learning becomes part of the weekly rhythm — not a special event.
Can families be involved?
Yes. Share classroom learning goals in simple language and invite families to extend conversations at home with everyday examples from your community.

← Back to Blog