A teacher sets a take-home task, but what are they really assessing?
Generative AI tools can convincingly imitate advanced human outputs across many fields. How can educators rise to the challenge to ensure assessments are valid, reliable and authentic?
Scroll, decide, reveal. About 20 minutes.
01/ 10Activity· How students are using AI
Pressure-test a take-home task
Below are three tasks schools set precisely because they are supposed to be hard to fake. Send one to a free AI assistant and watch what comes back, in real time, at the speed a student would get it. Try all three: each is defeated in a different way.
Send a task →
[[ demoSeen ]]/3 sent[[ demoHint ]]
!
The safeguard this walks through[[ heroDefeats ]]
AINew chat
[[ chatPrompt ]]
AI
[[ chatLead ]]
[[ chatSample ]]
[[ chatNote ]]
Answered in [[ chatSeconds ]]. Ask again and it returns a different one.
Ask a follow-up…↑
The reasonable objection
“I would notice.”
It is the first thing most of us say, and it is a fair thing to say. It has also been tested.
Which is why the useful question is not “can AI touch this task?” It is “what does this piece of work let me claim about this student?”
Having seen that, where do your own tasks sit?
[[ voteTitle ]][[ voteBody ]]
02/ 10Activity· What a mark actually means
What does a score actually tell us?
A mark is a number standing in for a claim: this student can do this thing. Below is one Year 9 result, and four students who could each have produced it. Click each one to see what the same 92 would mean about the student in front of you.
Yr 9 Science · take-home
Research task: “Evaluate two energy sources for a NSW coastal town”
92%
Submitted 11:48 pm, night before dueNo supervision, no AI restriction stated
Marking criteria, as written
[[ c.name ]][[ c.marks ]]
Read it again and notice what is missing. Not one criterion says the thinking has to be the student's own, and none of them mentions AI. So a student who directed a tool well can satisfy every line of this rubric, honestly, and still leave you unable to say what they know.
Four students could be behind this 92. Click each one.
Each option is a different student. The panel underneath shows what the mark would then be telling you, and what it would be hiding.
[[ storyEyebrow ]][[ storyTitle ]][[ storyBody ]]
One mark, four completely different meanings. The 92 cannot tell them apart, and neither can the rubric that produced it. That is workable, as long as nobody downstream treats the number as though it already has.
03/ 10Research· What teachers and students told us
Teachers and students agree on the problem. Not always the remedy.
Gradeo’s 2026 Measure of Learning Report surveyed over 650 students and 160 teachers and school leaders from over 150 schools.15 Alongside a series of focus groups, the results show how sharply confidence in take-home evidence has shifted across schools.
Where this actually gets decided
Teachers described this to us as a faculty conversation, not a policy problem.
The hard part was rarely accepting that AI changes assessment. It was agreeing, across a faculty, what a mark is now allowed to claim, and then marking consistently against that.
[[ stat1 ]]of students
say their year group commonly uses AI to do homework or take-home assessment.
n=653 · 91% in Year 9
[[ stat2 ]]of teachers & leaders
no longer trust that take-home work reflects a student's own effort.
n=167
[[ stat3 ]]of teachers
still count take-home written or research tasks among the fairest formats. That is 2 people out of 167.
13% of students agree
[[ stat4 ]]of teachers & leaders
want assessment redesigned rather than policed by detection tools.
Only 12% prioritise detection
Guess before you look
How worried were the students?
Set both estimates, then reveal the survey results. The useful comparison is between recognising the problem and wanting to remove take-home work.
What share of students agree that AI has made it harder to tell what someone genuinely knows from take-home work?
your guess [[ g1 ]]%students said 68%
You were out by[[ gap1Val ]][[ gap1Note ]]
And what share of students would back pausing take-home assessment altogether while schools sort this out?
your guess [[ g2 ]]%students said 31%
You were out by[[ gap2Val ]][[ gap2Note ]]
[[ guessTitle ]][[ guessBody ]]
Harder to judge abilityStudents 68% · Teachers 84%
Pause take-home workStudents 31% · Teachers 52%
“
The clause said we could use AI for brainstorming, but not for the coding, which I guarantee you the entire cohort did anyway.
Year 12 studentGovernment school
“
Even though it says you can't use AI, everyone, including me, used AI to generate ideas for their project.
Year 12 studentGovernment school
That brings us back to assessment literacy: define the claim first, then decide what evidence could support it.
04/ 10Activity· Where a mark comes from
Every mark passes through six steps. It can only be as strong as the weakest one.
Between deciding what to teach and writing a number in your markbook, you make six separate decisions. When a mark turns out to mean less than it looked like, it is almost always because one of these six was skipped, rushed, or left implicit, and not because the marking was careless.
Walk the six with the Year 9 energy task from earlier. Click any step to see what the decision is, what it looked like on that task, and what goes wrong when it is weak.
›
Step [[ chainNum ]] of 06 · the decision
[[ chainLabel ]]
[[ chainWhat ]]
On the Yr 9 energy task, this was:[[ chainExample ]]
If you get this step wrong
[[ chainWeak ]]
What assessment people call it[[ chainTerm ]]: [[ chainTermDef ]]
Notice that only one of those six steps is about marking. Step 06 is the one that reaches the student, which is why the standards define validity as a property of the claim you make from a mark, not of the task itself.3 There is no such thing as a valid task in the abstract, only evidence strong enough, or too weak, for the particular thing you want to say.1
05/ 10Research· Picture your own class
An underperforming student using AI to suddenly outperform may be easy to spot, but others are more challenging.
Imagine you teach the following class of 44 students. You want to understand how your class' take-home marks compare with a later supervised task. The numbers are designed to show a pattern, not report a finding.
Evidence from professional-writing experiments suggests generative AI can disproportionately improve the outputs of lower-performing participants, narrowing visible differences in finished work.5 That finding does not automatically transfer to school assessment. The class below is synthetic: it illustrates how differential AI assistance could make one spectacular discrepancy obvious while smaller distortions across the middle remain difficult to see.
Keep scrolling. The chart on the right changes as you go.
The supervised task in this comparison. Same students, same course, same rubric as the take-home. Only the conditions differ.
First: the distributions
Results have been released and you notice that take-home marks are higher and more tightly bunched, whereas supervised marks are lower and more spread out.
There may be multiple reasons for this from difficulty, timing, anxiety, usage of AI or exam nerves. The data tells us there's a discrepancy but not why.
Higher average scores under unsupervised conditions are a long-standing finding in ability testing, and they predate AI.16 AI can further widen that gap, compress visible differences, or make the causes harder to interpret.
Then: the same students
Mapping each student, you notice a steep slope showing that the class all outperformed in the take-home compared with the supervised task.
The student everyone notices
One student stands out. You likely noticed without crunching the data given this student moved from 91 to 48. That warrants inquiry, but not an accusation.
A bad day, exam anxiety, outside help, AI use or illness could all fit the same line. The chart tells you where to ask, not what happened.
The students easy to miss
The middle shifts are harder to spot as they've all dropped by roughly a grade. No single line looks extraordinary, so no single student triggers concern.
This is the pattern a masking effect can create: polished performance narrows visible differences without showing whether underlying learning changed. It is an illustration of a possible masking effect, not a documented pattern in this class or a settled finding about school students.
Ask for another angle
The compression sketched here comes from an experiment on adult professional writing, not from school assessment and not from this class.5
That is not evidence about this class. It is a reason to triangulate: compare a second task, a short explanation or an observed performance before changing a judgement.
Synthetic class · 44 students
Take-homeSupervised
[[ vizMode ]]
[[ gradeViz ]]
None of this makes the exam mark the “true” one. The point is that conditions change the evidence. An exam can settle who did the work and still miss most of what you taught.
“
I've had friends score 100 out of 100 on a take-home task, then score really low on the next one. Some of them failed it.
Year 12 studentGovernment school
“
It is extremely evident in examinations, where their results greatly differ from take-home tasks and projects.
TeacherIndependent school
Students who describe themselves as high-performing are more likely to notice it
Asked whether AI has made it harder for their own work to stand out, the students who say they do well agree most.
High-performing63%
Average49%
Students self-rating as high-performing or above average, n=343; average, n=212. Agree or strongly agree, excluding unsure.
06/ 10Activity· Have a go yourself
Which page is authentic?
You're concerned about AI so you've asked your class to handwrite a response to the same Othello question. You've left some notes and can see different handwriting and edits. What can you conclude about the students from these submissions?
Book AYr 12 English Advanced · Module B · markedBook BYr 12 English Advanced · Module B · marked
Which page is the student’s own work?
[[ artTitle ]][[ artBody ]]Evidence note: Weber-Wulff et al. tested 14 AI-text detection systems and found none accurate and reliable across the test set.12
“
I don't think this person, splitting their time across all their other subjects, could have done what they did. And it had that look that it was AI.
Year 12 studenton a classmate's project
Students judge each other on the same cues teachers use: fluency, tidiness, how much time someone had. Those cues are worth noticing. They are not evidence of who wrote the work.
07/ 10Research· Six tasks schools report trying
Useful evidence. Weak proof on its own.
Teachers are trying an assortment of task types to authenticate student learning, including the six below. Flip each card to see what AI can imitate, and what the task might still contribute.
A 2026 scoping review reached the same basic caution for reflective writing: it may remain educationally useful while becoming weaker as authenticated summative evidence.7
[[ cardsOpen ]]/6 pressure-tested
Several traces, combined with a short follow-up or observed sample, may make the inference much stronger, but none of these tasks proves authorship alone.
“
If you say “I'll see you guys in seven weeks”, you're going to get an AI surprise. In-class check-ins are more work, but I can see the product developing.
TeacherIndependent school
“
The only way I can guarantee it's their own knowledge is the in-class test. And I hate saying that, because a lot of my students are great prac workers who suck at exams.
TeacherGovernment school
08/ 10Activity· Allowing AI in a monitored classroom
What if I watch them using AI in class?
While observing a student's interactions (or prompt logs) with AI can strengthen attribution, doing so sets the stage for construct-irrelevant variance if each task inadvertently measures AI-proficiency, rather than the underlying skill or knowledge intended within the syllabus.
What you can see from here
You stand next to the student and watch them write the prompt, notice a suspect claim and revise the response.
This is stronger process evidence than merely receiving a final document. But what is the task trying to measure?
When we change the construct, we may be changing what we can reasonably measure.
So does watching count as evidence? It depends entirely on what you were trying to measure.
Say you watched a student do the four things below. Whether each one is evidence or just activity changes completely depending on the skill or knowledge the task was meant to capture. Choose one and watch the four steps re-label themselves.
The skill or knowledge I am trying to measure:
What you watched the student do
STEP 1
Writes a careful prompt
Sets the audience, the scope and the format to come back.
[[ w0Label ]]STEP 2
Catches a wrong claim
Spots an unsupported statement and checks it against a source.
[[ w1Label ]]STEP 3
Explains it out loud
Answers your follow-up questions with the laptop closed.
[[ w2Label ]]STEP 4
Rewrites the draft
Restructures the AI answer and cuts what does not belong.
[[ w3Label ]]
Counts as evidencetells you something about the skill you pickedJust activityyou saw it happen, but it does not evidence that skill
You chose to measure: [[ conName ]]
[[ conTitle ]] [[ conBody ]]
“
I asked students to explain sections of their assessment, and it was very clear who actually did the work and who didn't.
TeacherIndependent school, after allowing AI in a coding project
09/ 10Challenge· But won’t they need AI in the real world?
True. It still does not follow that every task should allow it.
AI-supported performance and independent learning are different outcomes. A task can support one without establishing the other. If every task measures how a student uses AI, then we may be masking students’ true capabilities.
Working well with AI is a real skill
If your outcome includes choosing, directing, checking or improving what AI produces, teach it, and assess it deliberately, with criteria that say so.
You cannot check what you do not know
Reasoning and evaluation depend heavily on relevant domain knowledge. Comprehension, reasoning and critical thinking are not free-floating skills that transfer cleanly to any content: they draw on what is already held in long-term memory. Readers understand a text about cricket largely to the extent that they know cricket.8
Which is exactly why the “they will have AI anyway” argument cuts the other way. A student can only spot a fluent, confident, wrong answer in a field they know something about. Knowledge is a large part of what makes AI safe to use, so the case for building it, and for assessing it independently at least sometimes, gets stronger as the tools get better, not weaker.
A useful distinction from one field experiment
Better work now did not always mean more learning later
Nearly 1,000 high-school maths students; results are specific to this setting and intervention.13
With standard GPT during practice+48%
Later, without GPT−17%
A guardrailed tutor also raised practice scores, but removed the later penalty; it did not produce a later advantage over the control group.
Twenty-second demonstration
Three sentences from an AI answer. One of them is wrong.
Year 8 Science. Read them, then reveal the false one. This is a demonstration rather than a test: the writing is uniformly fluent and confident, so nothing in the prose marks out which sentence to check.
During photosynthesis, plants use light energy to convert carbon dioxide and water into glucose.
The process takes place mainly in the mitochondria.
Oxygen is released as a by-product.
[[ errTitle ]][[ errBody ]]
And here is the tension in one number. 40% of students say take-home work has become more useful for their learning since AI arrived; 24% say less. Of the students who find it more useful, 73% also say it has made real ability harder to judge. Help with learning and proof of learning have come apart, and students can see both halves at once.
“
Gained: a student stuck at 9pm can get unstuck instead of stopping. Lost: the productive struggle, which is where a lot of the learning happened.
TeacherGovernment school
“
Not everyone can afford a $600-an-hour tutor. Don't get it to give you the answer. Get it to explain why one answer is better than another.
Year 12 studentMeasure of Learning Report 2026
10/ 10Activity· Choosing a method
Every format has its own trade-offs
Compare attribution, meaning confidence that the student produced the evidence, with coverage of the intended construct. Choose a method, then change the construct.
Step 1. The skill or knowledge you want to measure
Every method below is good at capturing some of these and poor at capturing others. Change this first, and the whole map moves.
Step 2. The method you would use to assess it
[[ methodTitle ]]
[[ methodBody ]]
[[ methodMap ]]
How to read this: positions are an editorial prompt, not empirical scores. In practice they move with task design, conditions, students and implementation.
RESEARCH INSIGHT: WHAT THE SECTOR PICKED IN 2026
Which formats feel fairest now?
We asked students and teachers which assessment formats they felt were fairest in the age of AI. Teachers and students agree on the top three. Teachers were also asked which formats best show what students know. Observed practical tasks jump from 28% to 46% there, the biggest gap of any format.
[[ r.label ]]
[[ r.t ]]
[[ r.s ]]
Teachers & leaders (n=167)Students (n=653)
Asked what schools should prioritise over the next two to three years, teachers put teaching acceptable AI use first (58%), clearer policies second (45%) and redesigned tasks third (39%). Better detection tools came last of nine options (12%).
The method hardest to outsource
Two minutes, laptop closed, three questions about the work.
A brief live explanation provides unusually strong evidence of attribution and understanding. It samples narrowly, so the questions asked and the decision rules applied need to be consistent across a faculty.
The cost, stated plainly
None of this is free. Oral questioning, second samples and supervised checkpoints all consume teacher time, and a redesign that ignores that will not survive a term. The workable version is sampling rather than universal coverage: a few students per task, a short set of questions used consistently, and a decision rule agreed across the faculty so different markers reach comparable judgements.
RESEARCH INSIGHT: ONE THING YOU CONTROL THIS TERM
Students who know the rules trust the marking more
Switch between students who say they understand what their school allows with AI, and students who do not.
[[ rFair ]]say their assessments give them a fair chance to show what they know
[[ rTrust ]]still say AI has eroded trust in take-home work
[[ rJudge ]]still say real ability is harder to judge from it
[[ rulesCaveat ]]
Learning check· Eight decisions
How strong is the evidence behind your marks?
Eight decisions. After every answer, the strongest option is revealed and your specific choice is explained. This is a short learning check, not a validated psychometric test.
[[ qDomain ]]
[[ qPrompt ]]
[[ qTerm ]][[ qExpTitle ]][[ qExpBody ]]
Your profile
Your reasoning pattern
This is formative feedback from eight decisions. It is not a normed assessment-literacy score.
[[ resultTitle ]]
[[ p.name ]][[ p.score ]]%
[[ resultMsgTitle ]] [[ resultMsgBody ]]
Three questions, before you redesign anything
01What exactly do I want to be able to say about this student?
02What evidence would let me say it honestly?
03What else could be producing this result?
Evidence base
Sources and scope notes
Definitions come from standard educational measurement sources. AI studies are labelled by context, because findings from professional writing or university assessment are not automatic claims about school students.
Gradeo (2026). Measure of Learning: Insights from the CSSA and ITE Online Trial HSC Examinations 2026. Evaluation survey, 21 August – 4 September 2026: 653 students (Years 7–12, 188 schools) and 167 teachers and school leaders across Government, Catholic and Independent sectors in NSW, plus four focus groups. Quotes lightly edited for fillers and de-identified.
Editorial note: the 44-student class, the Yr 9 energy task and both photographed English books are synthetic examples created for this page. Gradeo survey percentages and quotations come from source 15; experimental figures are cited separately at the point of use.
Future
Explore the Future Today
Book a demo to see how Gradeo can help create, deliver, mark and analyse assessments in one secure platform.