Preview: we are still checking facts before launch.

Classroom practice

For teachers

AI and maths tests: the teacher writes, the software marks

The best split for a maths test is simple. You write the test, because you know what you taught, how you taught it and what you need to find out. Software helps after that: it computes the answer key exactly, takes the routine work out of marking, and groups wrong answers so you can see the misconceptions behind them. Your judgement stays in charge of the test and of every mark, and you get time back for the part of the job that needs you.

Why the teacher should write the test

A test is a set of decisions. You choose which skills matter this unit, which question shows real understanding rather than a remembered routine, and which wording your students will recognise because you used it in class. Nobody outside your classroom knows those things as well as you do.

School leaders see this too. At a school AI planning meeting, one principal put it plainly: “We still want teachers writing the maths tests themselves.”

That does not mean you type every question alone. You might borrow an item from a colleague, adapt one from a textbook or ask an AI tool for a few ideas. What matters is that you choose each question, check it and own it.

What software does well

Once the test exists, a lot of the work that follows is careful, repetitive checking. That is where software earns its place. It can work out the correct answer to each question, compare every student’s answer with it, and sort the class by which questions they got wrong and how.

A score tells you how many questions a student got right. A pattern in the wrong answers tells you what to teach on Tuesday.

Why a language model should not write the answer key

A language model, the kind of AI behind most chat tools, writes the text that is most likely to come next. It does not calculate the way a calculator does. It can set out neat working and still reach a wrong answer, and nothing on the page will warn you. A wrong answer key is worse than no answer key, because you mark against it with confidence.

Maths answers are exact, so the key should be computed exactly: by a calculator, a spreadsheet formula, or the computer algebra system inside good maths software. If a tool uses AI to suggest questions, ask how it produces the answers. You want to hear that it calculates them, not that the AI writes them.

What misconception analysis looks like on a Year 8 test

Say your Year 8 test includes this question: solve 3(x − 4) = 2x + 5. The correct answer is x = 17. You can check it: 3 × 13 = 39, and 2 × 17 + 5 = 39.

Now look at two common wrong answers. A student who writes x = 9 has probably expanded the bracket as 3x − 4, multiplying only the first term. A student who writes x = −7 expanded correctly to 3x − 12, then moved the 12 to the other side without changing its sign.

Those are two different misconceptions, and they need two different reteaching moments. Software can group the students who gave each answer across the whole test, so you can see that, for example, a cluster of students made the bracket error on several questions. You then look at their working to confirm the pattern before you act on it. The software can point to a pattern, and you decide what it means.

Keep student details out of the marking step

Marking with software means student work goes into that software. That is the moment to slow down.

In Victorian government schools, the department’s generative AI policy (read 25 September 2026) requires schools to direct staff and students to keep student names and other personal information, and sensitive school information such as student assessment data, out of generative AI tools. It also says staff must be directed not to use generative AI tools to directly make judgements about student learning achievement or progress. So the final call on each mark stays with you. How these rules apply to marking software that computes answers without generative AI is a question for your school’s leaders. Marking software that does not use generative AI still falls under the department’s Software and Administration Systems policy (read 25 September 2026), which lists assessment or grading tools among the systems it covers. Check for an ST4S report in Arc before you use it. Whether a marking tool that does use generative AI can ever take student work is still [CHECK CURRENT GUIDANCE: VIC]. Other schools: check your sector’s or state’s rules.

In practice, find out first whether the marking software uses generative AI. If it does, a Victorian government school should not upload student work or results to it. A student number in place of a name does not change that, because the number still identifies the student. If it does not use generative AI, check that your school has approved it and where the work is stored before you upload a class set. Our guide to where student data goes lists the questions to ask.

What to do on Monday

Write your next test as you normally would. Work out the answer key with a calculator or computer algebra system, or confirm that your software computes it rather than generating it with AI.

For two or three key questions, write down the wrong answers you expect and what each one would tell you. This takes a few minutes and makes any analysis, by hand or by software, far more useful.

After marking, sort the wrong answers into patterns before you record the scores. Plan one short reteach for each pattern you find.

Before any marking tool touches student work, find out what its vendor says about where student work goes. As we publish tool checks, our tool directory shows this for each tool.