SproutSprout
← Back to blog
parentsAustraliaVictoriaproduct

Why a Chatbot Cannot Mark Your Child’s Work

24 September 2026 · 9 min read · Dominic Mazur

Add Sprout as a preferred source on Google

When a chatbot tells your child they are right, a language model formed an opinion. When a marked question tells them they are right, a stored answer was compared to what they typed. Those are not the same event, and the difference shows up on exactly the questions where it matters most.

This is the least glamorous part of any tutoring product and the part that decides whether the record is worth anything. A score built from a model’s judgement is a score you have to trust. A score built from a key is one you can check.

The structural problem, not a policy one

Two lanes either side of a dashed line dividing the server from the child's screen. In the chatbot lane, the model which knows the answer sends a question and answer to the screen, marked with a key symbol and noted as one prompt away from the answer. In the Sprout Tutor lane, the server holds the question and answer key and sends the question only, noted as the key stays behind on the server.
A chatbot cannot withhold the answer while your child thinks, because the thing holding the answer is the thing they are talking to.

A chatbot has the answer available to your child for the entire time they are supposed to be working it out. Not because it is badly behaved, but because there is only one conversation and the model is in it. Any restraint is a rule it has been asked to follow, and children are persistent, inventive and have the whole internet full of phrasings that get around rules.

A marked question is a different shape. The question is generated with its answer key attached, and the key is stripped on the server before the question is sent to the browser. For a multiple-choice question the correct option id is removed; for a fill-in-blank the list of accepted answers; for an estimate the acceptable range; for a spot-the-error the id of the wrong step. Your child’s device never receives it, so there is nothing to talk it out of.

Ten ways to ask, ten ways to mark

The other thing a chat box cannot do is vary the shape of the question, because a chat box only has one input: type something. Sprout Tutor has ten question types, and each one has its own marking function:

  • Multiple choice and true or false, the familiar two.
  • Fill in the blank, marked against a list of accepted answers.
  • Multi-select, where every correct option must be found and no wrong one chosen, so a guess is expensive.
  • Estimate in range, which is marked as a band rather than a single value, because estimation is a skill and grading it to the exact digit would be grading the wrong thing.
  • Ordering, where the sequence must be exactly right.
  • Matching and categorise, where every pair or every item must land correctly.
  • Spot the error, which shows a worked solution with one wrong step and asks which one. This is the inverse of asking for the answer, and it is the type a chat box is worst at.
  • Multi true or false, several statements at once, all of which must be judged correctly.

Six of those ten are effectively guess-proof. There are 120 ways to order five items and one of them is right. A child who does not know cannot bluff their way through ordering, matching or categorising the way they can through a four-option multiple choice, which is why a record built only on multiple choice tells you less than it looks like it does.

Being marked wrong for the wrong reason

The obvious objection to automatic marking is that it is pedantic, and it usually is. A typed answer goes through a normaliser first, so the following are all accepted rather than counted wrong:

  • .525 and 0.525, because a missing leading zero is a typing habit, not a misconception.
  • “left over” and “leftover”, and anything else that differs only by a hyphen between letters.
  • A trailing per cent sign or degree symbol, included or not.
  • Capitals, spacing and trailing punctuation, none of which change an answer.

Two things deliberately stay different. Minus three is not three, because the hyphen there is a minus sign and not a hyphenation. And 5 never becomes 0.5, because a bare digit gaining a leading zero would turn a wrong answer into a right one. Those two exceptions are the whole reason the normaliser is a written rule rather than a model deciding case by case.

The one line that does the marking

In the code, a child’s attempt is marked by comparing their selection against the stored data for that question, and a revealed answer can never come out correct: revealing is checked first, so asking to be shown the answer and then entering it is recorded as a reveal, not as a right answer.

That is worth a sentence on its own, because it is the thing that keeps the record honest. If revealing counted as correct, accuracy would measure persistence at clicking. Because it does not, the adult view can show accuracy, hints used and answers revealed as three separate numbers, and two children with the same score can be told apart.

What deterministic marking cannot do

The honest limit, and it is a real one. A stored key can only mark a question with a bounded answer. It cannot mark a paragraph, an argument, a method written out longhand, or a child’s explanation of why something works. Those are exactly the things a good teacher marks and exactly the things a curriculum asks for at the top end.

A chatbot will attempt all of them, and sometimes do it well. It will also be confidently wrong occasionally, in a register that makes the error hard for a child to catch, and you will not know which occasion was which. That trade is the actual choice, and anyone telling you it is free is selling something.

Our position is that for daily practice, where the point is retrieval and the parent needs a record they can trust, narrow and checkable beats broad and plausible. For an essay, it does not.

What to ask any product

  1. Does the answer reach the device before my child answers? If yes, there is nothing stopping a determined child, whatever the instructions say.
  2. Is a right answer decided by a key or by a model? Both are legitimate; only one is reproducible.
  3. Is asking for the answer recorded, and does it count against the score? If revealing is free, the accuracy number is decoration.
  4. How many question types are there? If the answer is one or two, the record is mostly measuring guessing.
  5. What happens to a formatting difference? A product that marks 0.525 wrong because your child typed .525 will teach them to distrust it inside a week.

Where we sit, declared

Sprout Tutor went into public beta on 22 September 2026, covering the Victorian Curriculum only, Years 1 to 9, free on every tier while in beta. It is not on the pricing page and there is nothing to buy. Everything described above is how it works rather than a roadmap, and everything it cannot do is listed above as well.

The short version

A chatbot cannot withhold the answer while your child thinks, because the thing that holds the answer is the thing they are talking to. Any restraint is a rule to be argued with. A marked question strips the key on the server, so nothing to argue with ever arrives.

Ten question types, each with its own marking rule, six of them effectively guess-proof, and a normaliser so that a missing leading zero or a stray hyphen is never counted as a wrong answer. Minus three still is not three, and 5 still is not 0.5.

A revealed answer can never be marked correct, which is what makes accuracy, hints and reveals three numbers worth reading rather than one number worth doubting. And the limit is honest: a stored key cannot mark an essay, and a chatbot can attempt one.

Sprout Lessons builds curriculum-aligned lessons for Foundation to Year 10, with the code on every lesson. Start free.

Checked 24 September 2026 against the product code rather than marketing copy: the ten interaction types and their marking functions, the server-side removal of each type’s answer key before a question is sent to the browser, the normalisation rules and their two deliberate exceptions, and the rule that a revealed answer cannot be marked correct. Sprout Tutor is in public beta, covers the Victorian Curriculum only for Years 1 to 9, and is free on every tier while in beta. Other products work differently: ask them the five questions above rather than assuming this page describes them.

FAQ

Can an AI chatbot mark my child’s work reliably?

It can form an opinion about whether an answer is right, and it will often be correct. What it cannot do is withhold the answer while your child is working, because there is only one conversation and the model holding the answer is the thing your child is talking to. Any restraint is a rule the child can argue around, and children are persistent and inventive.

How does Sprout Tutor decide whether an answer is right?

By comparing the child’s selection against the stored data for that question, with one marking function per question type. There are ten types and ten marking functions. The model writes the question; it does not judge the answer. Revealing is checked first, so a revealed answer can never be recorded as correct.

What question types are there besides multiple choice?

Ten in total: multiple choice, true or false, fill in the blank, multi-select, estimate in range, ordering, matching, categorise, spot the error, and multi true or false. Estimate in range is marked as a band rather than an exact value, because grading an estimate to the digit would grade the wrong skill. Spot the error shows a worked solution with one wrong step and asks which, which is the inverse of asking for the answer.

Will my child be marked wrong for typing an answer differently?

Not for formatting. Answers are normalised first, so .525 and 0.525 match, left over and leftover match, a trailing per cent or degree sign is ignored, and capitals, spacing and trailing punctuation do not count. Two differences are kept on purpose: minus three is not three, because that hyphen is a minus sign, and a bare 5 never becomes 0.5, because that would turn a wrong answer into a right one.

What can deterministic marking not do?

It can only mark a question with a bounded answer. It cannot mark a paragraph, an argument, a method written out longhand, or a child explaining why something works, and those are exactly what a good teacher marks. A chatbot will attempt all of them and sometimes do it well, and will also be confidently wrong occasionally in a register a child cannot catch. That trade is the real choice, and anyone saying it is free is selling something.

Which curriculum does Sprout Tutor cover?

The Victorian Curriculum only, Years 1 to 9. It went into public beta on 22 September 2026 and is free on every tier while in beta, and it is not on the pricing page, so there is nothing to buy. It does not yet follow the Australian Curriculum, the NSW syllabuses or any North American framework.

Written by

Dominic Mazur

Dominic Mazur is the founder of Sprout Lessons.

More about the people behind Sprout →

Build a lesson around what your students love

Sprout turns any topic and a student’s interests into an interactive, standards-aligned lesson in seconds. The free plan gives you credits every month, no card needed.