When teachers say an AI grader "got it wrong," the cause is usually not the AI. It's a rubric criterion that two human teachers would also have scored differently. The good news is that the fixes are simple, and they make your rubric better for students too.
Start with the test that matters
Ask yourself: Could a substitute grade this stack with only my rubric, without asking me a single question?
If the answer is no, an AI will run into the same gaps. Everything below is about closing them.
1. Use an analytic rubric
Give each criterion its own score: thesis, evidence, organization, conventions, and so on. Analytic rubrics help AI grading in two ways:
- The model gets one focused job per criterion instead of a fuzzy overall judgment.
- You get criterion-level scores and comments that you can scan and verify in seconds.
Holistic rubrics ("4 = strong overall response") are faster to write but much harder to audit.
2. Describe observable evidence, not impressions
This is the single biggest improvement you can make.
| Vague | Observable |
|---|---|
| Strong thesis | States a specific, arguable position in one sentence in the first paragraph |
| Good use of evidence | Includes at least two relevant quotations or data points, each explained in the student's own words |
| Well organized | Each paragraph has one main idea, and transitions connect each paragraph to the thesis |
| Few errors | Errors in spelling or grammar do not interfere with meaning; no more than 3 per page |
The right-hand column turns scoring into matching evidence against a description. People and AI both do that consistently. For more examples, see writing rubrics that actually speed up grading.
3. Write levels as degrees of the same thing
Every level of a criterion should describe the same dimension at different strengths. A common mistake is for the top level to talk about "voice" while the bottom level talks about "length." Then a response that has a strong voice but is short fits no level cleanly.
A good pattern for an "Evidence" criterion:
- 4: Two or more relevant pieces of evidence, each clearly explained and connected to the claim.
- 3: Two relevant pieces of evidence, but at least one isn't explained or connected.
- 2: One relevant piece of evidence, or evidence that is only loosely related.
- 1: No relevant evidence.
4. Keep it to 4–6 criteria and 3–4 levels
More criteria rarely change the final grade but add review time. More than four levels creates distinctions that no one applies consistently. If two criteria always get the same score, merge them.
5. Put points on every level
AI graders return numbers, and your gradebook needs numbers. Give each level an explicit point value, and make sure the criterion maximums add up to the assignment total you expect in Google Classroom. Mismatched totals are the most common reason a grade looks "off" after import.
6. Add the assignment prompt and a short note on what matters most
If the rubric is generic ("Argument writing rubric"), tell the grader what this assignment asked for: the prompt, the text students read, or the required length. One or two sentences of context noticeably improves the relevance of the feedback.
7. Pilot on three papers before the whole stack
Grade one strong, one middle, and one weak submission yourself, then run the AI on the same three. Where you disagree, read the criterion again. Nine times out of ten, the wording allowed both interpretations. Fix the wording, not the AI, and the rest of the stack will line up.
This also works as a quick calibration exercise if you share rubrics with a department, much like the norming sessions described in standards-based grading explained.
A ready-to-adapt example
Short constructed response (10 points)
- Claim (3 pts): 3 = directly answers the question in a clear sentence; 2 = answers the question but vaguely; 1 = restates the question without answering; 0 = missing.
- Evidence (4 pts): 4 = two relevant text details, accurately cited; 3 = two details, one inaccurate or not cited; 2 = one relevant detail; 0 = none.
- Reasoning (3 pts): 3 = explains how each detail supports the claim; 2 = explains one detail; 1 = lists details without explanation; 0 = missing.
That's specific enough for a substitute, and specific enough for an AI.
Using your rubric with ClassGrade Pilot
In ClassGrade Pilot you can upload a rubric as a PDF or Word document. It's converted into criteria, levels, and points that you can reuse across sections. Then grade a Google Classroom assignment from the side panel and review every proposed score before approved grades go back to Classroom. New to the idea? Start with our practical guide to AI rubric grading.