evidoria

← Back to browse

Good practice Imported

Independent Audit of Fobizz's 'AI Grading Assistant' — Osnabrück Study Finds Erratic Scores in a Tool Licensed by German States

Germany · Osnabrück · See the Germany profile

Top 56% 47/100 · Ask Evidence Copilot about this practice

Mühlhoff and Henningsen (Osnabrück) tested Fobizz's AI grading tool in two test series: scores and feedback looked arbitrary, one text ranged 1–14 points across runs, and a 2026 follow-up on two other tools found similar variation.

Independent Audit of Fobizz's 'AI Grading Assistant' — Osnabrück Study Finds Erratic Scores in a Tool Licensed by German States

Details

Promoter
University of Osnabrück (Mühlhoff & Henningsen) — audit of Fobizz
Period
2024-2026
Keywords
education, assessment, automated grading, algorithmic audit, schools

Description

In December 2024 Rainer Mühlhoff and Marte Henningsen (University of Osnabrück) published, on arXiv (2412.06651) and at the 38th Chaos Communication Congress, an independent test of the 'AI Grading Assistant' sold by the German company Fobizz, a tool licensed to schools in several German states (Rhineland-Palatinate, Mecklenburg-Western Pomerania, Saxony). Fobizz itself reports more than 500,000 teachers at over 7,500 schools.
Using two test series, the authors found that grades and written feedback often appeared random and did not reliably improve when the tool's suggestions were applied; only ChatGPT-generated texts received the top ratings; false claims and nonsensical submissions often went unnoticed; and some criteria were applied opaquely. In the example reported by heise, a single text received between 1 and 14 points across repeated runs. The authors argue that the vendor's 'objective and time-saving' marketing is misleading and call for systematic, subject-specific pedagogical evaluation before such tools enter schools. Fobizz adjusted its prompt after publication but did not share the original prompt.
A 2026 follow-up by Mühlhoff and Quägwer (SEMINAR – Lehrerbildung und Schule 2/2026) tested FelloFish and Edaira with simulated student texts and again found considerable variation for identical inputs, poor reproducibility and a tendency to rate verbatim adoption of suggestions higher than independent revision; they conclude such tools should only support teachers under critical human supervision.
Caveats: the studies use simulated texts and do not measure learning outcomes; the exact number of texts and runs is in the papers' appendices and was not independently re-verified here. A small independent pilot (IADIS) found that three experienced teachers also changed at least one grade between rounds in 73% of cases, so human grading is not a flawless benchmark either. The practice recorded here is the audit itself: an open, replicable method for testing AI assessment tools before procurement.

Read the full analysis: https://rainermuehlhoff.de/en/fobizz-AI-grading-assistant-test-studie/

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful