Jisc's National AI in Marking and Feedback Pilot — 38 UK Colleges and Universities
United Kingdom
Jisc ran a year-long national pilot of AI marking tools (Graide, Keath, TeacherMatic) across 38 UK colleges and universities, finding …
United States of America · New York City · See the United States of America profile · See the New York City profile
Evidence: Descriptive / self-reported Top 89% 27/100 · Ask Evidence Copilot about this practice
NYU Stern's Panos Ipeirotis replaced written assignments in an AI/ML Product Management course with AI-run oral exams via ElevenLabs voice agents, costing 42 cents per student versus roughly $750 in TA wages, with 70% of students calling it a more valid test of understanding.
At NYU Stern School of Business, professor Panos Ipeirotis noticed that written submissions in his 'AI/ML Product Management' course had started reading like polished consulting memos; when he probed authors with follow-up questions, many could not explain their own submitted work.
Test whether students actually understood what they had submitted, without relying on AI-detection software, by building an AI-run oral exam with co-instructor Konstantinos Rizakos.
Built on ElevenLabs' conversational voice AI alongside Anthropic Claude, Google Gemini and OpenAI models, the system reads each student's submitted project in advance and dynamically tailors questions to their specific claims. It was tested across a 36-student cohort over nine days, averaging 25 minutes per student.
The exam cost $15 in total, about 42 cents per student, versus an estimated $750 for two human graders at $25/hour. In student surveys, 83% rated the AI oral exam more stressful than a written test, yet 70% agreed it better measured their real understanding, and only 13% said they would prefer a human examiner. Three independent AI models grading the same responses agreed exactly 29% of the time and were within one point 60% of the time. Students scored higher on 'problem framing' (3.39/4 average) than on 'experimentation' (1.94/4), with three unable to discuss their own experimentation choices at all.
This is a single-course pilot self-reported by the instructor via NYU's student newspaper and independent tech press, not an independently audited or peer-reviewed study, so its small, one-off scale should temper expectations of generalisability.
Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.
Where this practice's information was retrieved from, and when.
United Kingdom
Jisc ran a year-long national pilot of AI marking tools (Graide, Keath, TeacherMatic) across 38 UK colleges and universities, finding …
Dominican Republic
SARA pairs a validated Spanish reading battery with AI-automated scoring on a tablet. Validated on 1,860 Santo Domingo pupils, its …
France
The Île-de-France region piloted the Ed.ai tool in 20 high schools, proposing AI-generated grades and remediation exercises for 7,600 exam …
Australia
Instead of an unwinnable AI-detection arms race, Sydney redesigned assessment into two lanes — secure in-person assessment of core capability, …
Open full copilot Grounded in cited practices — always check the sources.