evidoria

← Back to browse

Good practice Imported

AI-Powered Oral Exams — NYU Stern's Cheat-Resistant Alternative to Written Assessment

United States of America · New York City · See the United States of America profile · See the New York City profile

Evidence: Descriptive / self-reported Top 91% 27/100 · Ask Evidence Copilot about this practice

NYU Stern's Panos Ipeirotis replaced written assignments in an AI/ML Product Management course with AI-run oral exams via ElevenLabs voice agents, costing 42 cents per student versus roughly $750 in TA wages, with 70% of students calling it a more valid test of understanding.

15 USD
Total exam cost (one cohort)
0.42 USD
Cost per student (one cohort)
750 USD
Estimated cost for human graders (one cohort)
36 students
Students in pilot cohort (one semester)
83 %
Students rating exam more stressful than written test (one cohort)
70 %
Students agreeing it better measured understanding (one cohort)
13 %
Students preferring a human examiner (one cohort)
29 %
Exact agreement across 3 AI graders (one cohort)
60 %
Agreement within one point across 3 AI graders (one cohort)
3.39 out of 4
Average score, problem framing (one cohort)
1.94 out of 4
Average score, experimentation (one cohort)
AI-Powered Oral Exams — NYU Stern's Cheat-Resistant Alternative to Written Assessment

Details

Maturity
Pilot
Promoter
NYU Stern School of Business (Panos Ipeirotis, Konstantinos Rizakos)
Period
2025-2026
Keywords
education, higher education, assessment technology

Context

At NYU Stern School of Business, professor Panos Ipeirotis noticed that written submissions in his 'AI/ML Product Management' course had started reading like polished consulting memos; when he probed authors with follow-up questions, many could not explain their own submitted work.

Objectives

Test whether students actually understood what they had submitted, without relying on AI-detection software, by building an AI-run oral exam with co-instructor Konstantinos Rizakos.

Activities

Built on ElevenLabs' conversational voice AI alongside Anthropic Claude, Google Gemini and OpenAI models, the system reads each student's submitted project in advance and dynamically tailors questions to their specific claims. It was tested across a 36-student cohort over nine days, averaging 25 minutes per student.

Results

The exam cost $15 in total, about 42 cents per student, versus an estimated $750 for two human graders at $25/hour. In student surveys, 83% rated the AI oral exam more stressful than a written test, yet 70% agreed it better measured their real understanding, and only 13% said they would prefer a human examiner. Three independent AI models grading the same responses agreed exactly 29% of the time and were within one point 60% of the time. Students scored higher on 'problem framing' (3.39/4 average) than on 'experimentation' (1.94/4), with three unable to discuss their own experimentation choices at all.

Conclusions

This is a single-course pilot self-reported by the instructor via NYU's student newspaper and independent tech press, not an independently audited or peer-reviewed study, so its small, one-off scale should temper expectations of generalisability.

Implementation

Indicative cost
Low (< €50k) — Stated actual cost of $15 total ($0.42/student) for the AI exam versus an estimated $750 for human TA grading.
Time to results
Short (< 1 year) — Exams administered over a nine-day window within one semester.
Staffing & skills
Panos Ipeirotis (instructor), Konstantinos Rizakos (co-instructor), NYU Stern School of Business

Conditions for success

  • AI reads each student's own submission in advance and tailors follow-up questions to their specific claims
  • Multiple independent AI models used for grading to expose reliability limits

Common failure modes

  • Only 29% exact grader agreement across three AI models, a real reliability concern
  • Small, one-off, single-course pilot (n=36), self-reported and not independently audited
  • No formal data-privacy or AI-governance framework described for recorded student voice data

Where it fits

Governance type
single instructor, single course
Scale
36 students, one semester
Income level
high-income

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful