evidoria

← Back to browse

Good practice Imported

GenAI 24x7 Tutor — University of Wollongong's Human-Simulated Benchmark of ChatGPT, Wolfram GPT and Tutor Me GPT for Engineering and Maths

Australia · Wollongong · See the Australia profile

Top 95% 20/100 · Ask Evidence Copilot about this practice

University of Wollongong researchers had assistants role-play average and struggling students across 35 engineering/maths topics to benchmark ChatGPT-4, ChatGPT-4o, Wolfram GPT and Tutor Me GPT as tutors, finding frequent errors and near-zero accuracy on some mechanical topics.

GenAI 24x7 Tutor — University of Wollongong's Human-Simulated Benchmark of ChatGPT, Wolfram GPT and Tutor Me GPT for Engineering and Maths

Details

Promoter
University of Wollongong, Australasian Artificial Intelligence in Engineering Education Centre (AAIEEC)
Period
Second half of 2024 (data collection); published 21 January 2026
Keywords
engineering education, mathematics education, GenAI tutoring, accuracy benchmarking, higher education

Description

Researchers at the University of Wollongong's Faculty of Engineering and Information Sciences, working through the Australasian Artificial Intelligence in Engineering Education Centre (AAIEEC), tested whether ChatGPT-4, ChatGPT-4o, Wolfram GPT and Tutor Me GPT (Khanmigo Lite) could function as accurate, pedagogically sound 24/7 tutors for engineering and mathematics. Because their human research ethics committee would not approve trials with real students without prior evidence on risk, the team instead used three experienced research assistants (two PhD candidates, one doctorate holder) to role-play an average student and a consistently struggling student across 7 subjects and 35 topics in electrical engineering, mechanical engineering and mathematics.

Each of the roughly 280 sessions (35 topics × 4 tools × 2 student profiles) lasted at least 20 minutes and was scored 0–4 on accuracy and on six pedagogical-experience dimensions (relevance, pedagogical effectiveness, engagement, progression, contextual understanding, use of examples), using a rubric adapted from Merrill et al.'s human-vs-computer tutoring framework.

Accuracy was strong in electrical engineering and mathematics (models generally scored 3–4, meaning no more than one minor error) but weak in mechanical engineering, where average scores fell below 3 and, in one thermodynamics topic, dropped to 0–1.5; a water-pipe pressure-drop question was answered correctly in only 3 of 8 attempts. A Friedman test found no statistically significant difference in accuracy between the four tools. Tutor-experience scores were more consistently positive (lowest model average 2.97/4), with ChatGPT-4 rated the strongest overall tutor and Tutor Me GPT the weakest.

The authors are explicit that this was a simulation, not a trial with real students, and that it 'offers no direct evidence of GenAI's actual effectiveness in supporting learning.' They recommend against broader classroom implementation until accuracy improves and a supervised student trial is run, and disclose that the lead author sits on the journal's editorial board (and was recused from the editorial decision). The paper is open access (CC-BY-4.0) with a full appendix of the exact tutoring prompt used.

Read the full analysis: https://aimspress.com/article/doi/10.3934/steme.2026004

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful