evidoria

← Back to browse

Good practice Imported

URAG — Ho Chi Minh City University of Technology's Hybrid RAG Admission Chatbot Outperforms GPT-4o on Real Applicant Questions

Vietnam · Ho Chi Minh City · See the Vietnam profile · See the Ho Chi Minh City profile

Top 82% 33/100 · Ask Evidence Copilot about this practice

HCMUT built URAG, a two-tier retrieval-and-rules chatbot for admissions questions, that scored 62.8% accuracy on 500 real applicant queries versus 54.4% for GPT-4o, 50.2% for Claude 3.5 Sonnet and 44.0% for Gemini 1.5 Pro.

URAG — Ho Chi Minh City University of Technology's Hybrid RAG Admission Chatbot Outperforms GPT-4o on Real Applicant Questions

Details

Promoter
Ho Chi Minh City University of Technology (HCMUT) — URA Research Group
Period
2024-2025
Keywords
higher education, university admissions, generative AI, retrieval-augmented generation

Description

Vietnam's Ho Chi Minh City University of Technology (HCMUT) faced a familiar risk for university admissions chatbots: large language models can hallucinate on sensitive, time-bound questions such as application deadlines and program requirements, where a wrong answer has real consequences for applicants.
HCMUT's URA research group built URAG, a two-tier hybrid system that first checks a curated FAQ bank (built and paraphrased from official admissions documents via a preparation step called URAG-F) and falls back to retrieval-augmented generation, drawing on semantically chunked source documents (URAG-D), only when no FAQ match is found.
The chatbot was deployed live at ura.hcmut.edu.vn for a four-month admissions cycle, with usage peaking in early June and late August as Vietnamese high-school students approached exam and enrollment deadlines. Researchers then benchmarked it against 500 real admission questions collected from high-school applicants: URAG answered 62.8% correctly, ahead of GPT-4o (54.4%), Claude 3.5 Sonnet (50.2%) and Gemini 1.5 Pro (44.0%). The system was recognised at Ho Chi Minh City's Week of Innovation, Startup and Entrepreneurship (WHISE) 2024.
The comparison also exposes a limit worth stating plainly: URAG's advantage was concentrated in factual, FAQ-style questions (218 of 268 correct); on questions requiring multi-step reasoning it underperformed the general-purpose commercial models. An accuracy ceiling around 63% also means roughly one in three answers was still wrong, so the tool is best read as a triage layer that reduces routine load rather than a full replacement for human admissions staff.

Read the full analysis: https://www.ura.hcmut.edu.vn/hcmut-chatbot-uras-ai-innovation-takes-the-spotlight-at-whise-2024/

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful