Singapore's IMDA and AI Verify Foundation launched a Generative AI Evaluation Sandbox on 31 October 2023, with 16 organisations — Google, Microsoft, AWS, Singtel, OCBC, Deloitte, EY and regulator PDPC — testing a shared evaluation catalogue for large language models.
16 organisations
Organisations participating at launch (31 October 2023)
11 principles
Principles in the testing framework
Details
Maturity
Scaling
Promoter
Infocomm Media Development Authority (IMDA) & AI Verify Foundation
Period
2023-present
Keywords
AI governance, digital economy regulation, technology assurance
Context
On 31 October 2023, Singapore's Infocomm Media Development Authority (IMDA) and the independent AI Verify Foundation launched a Generative AI Evaluation Sandbox, billed as a first-of-its-kind controlled testing environment for large language model (LLM) applications before wider deployment.
Activities
Sixteen organisations confirmed participation at launch: model developers Google, Microsoft and Amazon Web Services; app developers Singtel and OCBC testing specific use cases; external testers Deloitte and EY; and Singapore's Personal Data Protection Commission (PDPC) as observing regulator. Each sandbox use case pairs an upstream model developer, a downstream application deployer and a third-party tester, so evaluation happens across the AI supply chain rather than by a single actor self-certifying. The sandbox is built around a shared Evaluation Catalogue setting out baseline testing methods, aimed at surfacing evaluation gaps in domain-specific contexts and languages and generating input for technical testing standards.
Results
According to IAPP's later reporting, the testing framework spans 11 principles (including transparency, safety, repeatability, fairness and human oversight) aligned to the NIST AI Risk Management Framework, the Hiroshima Process International Code of Conduct and ISO/IEC 42001, and by mid-2025 Standard Chartered had piloted a generative-AI application through the sandbox with independent third-party evaluation. IMDA and the AI Verify Foundation have publicised the participant list, framework and objectives, but no independently published data yet reports pass/fail rates, vulnerabilities found, or measurable changes to deployed systems as a result of sandbox testing.
Implementation
Indicative cost
Medium (€50k–€500k) — Not quantified in sources; run jointly by a statutory regulator and an independent foundation with participating firms bearing their own testing costs.
Time to results
Medium (1–3 years) — Launched 31 October 2023 with 16 organisations; framework and participation reported as continuing through at least mid-2025 (Standard Chartered pilot).
Staffing & skills
IMDA (statutory regulator) and the AI Verify Foundation (independent non-profit) jointly run the sandbox, Singapore's PDPC observes as regulator; Deloitte and EY act as third-party testers
Conditions for success
A supply-chain evaluation design pairing a model developer, an application deployer and a third-party tester for each use case, rather than self-certification
Explicit alignment of the testing framework to international standards (NIST AI RMF, Hiroshima Process, ISO/IEC 42001) to support cross-border credibility
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.