evidoria

← Back to browse

Good practice Imported

Generative AI Evaluation Sandbox — Singapore's IMDA/AI Verify Testbed for Trusted LLM Deployment

Singapore · Singapore · See the Singapore profile · See the Singapore profile

Evidence: Descriptive / self-reported Top 66% 53/100 · Ask Evidence Copilot about this practice

Singapore's IMDA and AI Verify Foundation launched a Generative AI Evaluation Sandbox on 31 October 2023, with 16 organisations — Google, Microsoft, AWS, Singtel, OCBC, Deloitte, EY and regulator PDPC — testing a shared evaluation catalogue for large language models.

16 organisations
Organisations participating at launch (31 October 2023)
11 principles
Principles in the testing framework
Generative AI Evaluation Sandbox — Singapore's IMDA/AI Verify Testbed for Trusted LLM Deployment

Details

Maturity
Scaling
Promoter
Infocomm Media Development Authority (IMDA) & AI Verify Foundation
Period
2023-present
Keywords
AI governance, digital economy regulation, technology assurance

Context

On 31 October 2023, Singapore's Infocomm Media Development Authority (IMDA) and the independent AI Verify Foundation launched a Generative AI Evaluation Sandbox, billed as a first-of-its-kind controlled testing environment for large language model (LLM) applications before wider deployment.

Activities

Sixteen organisations confirmed participation at launch: model developers Google, Microsoft and Amazon Web Services; app developers Singtel and OCBC testing specific use cases; external testers Deloitte and EY; and Singapore's Personal Data Protection Commission (PDPC) as observing regulator. Each sandbox use case pairs an upstream model developer, a downstream application deployer and a third-party tester, so evaluation happens across the AI supply chain rather than by a single actor self-certifying. The sandbox is built around a shared Evaluation Catalogue setting out baseline testing methods, aimed at surfacing evaluation gaps in domain-specific contexts and languages and generating input for technical testing standards.

Results

According to IAPP's later reporting, the testing framework spans 11 principles (including transparency, safety, repeatability, fairness and human oversight) aligned to the NIST AI Risk Management Framework, the Hiroshima Process International Code of Conduct and ISO/IEC 42001, and by mid-2025 Standard Chartered had piloted a generative-AI application through the sandbox with independent third-party evaluation. IMDA and the AI Verify Foundation have publicised the participant list, framework and objectives, but no independently published data yet reports pass/fail rates, vulnerabilities found, or measurable changes to deployed systems as a result of sandbox testing.

Implementation

Indicative cost
Medium (€50k–€500k) — Not quantified in sources; run jointly by a statutory regulator and an independent foundation with participating firms bearing their own testing costs.
Time to results
Medium (1–3 years) — Launched 31 October 2023 with 16 organisations; framework and participation reported as continuing through at least mid-2025 (Standard Chartered pilot).
Staffing & skills
IMDA (statutory regulator) and the AI Verify Foundation (independent non-profit) jointly run the sandbox, Singapore's PDPC observes as regulator; Deloitte and EY act as third-party testers

Conditions for success

  • A supply-chain evaluation design pairing a model developer, an application deployer and a third-party tester for each use case, rather than self-certification
  • Explicit alignment of the testing framework to international standards (NIST AI RMF, Hiroshima Process, ISO/IEC 42001) to support cross-border credibility

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful