evidoria

← Back to browse

Good practice Imported

JudgeGPT — a Nationwide Randomized Trial of Generative AI for Pakistan's Trial-Court Judges

Pakistan · Lahore · See the Pakistan profile · See the Lahore profile

Evidence: Randomised controlled trial Top 15% 80/100 · Ask Evidence Copilot about this practice

A randomized trial gave ~1,559 Pakistani trial-court judges (about half the bench) access to JudgeGPT, a GPT-4 tool searching 129,235 rulings and laws. Trained judges' districts resolved 6.3% more cases, rulings rated better 59% of the time, and no rise in bias.

1559
Judges covered by the trial
118
Courts covered
129235 documents (128,292 rulings + 943 laws)
Documents in the retrieval database
6.3 %
Increase in cases resolved
1848
Additional cases resolved per year
616
Additional cases resolved in the lowest-performing quartile
59 % of pairwise comparisons
Rulings rated better (trained judges) (vs 42% baseline)
42 %
Baseline share of rulings rated better
10-38.50 $ per $ invested (authors' estimate)
Estimated return on investment
2.26 million
Pending cases in Pakistan's trial courts (end of 2024)

Details

Maturity
Pilot
Promoter
Pakistan trial court judiciary, with ETH Zurich, Imperial College London and the New Economic School
Period
2025–2026
Keywords
judiciary, case backlog, generative AI, legal research

Context

Pakistan's trial courts faced 2.26 million pending cases at the end of 2024, with fewer than two judges per 100,000 residents compared with 22 in the EU, prompting researchers from ETH Zurich, Imperial College London and the New Economic School to build JudgeGPT, a GPT-4-based retrieval-augmented assistant drawing on 129,235 documents (128,292 court rulings and 943 Pakistani laws).

Objectives

The team designed a randomized controlled trial covering 1,559 judges across 118 courts - roughly half of Pakistan's trial-court bench - randomised into a JudgeGPT-plus-targeted-training arm, a JudgeGPT-plus-general-seminar arm, and a seminar-only control arm.

Activities

Trained judges used the tool far more intensively than the comparison arms, logging around 60 logins and more than 200 prompts each, versus about 20 logins and under 50 prompts in the other arms.

Results

Over the 40 weeks following the intervention, districts with more trained judges resolved about 1,848 additional cases per year - a 6.3% increase - with the largest effect, about 616 extra cases, in the lowest-performing quartile of districts; blind quality assessments rated trained judges' rulings as better in 59% of pairwise comparisons, up from a 42% baseline, and the study found no evidence that AI assistance increased gender or religious bias in judicial language.

Conclusions

The research team's own cost-benefit estimate puts returns at roughly $10-38.50 per dollar invested, though that figure comes from the study's authors rather than an independent government audit, and the programme remains a large-scale research trial rather than a codified permanent institution, with Pakistan's Supreme Court signalling openness to broader judicial AI use while noting the absence of a comprehensive legal framework governing it.

Implementation

Indicative cost
High (€500k–€5M)
Time to results
Medium (1–3 years)
Staffing & skills
A research team from ETH Zurich, Imperial College London and the New Economic School worked directly with Pakistan's trial-court judiciary to design and run the trial.

Conditions for success

  • Judges in the training arm received targeted training rather than only a general seminar, driving far higher tool usage (about 60 logins and 200+ prompts vs ~20 logins and under 50 prompts).
  • A randomised trial design across a large share of the national judiciary allowed a causal comparison between arms.

Common failure modes

  • The programme remains a research trial rather than a codified permanent institution.
  • Pakistan lacks a comprehensive legal framework governing judicial AI use, per the Supreme Court's own comments.
  • The cited $10-38.50 per dollar return is the research team's own estimate, not an independently audited figure.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful