How the quality score is computed
Every practice in Evidoria carries a single quality score from 0 to 100. It is calculated the same way in every catalogue, so scores are comparable and explainable.
The method, in one line
score = 100 × the (weighted) average of each evaluation dimension, after rescaling every dimension to 0–1.
- Each catalogue has a rubric: a set of evaluation dimensions (e.g. transferability, impact, sustainability).
- Each dimension's raw value is normalised to a 0–1 scale:
(value − min) / (max − min). - The normalised values are summed over the rubric's FULL set of dimensions (each weighted equally unless stated otherwise), divided by the rubric's total weight, and multiplied by 100. A dimension that has not been evaluated counts as zero — absence of evidence is not evidence of merit.
Because the arithmetic is identical everywhere, a score of 80 means the same thing in any catalogue: the practice meets, on average, 80% of its rubric's evaluation criteria. Catalogues differ only in which dimensions they assess, because their source data differs.
Reading a practice's score
A single number hides a lot. Most practices score high (often 90+), so the headline figure alone barely separates them. Two things on each practice page put the score in context.
- The score profile shows the per-dimension breakdown behind the score, so you can see where a practice is strong or weak rather than just its average.
- The catalogue position (e.g. “Top 10%”) shows where the practice sits among scored peers in the same catalogue. We count how many peers score at or below it (an “at or below” tie convention). It is shown only when a catalogue has at least five scored practices, so a position is meaningful.
Both are ways of reading the score — they never change the stored value. The score itself is always computed exactly as described above.
Dimensions assessed, by catalogue
AI in Education
Evaluated by C-NAPSE
| Dimension | Scale | Weight |
|---|---|---|
| Evidence of learning impact | 0–3 | 1 |
| Equity & inclusion | 0–3 | 1 |
| Transparency, ethics & data protection | 0–3 | 1 |
| Teacher capability & pedagogy | 0–3 | 1 |
| Scalability & sustainability | 0–3 | 1 |
AI in the Public Sector
Evaluated by C-NAPSE
| Dimension | Scale | Weight |
|---|---|---|
| Evidence of impact / measured public value | 0–3 | 1 |
| Transparency, fairness & accountability | 0–3 | 1 |
| Transferability / demonstrated replication | 0–3 | 1 |
| Scalability beyond pilot | 0–3 | 1 |
| Governance, capability & sustainability | 0–3 | 1 |
City Innovation Library
Evaluated by C-NAPSE
| Dimension | Scale | Weight |
|---|---|---|
| Evidence of impact / measured results | 0–3 | 1 |
| Transferability / demonstrated replication | 0–3 | 1 |
| Scalability | 0–3 | 1 |
| Sustainability / continuity | 0–3 | 1 |
| Innovation | 0–3 | 1 |
| Inclusiveness / equity | 0–3 | 1 |
| Multi-stakeholder collaboration | 0–3 | 1 |
Digital Inclusion
Evaluated by C-NAPSE
| Dimension | Scale | Weight |
|---|---|---|
| Innovation level | 0–5 | 1 |
| Sustainability (ongoing) | 0–1 | 1 |
| Evaluation evidence | 0–1 | 1 |
| Demonstrated reach | 0–1 | 1 |
| Documented learning | 0–1 | 1 |
Digital Inclusion
Evaluated by MEDICI consortium
| Dimension | Scale | Weight |
|---|---|---|
| Innovation level | 0–5 | 1 |
| Sustainability (ongoing) | 0–1 | 1 |
| Evaluation evidence | 0–1 | 1 |
| Demonstrated reach | 0–1 | 1 |
| Documented learning | 0–1 | 1 |
Gender Equality
Evaluated by ProPEGE consortium (Equal Leadership)
| Dimension | Scale | Weight |
|---|---|---|
| Transferability / replicability | 0–3 | 1 |
| Impact on gender equality | 0–1 | 1 |
| Effectiveness | 0–1 | 1 |
| Efficiency | 0–1 | 1 |
| Evaluated outcomes | 0–1 | 1 |
| Sustainability | 0–1 | 1 |
| Achievement / evidence | 0–1 | 1 |
| Gender-mainstreaming embedding | 0–1 | 1 |
| Curator validation | 0–1 | 1 |
Gender Equality
Evaluated by C-NAPSE
| Dimension | Scale | Weight |
|---|---|---|
| Transferability / replicability | 0–3 | 1 |
| Impact on gender equality | 0–1 | 1 |
| Effectiveness | 0–1 | 1 |
| Efficiency | 0–1 | 1 |
| Evaluated outcomes | 0–1 | 1 |
| Sustainability | 0–1 | 1 |
| Achievement / evidence | 0–1 | 1 |
| Gender-mainstreaming embedding | 0–1 | 1 |
| Curator validation | 0–1 | 1 |
PES & Nature-Based Solutions
Evaluated by C-NAPSE
| Dimension | Scale | Weight |
|---|---|---|
| Carbon — sequestration / storage, evidenced | 0–3 | 1 |
| Biodiversity — conservation / improvement, evidenced | 0–3 | 1 |
| Water & soil — retention, infiltration, erosion control | 0–3 | 1 |
| Fire resilience — risk reduction, evidenced | 0–3 | 1 |
| Governance, certification & transferability | 0–3 | 1 |
One canonical spine, many sources
The rubric method above is the canonical spine of every score on this platform. Several catalogues began life by importing practice collections built by earlier projects and consortia (for example ProPEGE for gender equality, MEDICI for digital inclusion). Their original evaluations are respected as attributed priors: where a source consortium assessed a practice, that assessment seeds the corresponding rubric dimensions and the source is credited on the practice page — but the dimensions, the arithmetic and the published score are always Evidoria's. Where something came from never decides how it is organised or ranked: provenance is metadata, recorded per record — never structure.
Validation levels
Every practice page shows how the record was produced and checked. The ladder, from most to least validated:
- Expert-evaluated — a named evaluator completed the catalogue's rubric for this practice.
- Curator-reviewed — a curator has reviewed the record's content.
- AI-enriched, pending review — structured content was added by the keyless AI routine and awaits curator review.
- Imported — brought in from a source catalogue; content as published by the source.
- Editor-added — created directly by an editor and not yet separately reviewed.
How our evidence ladder maps to other standards
Evidoria classifies each practice's evidence by study design. The table below maps that ladder onto three frameworks widely used by what-works units — the Maryland Scientific Methods Scale, Nesta's Standards of Evidence and EMMIE. Correspondences are indicative (a specific study can sit higher or lower), and EMMIE is multi-dimensional: only its Effect dimension maps onto a design ladder — its Mechanism, Moderators, Implementation and Economics readings correspond to the implementation dossier on each practice page.
| Evidoria evidence strength | Maryland SMS | Nesta | EMMIE |
|---|---|---|---|
| Randomised controlled trial | Level 5 — randomised assignment to treatment and control | Level 3–4 — causal impact demonstrated; 4 where independently replicated | Strong Effect evidence — direct estimate of impact with a credible counterfactual |
| Quasi-experimental | Level 3–4 — comparison group without full randomisation | Level 3 — causal comparison against a control or matched group | Moderate-to-strong Effect evidence — counterfactual present, selection risks remain |
| Observational / pre–post | Level 2 — before/after measures without a comparison group | Level 2 — data showing change among recipients | Weak Effect evidence — change observed, attribution not established |
| Descriptive / self-reported | Level 1 — correlation or descriptive account only | Level 1–2 — a logical account of impact, possibly with descriptive data | No Effect estimate — EMMIE's Mechanism/Implementation reading may still apply |
Where the data comes from
Every practice lists its data sources — the catalogue or website each piece of information was retrieved from — together with the date it was accessed. You'll find these on each practice's page under “Data sources”.