How the Procurement Integrity Risk Score works.
PIRS is an explainable 0–100 score that flags tenders and awards for human review. It is a prompt for scrutiny, never a finding of misconduct. Every flag is an inference with graded severity and human-readable evidence, and every flag can be contested — contesting adjusts the score. We publish the full method, including the parts that are standard and off-the-shelf, because a scoring model you can’t inspect isn’t one you should trust.
The core discipline: no fabricated zeros
A signal that cannot be computed for a record returns null and is excludedfrom the score — it is never silently scored as “clean.” A tender with no disclosed bidder count does not get a reassuring single-bidder pass; the indicator simply does not participate, and the record’s data-quality (the share of indicators that could actually run) is shown alongside its score. This is why an honest integrity panel can look sparse on thin data — and why a flag that does fire is backed by real, present data.
The ten indicators
Weights sum to 100. Each indicator’s severity (0–1) is multiplied by its weight and summed into the score. Weights below are read directly from the scoring engine.
| Indicator | Weight | Type | Needs | Rule |
|---|---|---|---|---|
| Single / no competitive bid OCP Cardinal | 16 | Rule-based | Disclosed bidder count | 1 or 0 bidders → maximum severity; exactly 2 → half; 3+ → clean. Fires only where the source discloses a real bidder count. |
| Short bidding window OCP Cardinal | 12 | Rule-based | Publication + submission deadline | Severity scales as the window falls below the 30-day norm; a same-day or negative window → maximum. |
| Non-competitive procedure OCP Cardinal | 12 | Rule-based | Procurement method | Direct / single-source / limited / restricted / emergency → maximum. Open / ICB / NCB / RFP → clean. Unknown method is excluded, not scored. |
| Threshold avoidance (contract splitting) OCP Cardinal | 8 | Rule-based | Contract value | Value within 10% below a common competitive-bidding threshold; the closer it sits under the line, the higher the severity. |
| Repeat-winner concentration OCP Cardinal | 12 | Corpus-calibrated | Historical buyer→supplier awards | One supplier repeatedly winning from one buyer; severity climbs with the count of prior awards (5+ → maximum). |
| Value outlier Statistical baseline | 10 | Corpus-calibrated | Sector×country 90th-pct baseline | Contract value above the sector/country 90th percentile; 2× the p90 → maximum severity. |
| Buyer burst Statistical baseline | 10 | Corpus-calibrated | Rolling buyer notice count | Abnormal clustering of notices from one buyer in a short window (12+ → maximum). |
| Award-vs-estimate gap OCP Cardinal | 8 | Rule-based | Award value + tender estimate | Award deviates from the pre-tender estimate; a 50%+ deviation → maximum severity. |
| Short decision period OCP Cardinal | 6 | Rule-based | Deadline + award date | Award made suspiciously soon after the submission deadline (little time to evaluate); same-day → maximum. |
| Disclosure opacity Open Contracting | 6 | Rule-based | Notice completeness | Key fields (value, sector, buyer) missing from a notice; severity scales with how many are absent. |
Bands
Score below 30: low. 30–60: elevated. 60 and above: high. A band is a triage cue for a reviewer, not a verdict.
What is off-the-shelf, and what isn’t
Off-the-shelf (and we won’t pretend otherwise): the individual indicator primitives are grounded in the open Open Contracting Partnership / Cardinal red-flag canon. Threshold avoidance, single-bidder, short window, and price gap are established red flags — not our invention.
Genuinely hard to replicate: (1) the cross-jurisdiction historical award corpus that makes the corpus-calibrated indicators meaningful; (2) the capital ↔ procurement join — tying a credit/bond view to procurement-integrity signals for the same entity; (3) the explainable, contestable scoring frame, with a tamper-evident audit trail over every score and contest. The defensibility is the corpus, the join, and methodology credibility — not a secret algorithm.
Does it work? How we back-test
A defensible prior is not the same as a measured one. To turn “the weights are reasonable” into evidence, we back-test the award-stage indicators (single-bidder, award-vs-estimate gap, and repeat-winner concentration) against an independent proxy for a known integrity problem: suppliers formally debarred by the World Bank and other multilateral development banks. If the score carries real signal, awards to debarred suppliers should concentrate more, and higher-severity, red flags than the corpus base rate. The harness reports lift, top-decile precision and recall, and — critically — its own statistical power.
- Debarment is a supplier-level, lagging, and sparse proxy — never a per-contract verdict, and “not debarred” is not “clean.”
- Supplier↔debarment matching is deterministic name canonicalization, so both false matches and misses are possible.
- Only award-stage indicators are exercised; notice-stage indicators are out of scope for a supplier-linked label.
- Below 10 debarred-linked awards the result is reported as underpowered — indicative, not validating.
Because the lift figure depends on the live award corpus and the current debarment label set, we do not print a fixed accuracy number here — a hard-coded metric on a static page would be stale the day it shipped and could not be trusted. Live back-test metrics — lift, precision, recall, coverage, and the underpowered flag — are computed against the current corpus and published in-platform, alongside the limits above.
Contesting a score is a first-class right
Every flag is an inference, and inferences can be wrong — a legitimate sole-source, a data-entry artefact, a canonicalization mismatch. Any scored entity (a firm, a notice, or an award) can dispute its score. A contest is reviewed by our team, and a well-founded contest adjusts the score and is recorded in the tamper-evident audit trail. Contesting is not a favour we grant; it is part of the method.
Contest a score