annual review · donor reporting · logframe · portfolio management · scoring
What impact weights hide in a project score
How the UK's output weights turned five results into one grade, what that grade could conceal, and what to record so it cannot.
From 2012, the UK's Department for International Development scored every project it funded once a year, and it published a worked example of how the score was built. A project has five outputs. Output 1 carries 30 per cent of the weight and is scored A+. Outputs 2 and 3 carry 10 per cent each and score A and B. Output 4 carries 20 per cent and scores A. Output 5 carries the other 30 per cent and scores C, meaning outputs substantially did not meet expectation.
The project score comes out at 90. Under the department's published grading bands, anything from 87.5 to 112.5 is an A. The project grades A, outputs met expectation, with one of its two most heavily weighted outputs having substantially failed.
The arithmetic is correct. What it assumes is the interesting part. A weighted average lets a strong result on one output make up for a weak result on another, and the example shows it doing exactly that: the A+ on output 1 absorbs the C on output 5. Whether that is acceptable depends on something the score does not say, which is whether output 1 can actually stand in for output 5 in the programme's logic.
A weighted grade is useful only when its readers can see what it allows to compensate for what. Most of the time they cannot.
A weighted average assumes outputs can substitute
Weighting answers one question: how much does each output count towards the total? It is silent on a second: can more of one replace less of another?
For some outputs the answer is yes, at least roughly. Two training streams aimed at the same district health workforce can trade off against each other; if one overdelivers and the other falls short, the workforce may still end up in about the same place. For others the answer is no. Take an illustrative maternal health project with two outputs, midwives on duty overnight and transport to reach them at night. Both might deserve 30 per cent. Neither can replace the other. A fully staffed facility that nobody can reach after dark delivers no more births than an empty one, and overdelivering on staffing does nothing to repair the missing transport.
Giving those two outputs equal weight says they matter equally. It does not say they are interchangeable, and the weighted sum treats them as if they were. A project could deliver the staffing in full, deliver almost none of the transport, and still score close to an A. The grade would be arithmetically sound and say almost nothing about whether women were giving birth in facilities.
This is the distinction the opening example hides. Thirty per cent on output 5 can mean "this output is worth thirty points" or "this output is essential and the rest cannot cover for it". The scoring system only knows the first meaning.
What a weight records
The weights themselves are set at design. The department's guidance on reviewing and scoring projects, issued in November 2011 for reviews from January 2012, sets out the rules. Each output carries an impact weight on the logframe, alongside a risk rating. No weight may be below 10 per cent. Changing one means changing the others so the total stays at 100. At every annual review the team is asked whether a weight has been revised since the last review and, if so, why.
It is tempting to read a weight as a causal claim, so that 30 per cent against 10 means the outcome depends three times as much on the first output. The system does not support that reading. What a 30 to 10 split establishes is that the team gave the first output three times the importance in the score. Whether that ranking reflects the programme's causal logic, or the team's confidence in delivery, or simply a convenient division of 100, is exactly what a reviewer would need to examine. The review template asks why a weight was revised, not why it was set.
Still, it is the only place in the framework where a team has to rank its outputs in writing. That makes the column worth reading, provided it is read as a choice about scoring rather than as evidence about causation.
The same number sets the grade
Under the 2011 guidance, each output is scored A++ to C at every annual review, on whether results achieved to date match those expected in the logframe. Those letters convert to numbers, 150 for A++ down to 50 for C, and the department's system multiplies each by its weight and sums them; the guidance states that this overall output score is calculated automatically. The outcome is assessed at annual review but not scored. It receives a score only at the project completion review.
Project scores then fed a portfolio quality index, weighted by budget, reported to the department's executive management committee, investment committee and board. At the baseline in the methodology note, 1,165 projects with lifetime budgets of £48.1bn scored 103.0, an A.
The department's results note for 2017 to 2018 is candid about the limits. It says the weights are assigned by project staff and the output scores are subjective and self-reported, and that a rise or fall in the index may reflect changes in performance or changes in how staff systematically rate projects.
Those are the department's acknowledgements. What follows is this article's inference, and it should be read as one. A number that both records a design choice and determines a grade reported to a board creates a possible incentive: a team uncertain about its hardest output has a reason to give it less weight, and a team ahead on an easy output has a reason to give it more. The pressure is the same one described in the piece on proxy indicators, where a measure that becomes a target starts to bend. The sources reviewed here do not show that any team reweighted strategically. They show that the system made it possible and that the scoring rested on judgement.
Leaving the weights out does not remove them
The obvious alternative is not to weight at all. The Global Fund's grant ratings before its 2013 funding reforms show why that does not solve the problem.
The Fund averaged each grant's percentage achievement against target, once across all indicators and once across a top-ten subset, capping any single indicator at 120 per cent, and converted the averages to letters: A1 above 100, A2 from 90 to 100, B1 from 60 to 89, B2 from 30 to 59. Within each average, every indicator counted equally. Equal weighting is still a weighting, and averaging makes every indicator compensable.
Victoria Fan and colleagues at the Center for Global Development reconstructed a sample of those scorecards in 2013. One was an HIV grant in Ethiopia. Its indicator for HIV-positive pregnant women receiving a complete course of antiretroviral prophylaxis reached 4,910 against a target of 25,000, under 20 per cent. Condoms distributed reached 125.6 per cent, capped at 120. Across all indicators the average was 96 per cent, which converts to A2. The paper reports the top-ten average as 84 per cent in its text and as 91 per cent in its appendix table; either figure sits well above the B2 band. The grant was rated B2 and received 17 per cent of its originally allocated phase 2 amount.
The authors read the scorecard as suggesting that the prophylaxis shortfall drove the outcome. If so, the average had allowed condoms to compensate for prevention of mother-to-child transmission, and someone overrode that compensation by hand, after the scores were in. Across their sample the authors found that at least a third of ratings could not be reproduced from the indicators using the published conversion table.
Set beside the UK system, the comparison is not between weighting and neutrality. The UK wrote its judgement down in advance and let the sum run. The Global Fund left the judgement unwritten and applied it afterwards through discretion. Neither told a reader, before the score arrived, which results the score was allowed to trade against each other.
What to record at the next review
None of this requires a new scoring system. It requires a few lines of record-keeping alongside the weights, and a team preparing a single annual review can do it.
- Record why each weight was chosen, in a sentence: the output's place in the theory of change, the team's confidence in delivery, or both.
- Keep the previous values. When a weight changes, the old figure should stay visible next to the new one, with the date.
- Explain every revision, and say whether it followed a weak score on that output.
- Name any output whose failure the overall grade must not conceal, and report its score beside the grade rather than only inside it.
The fourth is the one that answers the compensation problem directly. A project score accompanied by "output 5, essential, scored C" can no longer pass as an unqualified A.
At portfolio level the same records support a set of questions worth asking, though the answers are prompts for investigation rather than findings. Where does the heaviest weight sit relative to the theory of change, and if they differ, why? Do weights tend to fall after weak scores on the same output? Where weights are split evenly, was that a considered judgement that the outputs matter equally, or a default? Where an output sits at the 10 per cent floor review after review, is that because it was always minor, or because expectations about it changed? Each pattern has innocent explanations, and a single project tells you little.
Where this stops
The sources reviewed for this article do not establish how weights were actually distributed across the UK portfolio, how often they were revised, or whether revisions followed scores. The department published its rules, its index, its caveats, and committed to publishing every annual and completion review, which makes the analysis possible. Until someone does it, the incentive described above is a plausible mechanism, not a measured effect.
There is also a real cost to the obvious fix. If weights were stripped of any effect on the grade, they would record intentions nobody had much reason to keep current, and the department linked them to scoring precisely so that the column would be taken seriously. Flagging non-compensable outputs carries a similar risk: a team that knows an essential output will be reported on its own has a reason to be cautious about which outputs it calls essential.
The open question is whether a weight can carry real consequence without bending towards the grade it produces. What does seem clear is the minimum a reader needs. A weighted score that does not say what it allowed to compensate for what is a number whose meaning depends on information the reader was never given.
Sources
How the UK weighted and scored outputs
- Reviewing and Scoring Projects, How to Note, Department for International Development (November 2011). Source for the weight rules (on the logframe, 10 per cent minimum, totalling 100, revisions explained), the A++ to C scale, the automatic output score, outputs scored and the outcome assessed but not scored at annual review, and the commitment to publish reviews
- Methodology note: Portfolio Quality Index, Department for International Development. Source for the worked example scoring 90 and grading A, the 150 to 50 score values, the grading bands, the budget-weighted index, its reporting lines, and the baseline of 103.0 across 1,165 projects and £48.1bn
- Portfolio Quality Index, Single Departmental Plan results 2017 to 2018, Department for International Development. Source for weights assigned by project staff, scores described as subjective and self-reported, and the caution that movements may reflect how staff rate projects
The unweighted alternative
- Grant Performance and Payments at the Global Fund, Victoria Fan, Denizhan Duran, Rachel Silverman and Amanda Glassman, Center for Global Development Policy Paper 031 (August 2013). Source for the pre-2013 rating method and conversion bands, and the Ethiopia HIV scorecard (prophylaxis 4,910 of 25,000, condoms 125.6 per cent, all-indicator average 96 per cent, rating B2, 17 per cent of phase 2). The top-ten average is given as 84 per cent on printed page 14 and 0.910 in Appendix 10. Also the source for the authors' reading of the downgrade and the finding that at least a third of ratings could not be reproduced