data quality · education · indicators · measurement · proxy indicators
Proxy indicators and the cost of measuring the wrong thing
What a proxy indicator actually asserts, why the convenient one wins, and the two separate ways the link to the real thing breaks.
Almost every indicator is a proxy for something. Nobody measures dignity, resilience, empowerment or institutional capacity directly, because those are constructs rather than quantities, and what gets counted instead is something believed to move with them. The question is never whether to use a proxy. It is whether anybody checked that the thing being counted is still connected to the thing being claimed.
That check is rarely made, and when it fails the failure is quiet. A proxy does not announce that it has come loose. It keeps producing numbers, the numbers keep improving, the reports keep validating, and the programme keeps reporting progress against something it may have stopped achieving several years ago.
The largest worked example the sector has is the global push on school enrolment, and it is worth walking through properly, because it shows both ways a proxy fails and the second one is much harder to see than the first.
What a proxy actually claims
The clearest statement of the conditions comes from Lant Pritchett and Michelle Kaffenberger's work on learning profiles for the RISE Programme, and it is worth quoting the structure of the argument rather than the conclusion.
If the relationship between schooling completed and assessed skills is strong, meaning learning increases substantially with more schooling, and tight, meaning learning increases for nearly all students, then schooling is an adequate proxy for learning. Two conditions, and they are separable.
Strong is about the size of the effect. Tight is about its variance. A relationship can be strong on average and loose in distribution, which means the proxy works for the population and fails for any particular child, and a programme reporting at district level will never see the difference. It can also be tight and weak, where everyone moves together but hardly at all, which produces a proxy that looks reliable precisely because it is uninformative.
Most proxy justifications in programme documents assert neither condition. They assert plausibility: enrolment is related to learning, so counting enrolment tells you something about learning. That is true and it is not the claim the indicator is making. The indicator claims that a rise in the count corresponds to a rise in the construct, at a magnitude worth reporting, for the population being served.
What the enrolment case actually cost
The global push worked, on its own terms. Most countries met or nearly met the Millennium Development Goal that every child complete a full course of primary schooling.
Then the learning data arrived. Pritchett's Rebirth of Education recorded that a child entering fifth grade in Pakistan without knowing simple division had a one in six chance of learning it over an entire year of schooling. The learning profiles work found that in six of ten countries studied, half or fewer of the cohort of 18 to 37 year olds who had completed primary school could read.
The proxy had not been wrong at the outset. It had been weak and loose all along, and nobody had tested which, because the number was rising and a rising number does not prompt an audit.
Pritchett is direct about why enrolment was chosen. Speaking about the Millennium Development Goals, he describes agendas that prioritised measurable, easily communicable targets such as enrolment rates and years of schooling, which were politically and administratively convenient. The convenience is the mechanism. Enrolment is countable from an administrative register that already exists. Learning requires an assessment instrument, trained enumerators, a sampling frame and a budget.
So the proxy was selected on availability and then reported as though it had been selected on validity. That substitution is the thing to watch for, and it is not a failure of rigour by any individual. It is what happens when the design question "what can we measure?" arrives before "what are we claiming?"
Availability is the stated criterion, not a secret one
It is worth being precise here, because the convenient proxy is not usually chosen dishonestly. Funders say openly that availability governs.
The European Commission's working document on proxy indicators for rural development defines a proxy as an approximation to a common context indicator providing sufficient information to allow assessment of a relevant contextual aspect. Then it states one of the selection principles plainly: data should be available from EU sources at least at national level.
That is a reasonable rule for a policy framework that has to compare twenty-seven member states, and it makes availability a first-order criterion rather than a compromise. The UK's SDG reporting does the same thing in the open, flagging indicators as proxy where the national data differs from the UN specification and describing the substitute as the most suitable match currently available.
Both are honest. Neither says anything about strength or tightness, because at that level of aggregation nobody is claiming a causal link to an individual outcome. The problem arrives when a proxy chosen under those conditions is inherited by a programme that is making exactly that claim, about a specific population, in a specific district, with a target attached.
The second failure, which is different
Everything above concerns a proxy that was loosely connected from the start. There is a separate failure that happens to proxies which began well, and it arrives with the target.
Charles Goodhart's 1975 observation about monetary policy was that any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes. Marilyn Strathern's later compression is the version everyone knows: when a measure becomes a target, it ceases to be a good measure.
Donald Campbell stated the social-programme version in 1979 and it is the one that belongs in a results framework discussion. The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor. His own example was educational testing: achievement tests may be valuable indicators of general school achievement under normal teaching aimed at general competence, but when test scores become the goal of the teaching process they both lose their value as indicators and distort the educational process.
Read that against a logframe. An indicator with a target attached is a quantitative social indicator being used for decision-making, by definition, because the next tranche depends on it. Campbell's law is not a warning about badly designed programmes. It is a description of what a target does to a measure, in general, and a logframe is a machine for attaching targets to measures.
The practical consequence is that a proxy validated at design time is not validated for the period after the target is set, because setting the target changed the system. Nobody revalidates.
The fetishism problem
Des Gasper, writing at ISS The Hague in 2000, named the failure mode that follows and the name is well chosen. Fetishism, in his account, means forgetting that an indicator is only that, and forgetting that its validity should be re-examined regularly. He then identifies exactly what prevents the re-examination: it will not happen if consistency of the data series becomes treated as more important than its relevance.
That is the trap, and it is structural rather than cultural. A programme that has reported the same indicator for four years has a time series. Changing the indicator breaks the series, which means the fifth year cannot be compared to the first four, which means the change has to be explained to the funder, which means it is easier to keep reporting a measure you have stopped believing in.
Gasper also makes a point that cuts against the instinct to solve this with better indicator libraries: in many areas of work, capacity building among them, there is and can be no adequate standard list of indicators. A library gives you comparability. It does not give you validity for your programme, and a canonical indicator adopted because it was in the list is a proxy chosen on availability by another route.
What would actually help
Three things, none of which require agreement across the sector, and the first is nearly free.
Record that it is a proxy, and for what. An indicator row that says "number of health workers trained" is a count. The same row marked as a proxy for service quality, with the construct named, is a claim somebody can evaluate. Most templates have nowhere to put this, which is why it lives in the proposal narrative and disappears the moment the logframe is extracted from it.
Record the justification, and what would falsify it. UK Aid Match and several other funders ask for evidence that a proxy is appropriate. The useful version of that evidence is not a citation showing the two things are related in the literature. It is a statement of what relationship is assumed, how strong, and what observation would show it had broken. "We assume trained workers deliver measurably better consultations, and would expect that to show in the supervision checklist scores" is checkable. "Training improves quality" is not.
Give it a review date. Gasper's argument implies that a proxy has a shelf life, because the system it measures changes and the target changes it faster. A proxy with a review date is one somebody will look at again; a proxy without one is a permanent fixture by default, and the data series will defend it.
Where this stops
Two limits, and the second is a genuine objection to the whole argument.
The first is that revalidating a proxy usually requires measuring the thing the proxy was standing in for, which is the expense the proxy existed to avoid. The circularity is real. It is not total: a one-off validation study on a sample is far cheaper than continuous direct measurement, and it answers the strength and tightness questions for that population. But a programme that cannot afford the sample study cannot check its proxy, and telling it to do so anyway is advice that costs the adviser nothing.
The second is harder. The enrolment story is told here as a cautionary tale, and it is also a story about a global effort that put hundreds of millions of children into school who would not otherwise have been there. That is not nothing, and a measurement regime that had insisted on validated learning outcomes from the start might have produced a slower, smaller, better-measured push. Whether that would have been a better outcome for those children is not obvious, and anyone confident about it is reasoning from a counterfactual nobody has.
What can be said is narrower. The enrolment target was not wrong to exist. It was wrong to be read, for fifteen years, as a statement about education rather than a statement about attendance, and the reading was available to anyone who asked what the indicator actually claimed. Nobody asked, because the number was going up.
Sources
The strength and tightness conditions
- Learning Profiles: Schooling Versus Learning, Lant Pritchett and Michelle Kaffenberger, RISE Programme Working Paper 12. Source for the two conditions under which schooling is an adequate proxy for learning, strong and tight, and for the finding that in six of ten countries studied half or fewer of those completing primary school could read
- Schooling Ain't Learning, Center for Global Development, on Lant Pritchett's The Rebirth of Education. Source for the Pakistan fifth-grade division figure and for the near-achievement of the primary schooling Millennium Development Goal
- Schools are failing to deliver learning, Lant Pritchett, VoxDev. Source for the account of the MDGs prioritising measurable, easily communicable targets that were politically and administratively convenient
Proxy selection in funder guidance
- Working document: defining proxy indicators for rural development programmes, European Commission (2016). Source for the definition of a proxy indicator as an approximation providing sufficient information for assessment, and for the selection principle that data should be available from EU sources at least at national level
- UK SDG data, indicator 8.a.1. Source for the practice of flagging an indicator as a proxy where national data differs from the UN specification, described as the most suitable match currently available
The target effect
- Goodhart's law, summarising Charles Goodhart's 1975 formulation that any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes, and Marilyn Strathern's 1997 compression
- Campbell's law, in context. Source for Donald Campbell's 1979 statement on corruption pressures on quantitative social indicators, and for his own worked example on achievement testing losing its value as an indicator once it becomes the goal of teaching
On indicator validity over time
- "Logical Frameworks": Problems and Potentials, Des Gasper, Institute of Social Studies, The Hague (2000). Source for fetishism as forgetting that an indicator is only an indicator, for the requirement that validity be re-examined regularly, for consistency of a data series crowding out relevance, and for the argument that some areas of work admit no adequate standard indicator list