Design

New versus repeat: the hardest number in the sector

Summing reach across periods without knowing whether someone is new or returning produces a number that is wrong in a direction nobody can estimate.

The most reported number in humanitarian and development work is the number of people reached. Every donor asks for it. Every cluster reports it. Every annual review contains it. The number appears on dashboards, in parliamentary questions, in press releases announcing the scale of a response. It is the closest thing the sector has to a universal output metric, and it is almost always wrong.

It is wrong in a specific way. When a programme distributes food to a household in January and to the same household again in March, the programme has performed two deliveries. Whether it has reached one household or two depends on a single piece of information that most reporting systems do not record: was this household already counted? The answer determines whether the figure at the end of the year is a count of unique households served or a sum of service contacts, and the difference between those two numbers can be a factor of two, four, or twelve, depending on how often the programme delivers and how stable the recipient population is.

Nobody disputes that the distinction matters. What makes it the hardest number in the sector is that the distinction is easy to define and extraordinarily difficult to collect, and most guidance documents handle the difficulty by stating the requirement and leaving the method to the implementing partner.

What donors ask for, and what they get

The guidance is clear. The practice is not.

DG ECHO's Single Form instructions state that partners should provide the number of unique beneficiaries and avoid double-counting: if the same beneficiaries benefit from several interventions within the same sector, they should be counted only once. The 2021 guidelines add that the overall number per sector-specific category cannot exceed the total number of beneficiaries in a given sector. That is a ceiling rule, and it is enforceable only if the partner can identify which individuals appeared in more than one intervention.

FCDO's Afghanistan results factsheet (2024 to 2025) says results are collected on a unique beneficiary basis, with each beneficiary counted only once per result regardless of how many times they received assistance. To avoid double-counting across delivery partners, only the partner with the largest reach in each province is included in the aggregated total. The factsheet describes the result as a conservative estimate and says it should be interpreted as an "at least" figure.

An older DFID results framework methodology describes the same indicator type as "Peak Year (or cumulative if double-counting can be avoided)." The parenthetical is doing all the work. It says: we want a cumulative number, but only if you can deduplicate. If you cannot, give us the single highest year instead, because a peak-year figure is at least bounded. The methodology goes on to explain why it restricts the indicator to food, cash and vouchers: obtaining the total number of beneficiaries across different assistance types would result in a high level of double- and triple-counting, because the same person might receive food, shelter, and WASH services.

OCHA's guidance for the 2020 Global Humanitarian Overview draws a further distinction between "people reached" and "people covered." People reached is the number of people targeted who have benefited from one or several humanitarian activities at least once during the reporting period. People covered considers the type, quantity, quality and periodicity of the assistance received, and represents the number whose targeted needs have been met. "Reached" is the minimum: one contact counts. "Covered" is the standard: the planned assistance was delivered in full. Neither definition addresses whether a person reached in January and reached again in June is one person or two in the annual total.

Each of these documents states the requirement: count unique individuals. None of them resolves the operational problem that makes unique counting difficult, which is that most humanitarian data systems do not track individuals across deliveries.

Why the field cannot answer

The reason the new-versus-repeat distinction is so hard to collect is structural, not technical.

In a cash transfer programme using biometric registration, the system knows exactly which individuals have received assistance and when. UNHCR's PRIMES ecosystem, built around the proGres case management database, registers refugees and asylum seekers individually with photographs and biometrics, and its Global Distribution Tool verifies identities at the point of assistance. As of 2024, proGres held records on approximately 18 million registered refugees across more than 130 countries. WFP's SCOPE platform performs a similar function for food assistance. But even with dedicated registration systems, not every recipient is individually identified. The WFP External Auditor's report to the Executive Board (2021) noted that in 2019, 59 per cent of WFP's direct beneficiaries were participants, meaning individuals who physically attended an activity and received aid directly, and were potentially identified by name. The remaining 41 per cent were estimated, and their names were not known. Those are the figures for the two largest humanitarian organisations, operating with purpose-built registration infrastructure. For most implementing partners, the share of estimated rather than individually identified beneficiaries is higher.

In a blanket distribution, a health outreach, a community training, or a WASH intervention, no individual registration occurs. The field team counts the number of people who attended or the number of households that received supplies. If the same team returns to the same village the following month, they count again. Whether anyone in the second count was already in the first is unknown, because there is no identifier to match against.

The International Mine Action Standards technical note (2023) on beneficiary measurement states the position with unusual candour: for a given type of activity, double-counting of beneficiaries should be avoided where possible, but it may be inevitable in some cases, and the effort to avoid it may not be reasonable. Any incidences of potential double-counting should be made clear in reporting. That last sentence is the honest standard. Not: eliminate double-counting. Rather: disclose it.

The Food Security Cluster's beneficiary counting methodology for the 5W tracker in humanitarian operations asks partners to flag which activity reached the highest number of households per location, as a proxy for unique households reached. The method is a ceiling estimate: take the largest single activity's reach as the unique count, on the assumption that smaller activities are likely reaching a subset of the same people. It is an intelligent approximation, and it is an approximation, not a count. The underlying data does not say which households overlap.

Three states, not two

The question "is this person new or a repeat?" has three possible answers, and the third is the one that matters most for data integrity.

A person is new if they have not previously received assistance from this programme. A person is a repeat if they have. A person is unknown if the data system cannot tell. In most humanitarian datasets, the honest answer for every record is the third one.

The distinction matters because the three states aggregate differently. A cumulative flow indicator, one that answers "how many people have we ever reached?", should count only new individuals. If it sums all records, the total grows with every delivery cycle regardless of whether anyone new was served. A per-period flow indicator, one that answers "how many service contacts did we make this quarter?", should count all records, because every delivery is a real event even if the recipient was already counted. The first number tells the donor how broad the programme's coverage is. The second tells the programme how much work it did. They are both legitimate, they answer different questions, and they are only distinguishable if the system records which state each record is in.

When the state is unknown, neither number can be produced reliably. A sum of unknown records is not a count of unique individuals and not a count of total contacts. It is a number whose relationship to reality depends on the repeat rate, which is the piece of information the system does not have.

FCDO's Afghanistan methodology handles this by taking the conservative path: count only the largest partner per province, round down, and call the result "at least." That produces a defensible floor. It does not produce a count of unique individuals, because two partners operating in the same province may be reaching different populations, and excluding the smaller one discards real coverage. The alternative, summing both, risks counting the same household twice. Both choices are wrong in a known direction, and FCDO chose the one that understates.

What the datasets show

The operational datasets collected in the humanitarian 3W, 4W and 5W reporting systems confirm that this is not a theoretical concern.

The Colombia RMRP reporting system for 2025 carries 33,934 rows of delivery data across 76 columns, including 28 disaggregation columns covering eight dimensions: sex, gender (with a non-binary option), age bands, disability, ethnicity, pregnancy and lactation status, LGBTI+, and population group. It has two total-reached columns, both at 100 per cent fill. What it does not have is a field recording whether each row's beneficiaries were previously counted.

The Pakistan Monsoon Floods 5W for 2025 carries 20,164 rows with activity and reached figures at high fill rates. It has two total columns (Total Reached Individuals and Individuals) and four sex-and-age disaggregation columns filled at 71 to 78 per cent. Six mutually incompatible age-band schemes coexist in the same table. Again, no new-versus-repeat field.

The Venezuela 5W for 2025 is the largest in the corpus at 124,126 rows, with activity and reached both at 100 per cent fill. The Mozambique 4W from 2019 carries 19,338 rows. The Philippines Typhoon Rai 3W from 2022 has 14,214 rows. None of them carries a column distinguishing new from repeat beneficiaries.

In every one of these files, summing the reached column produces a number. That number is reported. It is almost certainly larger than the number of unique individuals served, because any programme that delivered more than once in the reporting period will have counted some recipients more than once. By how much the number is inflated, nobody knows, because the data to calculate the inflation factor was never recorded.

The direction nobody can estimate

The phrase "wrong in a direction nobody can estimate" is precise and worth unpacking.

If the repeat rate were known, the correction would be straightforward. A programme that distributes monthly to a stable population of 10,000 households has a repeat rate close to 100 per cent after the first month, and its annual sum of 120,000 household-contacts maps to roughly 10,000 unique households. A programme that distributes once to each of twelve different communities of 10,000 has a repeat rate close to 0 per cent, and its annual sum of 120,000 maps to roughly 120,000 unique households. Same number, same indicator, different meaning by a factor of twelve.

In practice, the repeat rate varies by sector, by modality, by context, by month, and by location within a single programme. A cash transfer programme targeting displacement-affected households in a camp has a high repeat rate, because the same households stay in the camp. A cash transfer programme targeting new arrivals at a transit point has a low one, because the population turns over. Both programmes report "number of individuals reached with cash assistance." The number means something different in each case, and the aggregation of the two produces a figure whose relationship to the number of unique individuals is determined by a weighted average of two unknown repeat rates.

This is why DFID's older methodology retreated to "peak year" as the fallback. A peak-year figure avoids cross-period aggregation entirely. It says: in our best single year, we reached this many people. It does not say how many of them were also reached the previous year. It sacrifices the cumulative story for a number that is at least defensible within a twelve-month window, though even within that window, monthly repetition is unresolved.

What would have to change

The gap is not in the guidance. The guidance is clear: count unique individuals. The gap is between what the guidance asks for and what the data systems record.

Closing that gap at the individual level, with biometric or digital registration of every recipient, is possible in some programme types and contexts, and unrealistic in others. A blanket supplementary feeding programme reaching 200,000 children cannot register each child before distributing a sachet of ready-to-use therapeutic food. The operational cost of universal registration exceeds the value of the precision it would provide, and in acute emergencies the delay it would impose is measured in lives.

But the gap could be narrowed without individual registration, by recording at the delivery level whether the recipients are believed to be new, repeat, or unknown. That field does not require identifying anyone by name. It requires the field team to answer one question at the point of data entry: are these the same people we served last time? The answer will often be "we think so" or "we do not know," and both of those are more informative than silence. A dataset where 80 per cent of rows are marked "unknown" and 20 per cent are marked "repeat" is more honest than a dataset where every row is silently treated as new, because it tells the reader that the cumulative total is an upper bound and that at least a fifth of the records are known to include people already counted.

Whether the sector will adopt that field is a governance question, not a technical one. Any spreadsheet can hold a column. The question is whether donors will require it, whether cluster coordinators will enforce it, and whether programme managers will fill it in honestly rather than defaulting to "new" because it makes the coverage number higher. The incentive structure points in the wrong direction: reporting more unique beneficiaries is better for the next funding proposal than reporting fewer, and the distinction between a genuinely large programme and a moderately large programme counting people twice is invisible in every dataset reviewed.

That incentive is the reason this is the hardest number in the sector, and it is the reason the problem has persisted for decades despite being well understood by everyone who works with the data.


Sources

Donor guidance and methodology

Operational methodology and audit

Practitioner analysis

Operational datasets referenced

  • Colombia RMRP/PRPC 2025 reporting (33,934 rows, 76 columns, 28 disaggregation columns across 8 dimensions). Dataset reviewed in evidence probe; not publicly linked here
  • Pakistan Monsoon Floods 5W 2025 (20,164 rows, 34 columns). Dataset reviewed in evidence probe
  • Venezuela 5W December 2025 (124,126 rows). Dataset reviewed in evidence probe
  • Mozambique 4W 2019 (19,338 rows). Dataset reviewed in evidence probe
  • Philippines Typhoon Rai 3W 2022 (14,214 rows). Dataset reviewed in evidence probe