capture · data quality · humanitarian data · indicators · M&E

Completeness is the only dimension most systems measure

Why the data quality dimension you can measure is the one that tells you least, and what the other two actually cost.

Somebody opens a dashboard and sees "data quality: 94%". It is green. Nobody asks what the 94 is a percentage of, and in most cases the answer is dull: 94 per cent of the cells that should have something in them do. That is a number about how full a spreadsheet is. It gets read as a number about whether the figures are right.

How that happens is partly a story about which quality checks are cheap. The research literature has never agreed on a standard list of data quality dimensions, and a 2019 survey of data quality tools says so without hedging. What it does record is which ones come up most: accuracy, completeness, consistency and timeliness.

Now take the frameworks that actually bind a programme in this sector. USAID's, set out in ADS 201, are validity, integrity, precision, reliability and timeliness. No completeness. No consistency. The Centre for Humanitarian Data, when HDX launched, chose accuracy, timeliness, accessibility, interpretability and comparability, under an overall test of relevance. Again, neither.

So two of the four dimensions researchers reach for first are missing from both. Neither organisation forgot them. The reason they left them out says a good deal about the green 94.

Some checks need more than the file

The survey borrows a distinction from Piro between hard dimensions, which a check routine can measure, and soft ones, which need someone's judgement. Useful, but it does not quite sort the four. A blunter question does: what do you need in your hands before you can compute this at all?

For completeness, the file. Count the cells, count the empty ones, divide. A script can do it in under a second on a workbook it has never seen.

Timeliness needs a second date. The Centre for Humanitarian Data's definition is the delay between when data is collected and when it becomes available, and most operational files carry one date at best.

Consistency needs something to compare against: another file and a key to join them, or at the cheap end, the same file checked against itself. Accuracy needs a person to go and look.

The survey then tested this against real software, which is its most useful part. Of the thirteen tools it evaluated in depth, none had a metric for timeliness and none for consistency. Completeness and uniqueness were the only dimensions with wide implementation. Thirteen tools in 2019 is a small sample and vendors have shipped features since, so this is a snapshot rather than a verdict. But the snapshot matches what anyone who has sat through a data quality review has seen: the dashboard leads with fill rate because fill rate is what the software can do.

A full column can still be missing most of the data

UNICEF's Data Quality Framework arranges the same ideas in a way the flat lists do not. Under "how accurate are the data?" it groups accuracy with completeness, precision, reliability and consistency. Timeliness sits under a separate question with punctuality, time since reference and frequency.

Completeness, in other words, is not a sibling of accuracy there. It is one of the things accuracy depends on.

The framework defines it carefully, too: how many of the events or individuals in the target population were actually documented. For birth registration, the share of births that should have been recorded and were. For a survey, the response rate.

Put that next to a fill-rate check and the two come apart. Take a monthly 5W for a nutrition cluster where every row has a figure in "people reached". By the cell count it is complete. Whether it covers every partner working in the district is a different question, and the file cannot answer it, because a partner that reported nothing leaves no row behind to be counted as blank.

That is the uncomfortable bit. The worst kind of incompleteness, the absent partner, the district nobody reported for, does not show up in a completeness score at all.

Timely by whose clock?

A typical 3W or 5W has a reporting period, which is the window the work fell in, and a file date, which is whenever somebody last pressed save. Neither tells you when the figures were collected. A row typed in March for work done in January, in a workbook last saved in June, has three candidate dates and nothing to say which one matters.

HDX handles this pragmatically. The State of Open Humanitarian Data 2026 first asks whether a dataset is available, then whether it is up to date, judged against the update frequency the contributing organisation set for itself. On that basis it estimates that 68 per cent of crisis data across 22 operations is available and up to date, down from 74 per cent a year earlier.

For a platform holding thousands of datasets from hundreds of sources, with no way to audit anyone's collection process, it is the call most people in that seat would make. It does mean the measure is about whether organisations keep their own promises. That is worth knowing. It is not the same as knowing whether the data describes the present.

Programmes do a smaller version of this all the time. "Submitted on time" means the report arrived before the deadline. Whether the numbers inside were collected last week or carried over from last quarter goes unrecorded, because the form has one date field and the submitter fills it in.

USAID's own definition sets a harder test than either. Data should be available at a useful frequency, be current, and be timely enough to influence management decision making. The last part ties timeliness to a decision rather than a calendar. It is the right test. No script can run it.

Most consistency checking is internal

Internal consistency is the affordable version: a field compared across rows, an entity across periods, totals against their parts. It is how you catch one district spelled four ways, or an age-band scheme that changes halfway down a sheet. Most automated checking lives here, and it earns its keep.

External consistency is where the expensive mistakes are. The figure sent to the cluster against the figure sent to the donor. A partner's delivery count against the warehouse dispatch log. A reached total that is somehow larger than the population it came from. A file can be internally spotless and still be wrong about the world, and only an outside comparison shows it.

It is rare because it needs two datasets at once, in a comparable shape, with something to join them on. The join is where it usually dies. The partner's name is spelled differently in each file, one uses admin 2 and the other admin 3, or the reporting months are offset by a fortnight and nobody wants to be the person who reconciles them.

Why the funder lists leave them out

Once you look at what each dimension costs, the gaps in the USAID and HDX lists stop looking like gaps.

Both lists describe what somebody has to answer for, not what a script can compute. The US Federal Committee on Statistical Methodology's framework is explicit about this. It says it differs from frameworks in computer and information science because it judges data for statistical use and decision making, not for operational purposes. Its eleven dimensions sit under utility, objectivity and integrity, and completeness is not among them.

Completeness and consistency, in their cheap forms, are exactly the checks you can run without anyone being accountable for the answer. Leaving them off a standard makes sense if the standard exists to tell a mission what it must verify before it reports externally. It makes less sense if readers take the list as a full account of what can go wrong, and most readers probably do.

Which brings us back to the 94.

Three small changes

Name the dimension. "94 per cent of required fields populated" tells a reader what was checked. "94 per cent data quality" tells them nothing while sounding like everything. It costs four words.

Add a collection date. With it, timeliness becomes something you can compute rather than assert, and it is one column. Expect some pushback from field teams, because close to a deadline the collection date is often the awkward one, and that is exactly why it is worth having.

Run one outside comparison. Not a consistency regime, nobody has budget for that, but one pairing that matters to you: what you reported upward against what you hold locally, or delivery counts against dispatch. Once a quarter is enough to catch a kind of error no internal check can see, and to give you a rough feel for the comparisons you are not running.

What this argument does not settle

Completeness has taken a beating here, and it deserves some defence. A column at 40 per cent fill is telling you something urgent about how data is being collected, and no amount of cross-checking replaces that signal. The problem is that the cheap measure gets over-read, not that it exists.

None of this touches accuracy either, which is the only dimension most readers really care about. A dataset can be complete, current and consistent with every other dataset while every figure in it is wrong, because they were all copied from the same bad source. Finding that out means going and looking, and it is why USAID's standards come with a Data Quality Assessment rather than a dashboard.

One relationship would settle a lot of this, and it does not seem to have been measured: whether a dataset that scores well on the cheap dimensions is any more likely to be accurate. It feels as though it should be. A team careless with empty cells is probably careless elsewhere. But that is precisely the kind of reasonable-sounding proxy that fails without anyone noticing, and as far as the published literature shows, nobody has checked.


Sources

Funder and platform frameworks

  • Activity Monitoring, Evaluation and Learning Plan Guidance Document, USAID (November 2017), hosted on Grants.gov as an attachment to a funding opportunity. Source for the five data quality standards of ADS 201, validity, integrity, precision, reliability and timeliness, with their definitions, including timeliness as available at a useful frequency, current, and timely enough to influence management decision making, and for the Data Quality Assessment as the examination of indicator data against those five standards. Hosted on Grants.gov because USAID's own pages are no longer reliably available
  • Measuring the quality of humanitarian data: An emerging framework, Javier Teran, OCHA Centre for Humanitarian Data. Source for the HDX dimensions of accuracy, timeliness, accessibility, interpretability and comparability under an overall test of relevance, and for timeliness defined as the delay between collection and availability
  • The State of Open Humanitarian Data 2026, OCHA Centre for Humanitarian Data (March 2026). Source for the availability-then-timeliness assessment method, for timeliness judged against the update frequency set by the contributing organization, and for the estimate that 68 per cent of crisis data across 22 operations is available and up to date, against 74 per cent the previous year
  • Data Quality Framework, UNICEF. Source for accuracy, completeness, precision, reliability and consistency grouped under the question of how accurate the data are, for timeliness grouped with punctuality, time since reference and frequency, and for completeness defined as the proportion of events or individuals in the target population actually documented

The academic position

  • A Survey of Data Quality Measurement and Monitoring Tools, Lisa Ehrlinger, Elisa Rusz and Wolfram Wöß, Johannes Kepler University Linz (2019). Source for the absence of consensus on a standard list of dimensions, for accuracy, completeness, consistency and timeliness as the four most frequently used, for Piro's distinction between hard and soft dimensions, and for the finding that none of thirteen evaluated tools implemented a timeliness or consistency metric while completeness and uniqueness were the only widely implemented ones
  • A Framework for Data Quality, Federal Committee on Statistical Methodology (September 2020). Source for eleven dimensions organised under utility, objectivity and integrity, for the framework's stated difference from computer and information science frameworks, and for timeliness defined as the time between the event described and the data's availability