What an independent program evaluator measures, layer by layer
Federal awards often require an independent evaluator. This is what that evaluator measures, layer by layer, from inputs and activities through outputs, outcomes and the harder question of what would have happened anyway.
Contents
A funded program had run for eighteen months, the team was proud of it, and the evaluation report said the program had delivered every activity it promised and could not yet show that anything had changed for the people it served. Both halves of that sentence were true. Understanding why requires knowing what an evaluator measures and in what order, which is the subject of this article. I write it from the standing of an executive at a firm whose operating divisions include a program evaluation practice, not as an evaluator myself, and I name no agency and no client.
Many federal awards require an evaluator who is independent of the team that delivers the program. The requirement is structural. The evaluator reports to the funder about the program rather than to the program about itself. What the evaluator reports on is a stack of measurements, and each layer answers a different question.
Inputs and fidelity
The first layer is whether the program that ran is the program that was funded. Inputs are the resources: staff hired against the positions proposed, dollars spent against the budget categories, equipment purchased, partners engaged. Fidelity is whether the activities were delivered as designed, at the dose designed, to the population designed. A curriculum written for twelve sessions that delivered seven, or a mentoring program that reached a different age group than proposed, has a fidelity finding before anyone asks whether it worked.
This layer is measured from the program’s own records, and the quality of those records is the first thing an evaluator learns about a program. Attendance logs, expenditure reports, staffing rosters and activity calendars are the raw material. A program that cannot produce them cannot be evaluated at any layer above this one.
Outputs
Outputs are the countable products of activities: people served, sessions delivered, materials produced, trainings completed, referrals made. They are the numbers most program reports lead with, and they are the layer most easily confused with success. An output tells the funder that the program did what it said. It does not tell the funder that anyone is better off.
The evaluator’s contribution at this layer is definitional. What counts as a person served, a single contact or a completed sequence. What counts as a training completed. The definitions are fixed in the evaluation plan before the program starts collecting, because a definition changed midway makes the two halves of the data incomparable.
Outcomes, short and long
Outcomes are changes in the people or systems the program set out to change: knowledge gained, behavior changed, a credential earned, a job obtained, a health indicator moved. Short term outcomes are measured during or immediately after participation. Longer term outcomes are measured months or years later, which is why so many evaluations of short programs end with the honest sentence that longer term outcomes could not yet be observed.
Measuring outcomes requires instruments: surveys, assessments, administrative data matched to participants, interviews. The instruments are chosen and written before the program starts, and they are administered at baseline and again later, because a change cannot be measured without a starting point. The most common weakness I see is a program that measured only at the end and then had nothing to compare to.
The counterfactual
The hardest question is whether the outcomes would have happened anyway. Participants in a job program may have found jobs without it. Students in a tutoring program may have improved because they were also in school. The evaluator’s job at this layer is to design a comparison that isolates the program’s contribution: a comparison group that did not receive the program, a randomized assignment where that is feasible and ethical, a matched group built from administrative data, or an interrupted time series where none of those are possible.
Each design has a strength of evidence attached to it, and funders increasingly specify the strength they expect. A descriptive study says what happened. A quasi experimental study says what happened relative to a plausible comparison. An experimental study says what the program caused. The evaluator states which of these the design supports and does not claim more.
Cost and the question of scale
A growing number of evaluations include a cost layer: what each output and each outcome cost to produce, and how that compares to alternatives. This is where evaluation and finance meet. The cost per participant served is an output cost. The cost per participant who achieved the outcome is an outcome cost, and the gap between the two numbers is often the most useful thing in the report. It tells a funder what scaling the program would buy.
The cost layer depends on the input records being clean, which brings the stack back to its foundation. An evaluator who cannot trust the expenditure data cannot produce a cost per outcome that anyone should act on.
Implementation and context
Alongside the numbers, an evaluator documents how the program was actually run and what surrounded it. Staff turnover, partner changes, a policy shift in the region, a pandemic. These are not excuses in the report. They are the explanation a reader will need for why a program that ran well in one site ran poorly in another, and they are the material a funder uses to decide whether the model travels.
This layer is qualitative, and it is written for a reader who was not there. Everything in it is built to survive that reader’s absence.
What the evaluator needs from the program
Each layer of the stack depends on records the program keeps, and the evaluator can only measure what was recorded. So the first meeting between an evaluator and a program team is usually about data: what the program will collect, in what form, at what points, and who owns it. A participant identifier that persists across sessions so that attendance can be linked to outcomes. A baseline instrument administered before the first activity. Expenditure coded against the budget categories in the award rather than against the organization’s general ledger. Consent for participants to be followed up later. None of this is exotic, and all of it has to exist from the first day, because the evaluator cannot go back and create a baseline in month nine. The programs that get the most from evaluation are the ones that treat the data plan as part of the program design rather than as a requirement imposed from outside. The ones that get the least are the ones that hand the evaluator a folder of sign in sheets at the end and ask what it shows.
What the layers mean for a program team
A team reading its first evaluation report often expects a verdict. What it receives is a stack: the program ran with this fidelity, produced these outputs, showed these short term outcomes, could or could not be distinguished from a comparison, and cost this much per result, in this context. Each layer can be strong while the next is weak. A program that delivered every activity and cannot yet show outcomes is not failing. It is at the layer its data allows.
The honest limit of the whole discipline is time. The outcomes that matter most take longest to appear, and most awards end before they do. The best evaluations plan for that from the start, by building the baseline, fixing the definitions and setting up the follow up while the program is still running and the participants can still be found.