Capability building guide
Evidence for capability building: match proof to the claim.
Participation, learning, workplace practice, reliable performance and system continuity are different results. Measure them separately.
Published and last updated: 4 August 2026 · By Consultancy Mantra
The short answer
Good evidence answers a defined decision.
Start by writing the claim you may need to make. If the claim is that people attended, an attendance record is relevant. If the claim is that they can perform a practice at work, attendance is not enough; the evidence must show the practice in an authentic or realistic context. If the claim is that an organisation can sustain the practice, evidence must also cover roles, routines, tools, governance and ownership.
An evidence plan is therefore not a list of every available metric. It is a small, traceable set of questions, indicators, sources, responsibilities and review points tied to an intended use.
Begin with a claim
Separate description, contribution and attribution.
Different claims require different designs. Clarity at this stage protects teams from over-interpreting weak evidence later.
What happened, for whom and in what context?
Descriptive evidence can document participation, delivery, assessment results, observed practice or system conditions. It can show a pattern without establishing why that pattern occurred.
Is the programme a credible part of the explanation?
A contribution claim tests the intended pathway, assumptions, implementation quality, contextual factors and plausible alternatives. Multiple sources can strengthen or weaken the explanation.
Did the intervention cause the observed difference?
A causal claim needs a design capable of estimating what would have happened without the intervention. The World Bank’s Impact Evaluation in Practice explains counterfactual methods, sampling, data collection and evaluation management.
An evidence ladder
Collect evidence at the level you intend to claim.
The levels below are a practical synthesis for capability work. They are not interchangeable, and a higher level does not make lower-level operational data unnecessary.
1. Reach and participation
Examples include invitations, enrolment, attendance, completion and participant characteristics. These records help teams understand who was reached and whether access differed across relevant groups. They do not show what participants learned or used.
2. Learning and demonstrated task performance
Use a knowledge check, simulation, demonstration, assessed assignment or structured rubric aligned with the stated learning outcome. A realistic task is usually more informative than recall alone when the intended result is applied practice.
3. Application in the work context
Review observations, work products, decisions, workflow records or other traces of the practice. State the quality standard and support available. Self-report can add context, but it should not be the only evidence for a claim about observed performance.
4. Reliability over time and situations
One successful task shows possibility, not yet dependable performance. Review repeated evidence across an appropriate period or set of cases. Record exceptions and variation rather than hiding them in a single average.
5. Organisational conditions and continuity
Look for active ownership, competent support roles, usable tools, review routines, decision rights, resources and governance. Evidence should show that these conditions operate, not merely that a document says they should.
6. Outcomes and wider effects
Specify the result for people, services or the organisation and the time needed for it to emerge. Consider intended and unintended effects. Be explicit about whether the evaluation describes change, supports a contribution story or estimates causal impact.
Design from a theory of change
Make the expected route and assumptions visible.
The 2026 UK government Magenta Book advises defining the intervention and developing a theory of change at the start of evaluation scoping. It describes a chain from inputs and activities through outputs and outcomes, alongside mechanisms, assumptions, context and alternative explanations.
For capability building, this means stating why the learning and support are expected to change practice, why that practice should affect an outcome, and what else must be true. The map should be revised when evidence challenges an assumption.
Inputs and activities
People, time, content, tools, facilitation, coaching and operational changes. Evidence at this stage supports delivery management, not an outcome claim.
Outputs
Completed sessions, resources, assessments, support contacts or revised workflows. Outputs show what the programme produced and for whom.
Intermediate outcomes
Learning, confidence when appropriately measured, application, quality or adoption. These can provide earlier feedback while longer-term results are still developing.
Final outcomes and assumptions
The intended change in performance, service or organisational function, together with the contextual conditions and causal assumptions connecting each stage.
Build the evidence plan
Six fields make every measure easier to interpret.
What decision will this evidence inform?
Name the user and timing. A measure without a decision owner often becomes reporting activity with no clear consequence.
What observable sign represents the result?
The updated CDC Program Evaluation Framework describes indicators as measurable statements that bridge broad constructs and specific measures.
Where will the evidence come from?
Identify the assessment, work product, observation, interview, system record or survey. Check whether the source can answer the question without creating unreasonable burden.
What counts as sufficient performance?
Define quality, completeness, timing, independence and acceptable support. Where judgement is involved, use a rubric and examples to improve consistency.
When can the result reasonably be observed?
Collecting too early can mistake recall for transfer. Waiting until the end can remove the opportunity to improve delivery. Use review points that match the expected pathway.
Who collects, interprets and acts?
Separate data ownership, assessment and decision roles where risk warrants it. Record access, consent, privacy, retention and quality requirements before collection.
Avoid common evidence errors
Do not ask one measure to prove a different result.
Attendance presented as capability
Attendance proves presence under the stated recording method. Pair it with task, application or system evidence when the claim goes further.
Confidence treated as competence
Self-reported confidence can reveal perception and help target support. It may not match performance. Use it alongside observed or assessed evidence and investigate meaningful differences.
A test that does not resemble the work
A recall quiz may be valid for factual knowledge but weak for judgement, tool use or collaboration. Match the assessment method to the capability definition and intended context.
Only averages are reported
Averages can hide variation by role, location, starting point or access. Use disaggregation where it is relevant, ethical and statistically responsible, while protecting privacy.
Outcome change is presented as programme impact
The CDC framework distinguishes outcome evaluation, which measures achieved outcomes but cannot attribute causality, from impact evaluation, which compares outcomes with and without the intervention. Use causal language only when the design supports it.
Make evaluation proportionate
Evidence quality depends on use, context and risk.
The OECD guidance on evaluation criteria presents relevance, coherence, effectiveness, efficiency, impact and sustainability as complementary lenses and stresses that they should be adapted to purpose and context. A small internal pilot and a high-stakes public programme should not automatically use the same design.
Proportionality does not mean accepting evidence that cannot answer the question. It means narrowing the question, selecting a defensible method, documenting limitations and spending effort where uncertainty matters most.
For improvement
Use rapid, timely evidence on participation, task performance, barriers and implementation. The priority is to learn early enough to change the programme or workplace support.
For assurance
Use agreed standards, traceable records, clear assessment roles and an appropriate degree of independence. Retain enough context for a reviewer to understand what the evidence does and does not establish.
For impact decisions
Seek evaluation expertise early. Define the counterfactual question, ethical constraints, sample, data access and decision timetable before implementation removes feasible design options.
Apply the evidence model
Make evidence part of delivery, not an end report.
Consultancy Mantra’s Data & Insights domain connects measurement, decision support and usable data with the operating context. The CA-PRAXIS capability framework uses agreed evidence to distinguish awareness, guided work, independent practice, reliability and sustained organisational conditions.
The CA-PRAXIS handbook provides a concise evidence view, while Consult → Build → Amplify explains how diagnosis, delivery, transfer and ownership can sit in one engagement. Evidence should remain honest about limitations; a framework does not turn weak data into proof.
Sources and further reading
Primary and authoritative references used in this guide.
Centers for Disease Control and Prevention
CDC Program Evaluation Framework, 2024 — evaluation types, indicators, evidence and the six-step framework.
UK Government
Magenta Book: Central Government guidance on evaluation — evaluation scoping, theory of change and design.
OECD
Applying Evaluation Criteria Thoughtfully — relevance, coherence, effectiveness, efficiency, impact and sustainability.
World Bank and Inter-American Development Bank
Impact Evaluation in Practice, second edition — causal inference, counterfactual methods, sampling and implementation.
International Labour Organization
Human Resources Development Recommendation, 2004 (No. 195) — evaluation of training and lifelong-learning policies against wider human-development goals.
Use the companion guide
Before choosing measures,
choose the intended result.
Review when training is sufficient and when wider capability building is required.
Read training vs capability building →