Home / Insights / Evidence for capability building

Capability building guide

Evidence for capability building: match proof to the claim.

Participation, learning, workplace practice, reliable performance and system continuity are different results. Measure them separately.

The short answer

Good evidence answers a defined decision.

Start by writing the claim you may need to make. If the claim is that people attended, an attendance record is relevant. If the claim is that they can perform a practice at work, attendance is not enough; the evidence must show the practice in an authentic or realistic context. If the claim is that an organisation can sustain the practice, evidence must also cover roles, routines, tools, governance and ownership.

An evidence plan is therefore not a list of every available metric. It is a small, traceable set of questions, indicators, sources, responsibilities and review points tied to an intended use.

Begin with a claim

Separate description, contribution and attribution.

Different claims require different designs. Clarity at this stage protects teams from over-interpreting weak evidence later.

Description

What happened, for whom and in what context?

Descriptive evidence can document participation, delivery, assessment results, observed practice or system conditions. It can show a pattern without establishing why that pattern occurred.

Contribution

Is the programme a credible part of the explanation?

A contribution claim tests the intended pathway, assumptions, implementation quality, contextual factors and plausible alternatives. Multiple sources can strengthen or weaken the explanation.

Attribution

Did the intervention cause the observed difference?

A causal claim needs a design capable of estimating what would have happened without the intervention. The World Bank’s Impact Evaluation in Practice explains counterfactual methods, sampling, data collection and evaluation management.

An evidence ladder

Collect evidence at the level you intend to claim.

The levels below are a practical synthesis for capability work. They are not interchangeable, and a higher level does not make lower-level operational data unnecessary.

1. Reach and participation

Examples include invitations, enrolment, attendance, completion and participant characteristics. These records help teams understand who was reached and whether access differed across relevant groups. They do not show what participants learned or used.

2. Learning and demonstrated task performance

Use a knowledge check, simulation, demonstration, assessed assignment or structured rubric aligned with the stated learning outcome. A realistic task is usually more informative than recall alone when the intended result is applied practice.

3. Application in the work context

Review observations, work products, decisions, workflow records or other traces of the practice. State the quality standard and support available. Self-report can add context, but it should not be the only evidence for a claim about observed performance.

4. Reliability over time and situations

One successful task shows possibility, not yet dependable performance. Review repeated evidence across an appropriate period or set of cases. Record exceptions and variation rather than hiding them in a single average.

5. Organisational conditions and continuity

Look for active ownership, competent support roles, usable tools, review routines, decision rights, resources and governance. Evidence should show that these conditions operate, not merely that a document says they should.

6. Outcomes and wider effects

Specify the result for people, services or the organisation and the time needed for it to emerge. Consider intended and unintended effects. Be explicit about whether the evaluation describes change, supports a contribution story or estimates causal impact.

Design from a theory of change

Make the expected route and assumptions visible.

The 2026 UK government Magenta Book advises defining the intervention and developing a theory of change at the start of evaluation scoping. It describes a chain from inputs and activities through outputs and outcomes, alongside mechanisms, assumptions, context and alternative explanations.

For capability building, this means stating why the learning and support are expected to change practice, why that practice should affect an outcome, and what else must be true. The map should be revised when evidence challenges an assumption.

Inputs and activities

People, time, content, tools, facilitation, coaching and operational changes. Evidence at this stage supports delivery management, not an outcome claim.

Outputs

Completed sessions, resources, assessments, support contacts or revised workflows. Outputs show what the programme produced and for whom.

Intermediate outcomes

Learning, confidence when appropriately measured, application, quality or adoption. These can provide earlier feedback while longer-term results are still developing.

Final outcomes and assumptions

The intended change in performance, service or organisational function, together with the contextual conditions and causal assumptions connecting each stage.

Build the evidence plan

Six fields make every measure easier to interpret.

Question

What decision will this evidence inform?

Name the user and timing. A measure without a decision owner often becomes reporting activity with no clear consequence.

Indicator

What observable sign represents the result?

The updated CDC Program Evaluation Framework describes indicators as measurable statements that bridge broad constructs and specific measures.

Source

Where will the evidence come from?

Identify the assessment, work product, observation, interview, system record or survey. Check whether the source can answer the question without creating unreasonable burden.

Standard

What counts as sufficient performance?

Define quality, completeness, timing, independence and acceptable support. Where judgement is involved, use a rubric and examples to improve consistency.

Timing

When can the result reasonably be observed?

Collecting too early can mistake recall for transfer. Waiting until the end can remove the opportunity to improve delivery. Use review points that match the expected pathway.

Responsibility

Who collects, interprets and acts?

Separate data ownership, assessment and decision roles where risk warrants it. Record access, consent, privacy, retention and quality requirements before collection.

Avoid common evidence errors

Do not ask one measure to prove a different result.

Attendance presented as capability

Attendance proves presence under the stated recording method. Pair it with task, application or system evidence when the claim goes further.

Confidence treated as competence

Self-reported confidence can reveal perception and help target support. It may not match performance. Use it alongside observed or assessed evidence and investigate meaningful differences.

A test that does not resemble the work

A recall quiz may be valid for factual knowledge but weak for judgement, tool use or collaboration. Match the assessment method to the capability definition and intended context.

Only averages are reported

Averages can hide variation by role, location, starting point or access. Use disaggregation where it is relevant, ethical and statistically responsible, while protecting privacy.

Outcome change is presented as programme impact

The CDC framework distinguishes outcome evaluation, which measures achieved outcomes but cannot attribute causality, from impact evaluation, which compares outcomes with and without the intervention. Use causal language only when the design supports it.

Make evaluation proportionate

Evidence quality depends on use, context and risk.

The OECD guidance on evaluation criteria presents relevance, coherence, effectiveness, efficiency, impact and sustainability as complementary lenses and stresses that they should be adapted to purpose and context. A small internal pilot and a high-stakes public programme should not automatically use the same design.

Proportionality does not mean accepting evidence that cannot answer the question. It means narrowing the question, selecting a defensible method, documenting limitations and spending effort where uncertainty matters most.

For improvement

Use rapid, timely evidence on participation, task performance, barriers and implementation. The priority is to learn early enough to change the programme or workplace support.

For assurance

Use agreed standards, traceable records, clear assessment roles and an appropriate degree of independence. Retain enough context for a reviewer to understand what the evidence does and does not establish.

For impact decisions

Seek evaluation expertise early. Define the counterfactual question, ethical constraints, sample, data access and decision timetable before implementation removes feasible design options.

Apply the evidence model

Make evidence part of delivery, not an end report.

Consultancy Mantra’s Data & Insights domain connects measurement, decision support and usable data with the operating context. The CA-PRAXIS capability framework uses agreed evidence to distinguish awareness, guided work, independent practice, reliability and sustained organisational conditions.

The CA-PRAXIS handbook provides a concise evidence view, while Consult → Build → Amplify explains how diagnosis, delivery, transfer and ownership can sit in one engagement. Evidence should remain honest about limitations; a framework does not turn weak data into proof.

Sources and further reading

Primary and authoritative references used in this guide.

Use the companion guide

Before choosing measures,
choose the intended result.

Review when training is sufficient and when wider capability building is required.

Read training vs capability building →