Evaluation and learning / 10 minutes
Turning evaluation findings into programme decisions
Learn how to design evaluations for use, connect evidence to conclusions, and turn recommendations into owned programme decisions.

An evaluation can meet its terms of reference, answer its questions, and still have little influence on a programme. The report may arrive after a planning decision. Recommendations may be too broad to assign. Stakeholders may dispute the interpretation. A long document may circulate without a clear moment for response.
These are not only communication failures. They often begin in the evaluation design.
Use does not mean forcing every finding into an immediate action. It means creating a credible, timely, and responsible route from evidence to judgement, choice, ownership, and follow-through.
Define the decision before defining the evaluation
“Assess programme performance” is not a sufficient use case. A stronger starting point names the decisions and users.
For example:
- A programme team must decide which delivery components to adapt for the next phase
- Leadership must decide whether a model is ready for expansion
- A public institution must decide which implementation constraint requires priority attention
- Partners must decide how responsibilities or resources should change
- A monitoring team must decide which assumptions need closer testing
The evaluation can serve accountability, learning, strategy, funding, and programme improvement at the same time, but tensions between these purposes should be made visible. Different users may require different questions, levels of independence, timing, and outputs.
Create a use map that lists each intended user, the decision they face, when it will occur, and the evidence they need. If a question has no identifiable user or decision, test whether it belongs in the evaluation.
Choose criteria thoughtfully
The OECD guidance on applying evaluation criteria thoughtfully explains that relevance, coherence, effectiveness, efficiency, impact, and sustainability should not be applied as a mechanical checklist. Their meaning and use depend on the intervention, purpose, users, context, and stage.
That matters for decision use.
Relevance can help determine whether objectives and design continue to respond to needs, policies, and priorities. Coherence can examine how the intervention fits with other actions, institutions, or systems. Effectiveness can assess progress toward intended results and the factors shaping it. Efficiency can consider how resources become results, including timeliness and tradeoffs. Impact can examine significant positive or negative, intended or unintended effects. Sustainability can explore whether benefits or capacities are likely to continue and under which conditions.
Not every evaluation needs equal treatment of all six. Select criteria because they clarify a decision. Then translate them into specific questions and judgement standards.
For example, “Was the programme effective?” is too broad. A useful question might ask which delivery components contributed to a defined outcome for which groups, and what contextual factors enabled or constrained that contribution.
Build a visible line from evidence to recommendation
Decision makers should be able to follow the reasoning.
A useful chain is:
Evidence: What do the verified sources show?
Finding: What pattern or observation emerges?
Interpretation: What does the pattern mean in context?
Conclusion: What reasoned judgement answers the evaluation question?
Recommendation: What response follows, for whom, and why?
Decision: What will the responsible actor do, defer, test, or decline?
When these stages are compressed, recommendations can feel subjective. When they are separated, users can identify where they agree or disagree.
An evaluation matrix can protect this chain. It links questions to criteria, indicators or lines of inquiry, sources, methods, analysis, and judgement standards. It should remain active during analysis rather than being filed after inception.
Use triangulation to explain, not only confirm
Programme decisions often involve evidence that does not align neatly. Monitoring records may show strong coverage while participant interviews describe access barriers. Staff may report consistent implementation while observation suggests local adaptation. An average result may hide substantial differences across locations or groups.
Triangulation should ask:
- Do sources address the same concept and period?
- Whose perspective does each source represent?
- What incentives or measurement conditions may affect the answer?
- Does disagreement reveal variation rather than error?
- Which source is best suited to the specific claim?
The goal is not to force convergence. Divergence may be the finding that matters most for a decision.
Quantitative estimates can describe magnitude or distribution. Qualitative evidence can help explain mechanisms, experience, and context. Programme records can show implementation processes. External evidence can test assumptions about the wider environment. Each has a role and a limit.
Bring stakeholders into interpretation without surrendering independence
Stakeholder participation can improve accuracy, context, ownership, and feasibility. It can also create pressure to soften findings or negotiate away inconvenient evidence.
Separate factual correction, contextual interpretation, and evaluative judgement.
Stakeholders should be able to identify factual errors, missing records, misunderstood processes, or contextual changes. They can offer alternative explanations and test whether a recommendation is operationally realistic. They should not be able to remove a supported finding simply because it is uncomfortable.
A structured sensemaking session can ask:
- What confirms or challenges existing understanding?
- Which finding matters most for upcoming decisions?
- Where do interpretations differ, and why?
- What can be acted on now?
- What requires more evidence or authority?
- What risks could follow from a proposed response?
Document important disagreements. Consensus is not always possible or desirable.
Protect people and data when communicating findings
Evaluation use must remain ethical. A compelling finding is not automatically safe to publish.
The OCHA Data Responsibility Guidelines and the ICRC Handbook on Data Protection in Humanitarian Action reinforce responsible handling of data in humanitarian contexts. The UNICEF Innocenti approach to ethical evidence generation places ethical reflection across evidence work.
Before sharing a finding, consider:
- Could a quote, location, small subgroup, or combination of attributes identify someone?
- Could publication stigmatize or expose a community, service, or staff group?
- Does the audience need the underlying detail to make the decision?
- Can the same point be communicated with aggregation or careful paraphrase?
- Are consent and original collection purposes consistent with the proposed use?
- Who should receive the full report, a restricted annex, or a public summary?
Data responsibility does not end when analysis is complete. Reporting, presentation, storage, sharing, and archiving all matter.
Write recommendations that can be decided
Weak recommendations often use verbs such as “strengthen,” “improve,” “enhance,” or “ensure” without specifying the change. They describe a desirable direction but not a decision.
A decision-ready recommendation should identify:
- The issue and evidence behind it
- The actor with authority to respond
- The proposed change or choice
- The intended result
- Important dependencies or risks
- A realistic time horizon
- Whether action is urgent, phased, or conditional
For example, “Improve monitoring” is not assignable. A stronger recommendation could identify the programme role responsible for revising two priority indicators, the decision those indicators must inform, the required data source, and the next review point.
Recommendations should also be proportionate to the evidence. A suggestive pattern may justify testing or closer monitoring, not a major redesign. A robust and repeated finding may support a firmer decision.
Limit the number of recommendations. A long list shifts the burden of prioritization back to the client. Group related actions, distinguish critical from desirable changes, and state tradeoffs.
Create a management response that records real choices
A management response should be more than an agreement column.
For each recommendation, record:
- Accepted, partly accepted, deferred, or not accepted
- The reason for the decision
- The responsible owner
- The action to be taken
- The expected timing or review point
- Dependencies and required resources
- How progress will be evidenced
Not accepting a recommendation can be legitimate. The important point is to make the reasoning visible. Circumstances may have changed, authority may sit elsewhere, costs may outweigh likely benefit, or the evidence may not be sufficient for the proposed action.
Partial acceptance should specify which part will be acted on. “Noted” should not be used as a substitute for a decision.
Match the output to the user
A single long report rarely serves every audience.
The technical report should preserve methods, evidence, limitations, reasoning, and supporting detail. A decision brief can focus on the choices, evidence, risks, and next steps. A presentation can support discussion. A workshop can test interpretation and ownership. A verified dataset or analytical annex can support appropriate secondary use. A public summary can communicate findings without exposing sensitive details.
Different formats must remain consistent. A short brief should simplify the presentation, not overstate the evidence.
Timing matters as much as format. Interim sensemaking may be appropriate when a decision cannot wait for final layout, provided findings are sufficiently checked and clearly labelled. Surprising or sensitive findings should not first reach responsible leaders through a public release.
Connect evaluation to programme learning
An evaluation is one moment in a wider evidence system. Its findings should connect to monitoring, reflection, planning, budgeting, risk review, and future research.
Ask:
- Which assumptions need continued monitoring?
- Which indicators need revision?
- Which questions remain unanswered?
- What adaptation should be tested rather than adopted immediately?
- Which finding should inform strategy or policy beyond the programme?
- When will the management response be reviewed?
This turns evaluation from a closing event into a source of disciplined learning.
The evaluation should also leave an audit trail. Future teams need to understand how conclusions were reached, which data and documents were used, what limitations applied, and why recommendations were accepted or rejected.
A decision-use test
Before finalizing an evaluation plan, check whether:
- Intended users and decisions are explicit
- Questions are prioritized around those decisions
- Criteria are selected thoughtfully
- Judgement standards are clear enough to support conclusions
- Evidence sources and limitations are mapped
- Ethical and data-sharing responsibilities cover reporting and use
- Stakeholders have defined roles in factual review and interpretation
- Independence and disagreement are protected
- Recommendations will name an actor and practical response
- A management-response process is scheduled
- Outputs match different users without changing the evidence
- Follow-through will connect to programme routines
If these conditions are absent, a better dissemination event will not solve the problem. Use must be designed into the evaluation.
Final principle
Evaluation findings become useful when the people with authority can understand the evidence, see its limits, examine the reasoning, make an explicit choice, and own the next action.
The evaluator's role is not to make every programme decision. It is to provide a credible path from question to evidence and from evidence to accountable judgement.
Further reading
- OECD: Applying Evaluation Criteria Thoughtfully
- OCHA: Data Responsibility Guidelines
- UNICEF Innocenti: Ethical Evidence Generation
- ICRC: Handbook on Data Protection in Humanitarian Action
Closing call to action
Heading: Planning an evaluation that must lead to a decision?
Body: ERC can help frame the users, questions, evidence chain, interpretation process, and decision route from the start.
Action: Discuss an evaluation