by Carlos Rodríguez-Ariza
The problem was not the methods
Evaluation faces a paradox that is hard to ignore. We have never had so many methodologies, tools, data sources and technical capabilities. However, the effective influence of evaluation on public decision-making remains uneven and, in many contexts, surprisingly limited, with institutional advances and setbacks, steps forward and backward, even in initiatives considered to be benchmarks.
For years, we have mainly discussed tools, designs and methodologies. However, the most significant challenges facing contemporary evaluation seem to be less and less related to methods and increasingly to the systems within which those methods operate. The problem no longer appears to be essentially technical. It is, above all, a matter of institutional design, governance, strategic relevance and collective capacity to transform evidence into learning, decision-making and action.
For decades, much of the evaluation field relied on a paradigm centred on measurement, standardisation and accountability. That approach enabled significant progress, but it also revealed growing limitations: fragmentation, ritualisation, under-utilisation of findings and a disconnect from actual decision-making processes (Dahler-Larsen, 2012; Bamberger et al., 2019).
More recent studies have identified additional challenges, such as evidence overload, the persistence of organisational incentives that are not conducive to learning, difficulties in addressing complex and adaptive problems, and the vulnerability of evaluation institutions to political and institutional changes (Patton, 2018; Crawford, 2021; OECD, 2023).
Furthermore, contemporary literature suggests that the main challenge no longer lies so much in producing more evidence as in strengthening institutional capacities to interpret it, deliberate on it, integrate it with other sources of knowledge, and use it effectively in contexts characterised by complexity, uncertainty and frequent shifts in priorities (Patton, 2018; OECD, 2023; Salas et al., 2022).
At the same time, public issues have become more complex, interdependent and politically contested. As a result, a fundamental question has resurfaced: what role should evaluation actually play in contemporary governance? (OECD, 2023).
We continue to carry out evaluations (and plan them) in many contexts as if the problems were linear and relatively stable. However, the realities in which we operate are increasingly dynamic, uncertain and interconnected.
Challenges relating to poverty, inequality, climate change, social cohesion, children or digital transformation cannot be properly understood or managed from fragmented perspectives. They require the ability to connect scattered information, integrate diverse perspectives and learn continuously. Therefore, the main challenge is not merely to improve evaluation as an isolated practice, but to rethink it as a system. This article proposes six interconnected ideas:
-
-
- The problem lies in design, not in methods.
- Collaboration is infrastructure (a prerequisite), not a value.
- We continue to evaluate sectors facing systemic problems.
- Evaluation involves redistributing power.
- Without care, there is no sustainability.
- True change consists of shifting the focus from the individual to the system.
-
The thesis underpinning these reflections is simple: evaluation will only be relevant when it ceases to be conceived as an isolated exercise in measurement and comes to be understood as an institutional framework capable of bringing together accountability, learning, decision-making and transformation, linking evidence, public policy, collective capacities and genuine processes of participation. Furthermore, its main contribution will not lie solely in producing knowledge about reality, but in strengthening the capacity of people, organisations and systems to learn, adapt and transform themselves on an ongoing basis.
The real problem is systemic
The debate on evaluation remains too entrenched in familiar dichotomies: learning versus accountability; qualitative versus quantitative; experimental versus adaptive; rigour versus use. These discussions remain important, but they are no longer sufficient. The accumulated evidence shows that many evaluations fail not because of a lack of methodological quality, but because they operate within systems that do not foster learning, use or evidence-based decision-making (Patton, 2018; Sanderson, 2002). It is common to find technically sound evaluations that have little influence on decisions, processes geared more towards accountability than learning, and an abundance of evidence with little political traction.
Relevance is not built at the end of an evaluation. It is built from the outset, when evaluation plans are designed that are strategically relevant and consistent with policies and strategies, and when evaluation questions, processes and mechanisms for utilising findings are linked to real priorities. This involves engaging, from the very beginning, those who produce, use and experience the policies and programmes being evaluated. Relevance stems both from technical quality and from the ability to integrate diverse knowledge, experiences and perspectives.
Collaboration is infrastructure (a prerequisite)
One of the most widely accepted views in evaluation is that we need to collaborate more. However, complex systems do not collaborate simply because their actors wish them to. They collaborate when there are incentives, rules, structures and capacities that make such collaboration possible (Axelrod, 1984; Ostrom, 1990). When these conditions are not present, non-collaboration (or the protection of resources, information and spheres of influence) often constitutes a rational response from the perspective of the actors involved.
When these elements are absent, recurring patterns emerge, such as organisational competition, institutional fragmentation, power asymmetries and conflicting individual incentives. Effective collaboration requires a redesign of evaluation systems to promote:
-
-
- the joint development of questions and learning agendas;
- the shared production of evidence;
- the integration of knowledge derived from evaluation, monitoring, research, administrative data, knowledge management and other institutional functions;
- the collective interpretation of results;
- the coordinated use of findings;
- the strengthening of evaluation and evidence-use capacities across the entire ecosystem, beyond evaluation units.
-
From this perspective, collaboration is not merely about coordinating actors or sharing information. It involves creating the conditions for different organisations, disciplines, functions and communities of practice to contribute jointly to the generation, interpretation and use of knowledge. The strength of an evaluation ecosystem depends less on the quality of each individual component than on the quality of the connections between them.
Participation is not merely an ethical or democratic issue. It is also a functional prerequisite for generating collective learning. Systems learn best when they incorporate diverse perspectives, distributed knowledge and complementary experiences. Without meaningful participation, evaluation runs the risk of producing technically correct but socially irrelevant responses.
In more mature evaluation ecosystems, the central question shifts from who produces the evidence to how different actors and functions collectively contribute to generating insights for public action. There is also a fundamental tension between independence and collaboration. The credibility of evaluation requires autonomy, but its use demands continuous interaction with those who design and implement policies and programmes. The challenge is not to choose between the two, but to create institutional conditions that allow them to be reconciled.
This involves broadening evaluative thinking beyond formal evaluations and strengthening capacities to generate, interpret and use evidence through complementary functions such as planning, monitoring, applied research, knowledge management, organisational learning, data analytics, citizen feedback mechanisms and foresight.
From this perspective, the ultimate aim is not to produce more evaluations, but to increase the collective capacity to generate, integrate and use evidence in order to learn, adapt and make better decisions. Evaluation is an essential part of this ecosystem, but it is not the only source of knowledge relevant to public policy (Nutley et al., 2007; Parkhurst, 2017).
We continue to evaluate sectors facing systemic problems
The main contemporary challenges (inequality, poverty, climate change, child wellbeing, migration and social cohesion) do not respect sectoral boundaries. However, evaluation systems are still frequently organised around isolated programmes, fragmented indicators and disconnected institutional structures. This tension becomes particularly apparent when we attempt to incorporate cross-cutting approaches relating to gender, equity, human rights or participation (Kabeer, 1999).
The problem is not the absence of these approaches. The problem is that all too often they continue to function as additional layers rather than as organising principles for analysis and action. Evaluation needs to move from cross-cutting approaches to integration; from silos to systems; and from the aggregation of results to an understanding of interdependencies. Understanding complex problems also requires incorporating the diversity of experiences of those who live through them from different social, territorial and institutional positions.
Evaluation involves redistributing power
Contemporary institutions face growing challenges to their legitimacy. At the same time, new forms of knowledge production, citizen participation and collective intelligence are emerging. In this context, evaluation cannot be limited to verifying results. It must contribute to strengthening spaces for deliberation, learning and informed decision-making. This involves recognising the plurality of knowledge and valuing situated knowledge.
Redistributing power does not merely mean broadening participation in data collection. It means expanding the capacity of different actors to formulate questions, interpret evidence and participate in decision-making. Participation thus ceases to be a specific technique and becomes a constitutive dimension of more democratic and legitimate evaluation systems.
However, this aspiration raises a significant tension. For decades, evaluation has promoted participation as a central principle, yet many of the organisations in which it operates continue to be characterised by hierarchical structures, asymmetrical flows of information and vertical accountability mechanisms (Cornwall, 2008; Parkhurst, 2017). In these contexts, participation risks becoming a symbolic exercise if it is not accompanied by genuine opportunities to influence the questions asked, the evidence considered and the decisions taken.
The issue is not simply to involve more stakeholders in the evaluation process, but to promote forms of participation that are relevant, functional and proportionate to the aims of the evaluation. Participation generates value when it enhances the quality of learning, improves understanding of complex problems, incorporates diverse perspectives and strengthens the legitimacy of decisions; not when it becomes a procedural requirement devoid of effective influence (Chambers, 1997; Cornwall, 2008).
From this perspective, evaluation also involves democratising the production and use of knowledge. This entails creating conditions in which different stakeholders can contribute meaningfully to the generation of evidence, the interpretation of findings and the development of collective responses to complex public problems.
Artificial intelligence offers new opportunities to expand analytical capabilities, process large volumes of information and facilitate new ways of generating evidence. However, it can also reinforce pre-existing dynamics of power concentration, opacity, bias and exclusion if the data, algorithms and decision-making criteria remain beyond public scrutiny or pluralistic debate (Crawford, 2021). In this sense, the challenges posed by artificial intelligence are not merely technological, but also institutional, ethical and democratic.
The central question is not who has access to more data or better algorithms, but who is involved in defining the questions, interpreting the evidence and using the results. Ultimately, the challenges posed by artificial intelligence bring into even sharper focus a fundamental issue for contemporary evaluation ecosystems: who has the capacity to influence what knowledge is produced, what evidence is considered valid and how it is used to guide public action.
These tensions reinforce the need to design evaluation ecosystems capable of combining technological innovation, a diversity of perspectives and appropriate mechanisms for governance, participation and accountability.
Without care, there is no sustainability
Evaluation systems increasingly operate in contexts characterised by pressure, acceleration and institutional fatigue. Organisational overload affects both those who produce evidence and those who must use it. Under these conditions, it is difficult to build genuine processes of learning, reflection and adaptation.
Care is not an incidental element. It constitutes a fundamental infrastructure for institutional sustainability (Puig de la Bellacasa, 2017). Without care, evaluation can become an extractive practice: it generates information, but does not necessarily strengthen capacities or bring about transformation.
Participation also runs this risk when people are invited to contribute without having the time, resources or appropriate conditions to do so meaningfully. Caring for systems also involves caring for those who participate in them.
However, there is a growing risk of interpreting care exclusively as an individual responsibility. The rise in initiatives focused on wellbeing, resilience or self-care has helped to highlight real problems, but may prove insufficient when the causes of distress are primarily organisational rather than individual.
Many of the challenges affecting people’s wellbeing today have systemic roots: excessively hierarchical structures, a lack of top-down accountability, deficits in internal communication, persistent power imbalances, poorly managed change processes, prolonged organisational uncertainty, or institutional cultures that prioritise constant urgency over reflection and learning. These dynamics not only affect people’s wellbeing; they also reduce organisations’ capacity to learn, innovate and adapt (Edmondson, 2019).
From this perspective, care cannot be limited to helping people adapt better to dysfunctional organisations. It must also encompass the ability of organisations to critically examine their own dynamics, identify structural sources of burnout and create healthier conditions for collaboration, learning and decision-making.
Psychological safety—understood as the ability to express doubts, mistakes, disagreements or ideas without fear of reprisal—is an essential condition for systems to be able to learn continuously (Edmondson, 2019). Similarly, learning organisations require spaces for collective reflection, feedback and the questioning of established assumptions (Senge, 2006).
The challenge lies not merely in promoting individual wellbeing, but in building organisational wellbeing. In other words, moving from an approach centred on the resilience of individuals to one centred on the resilience of systems.
Because, ultimately, sustainable evaluation systems do not depend solely on resilient individuals, but on organisations capable of caring for, learning from and continuously transforming themselves.
From the individual to the system
For decades, much of the effort to strengthen evaluation focused on training. Training remains important. But the accumulated evidence shows that robust evaluation systems require much more than individual training (OECD, 2023; UNDP, 2021).
For too long, we have assumed that improving people’s capabilities would automatically lead to better systems. However, accumulated experience suggests that individual knowledge rarely translates on its own into sustainable institutional change.
People can learn new methodologies, approaches or tools. But if organisations lack incentives to use evidence, if leaders do not demand it, if decision-making processes remain disconnected from learning, or if institutional structures hinder collaboration, much of that knowledge ends up being underutilised.
This reflection is consistent with recent trends in capacity-building for evaluation. Various frameworks promoted by UNEG and other international actors have highlighted the need to move beyond approaches focused exclusively on individuals, in order to simultaneously incorporate organisational capacities and enabling environments capable of sustaining learning, the use of evidence and informed decision-making (Baser & Morgan, 2008; Preskill & Boyle, 2008; UNDP, 2021).
Today, there is a need for—or a requirement for—policy frameworks, stable resources, institutional demand for evidence, coordination mechanisms, sustained organisational capacities, leadership and an institutional culture geared towards learning, and enabling environments that foster the use of evidence. Experience in Latin America confirms that the most significant progress has not stemmed solely from methodological innovations, but from efforts to strengthen institutions, build collective capacities and promote the systematic use of evidence in decision-making (Salas et al., 2022). Furthermore, many of the most sustainable advances have arisen from the creation of communities of practice, learning networks and national evaluation systems capable of bringing together multiple stakeholders around shared objectives.
Ultimately, the challenge is no longer simply to train better evaluators. It is to build organisations that learn better, institutions that make better use of evidence, and systems that enable knowledge to be transformed into decisions, adaptation and transformation. Because the strength of an evaluation system does not depend solely on how much its individual members know, but on how much the system as a whole is capable of learning.
Conclusion: what is really at stake
If we consider these six ideas together, one conclusion emerges: the main challenges of contemporary evaluation are less methodological than we tend to assume.

Reimagining evaluation means moving away from viewing it solely as a technical function and beginning to design it as an institutional infrastructure for learning, deliberation and transformation. Reimagining evaluation also means democratising it. Not merely so that more people take part in evaluations, but so that more stakeholders can influence the questions asked, the evidence deemed relevant, the interpretation of findings and the decisions ultimately taken. Because, ultimately, evaluation does not contribute to change through the reports it produces, but through the conversations it facilitates, the learning it generates and the collective transformations it helps to mobilise.
Experience in Latin America over recent decades shows that the most sustainable advances in evaluation have rarely stemmed solely from methodological innovations. They have arisen, above all, from efforts to strengthen institutions, build collective capacities, promote the use of evidence, consolidate communities of practice and broaden the participation of multiple stakeholders in learning and decision-making processes.
Seen from this perspective, the central question is no longer how to produce better evaluations. The question is how to build systems that learn better; that are capable of linking evidence and decision-making, where accountability and learning are intertwined; systems that combine independence and collaboration; that integrate innovation, participation and governance; and that transform information into collective intelligence for public action. Because, ultimately, the challenge of evaluation does not lie solely in producing knowledge about change. It lies in helping to create the conditions that enable people, organisations and societies to learn, adapt and transform themselves on an ongoing basis.
Questions to open the debate
-
-
- Are we designing evaluations, evaluation systems or ecosystems of evidence, learning and decision-making?
- What institutional conditions enable evidence to genuinely influence decisions, rather than merely appearing in reports?
- How can we build forms of collaboration that simultaneously strengthen the independence, credibility and use of evaluation?
- How can we integrate cross-cutting approaches such as gender, equity, rights, participation or sustainability without creating new institutional fragmentation?
- What does redistributing power in evaluation actually entail, and who is involved in defining the questions, interpreting the evidence and using the results?
- How can we move from participation as a procedural requirement to participation as a mechanism for collective learning and transformation?
- What capacities do organisations need to learn continuously and turn evidence into action?
- How can we build systems capable of learning, adapting and making decisions in contexts of increasing complexity, uncertainty and change?
-

REFERENCES:
Axelrod, R. (1984). The evolution of cooperation. Basic Books.
Bamberger, M., Vaessen, J., & Raimondo, E. (2019). Dealing with complexity in development evaluation. SAGE.
Baser, H., & Morgan, P. (2008). Capacity, change and performance: Study report. European Centre for Development Policy Management (ECDPM).
Chambers, R. (1997). Whose Reality Counts? Putting the First Last. Intermediate Technology Publications.
Cornwall, A. (2008). Unpacking Participation: Models, Meanings and Practices. Community Development Journal, 43(3), 269–283.
Crawford, K. (2021). Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press.
Dahler-Larsen, P. (2012). The evaluation society. Stanford University Press.
Edmondson, A. C. (2019). The fearless organization: Creating psychological safety in the workplace for learning, innovation, and growth. Wiley.
Kabeer, N. (1999). Resources, agency, achievements: Reflections on the measurement of women’s empowerment. Development and Change, 30(3), 435–464.
Nutley, S., Walter, I., & Davies, H. T. O. (2007). Using Evidence: How Research Can Inform Public Services. Policy Press.
OECD. (2023). Evaluation insights: Strengthening the use of evidence. OECD Publishing.
Ostrom, E. (1990). Governing the commons. Cambridge University Press.
Parkhurst, J. (2017). The Politics of Evidence: From Evidence-Based Policy to the Good Governance of Evidence. Routledge.
Patton, M. Q. (2018). Principles-focused evaluation. Guilford Press.
Patton, M. Q. (2011). Developmental Evaluation. Guilford Press.
Preskill, H., & Boyle, S. (2008). A multidisciplinary model of evaluation capacity building. American Journal of Evaluation, 29(4), 443–459.
Puig de la Bellacasa, M. (2017). Matters of care. University of Minnesota Press.
Salas, N., et al. (2022). Sistemas nacionales de evaluación en América Latina.
Sanderson, I. (2002). Evaluation, policy learning and evidence-based policy making. Public Administration, 80(1), 1–22.
Senge, P. (1990). The Fifth Discipline: The Art and Practice of the Learning Organization. Doubleday.
UNDP. (2021). Evaluation guidelines. United Nations Development Programme.