In our first conversation, we discussed why evaluating public policies can be politically risky. Starting with Simon Schwartzman's provocation about who speaks to the king, we saw that evaluations are neither produced nor received in a neutral space. Those who commission, evaluate, and decide respond to different incentives. For those in power, acknowledging that a policy has failed generates wear and tear and forces them to revise decisions, while the benefits are uncertain and delayed. Therefore, even rigorous evidence can be ignored.
In this second conversation, the focus is on the limits of evaluation: what can it reveal about a public policy? In theory, evaluating means measuring results and separating successful policies from those that failed. In practice, the same policy can be considered successful by one evaluation and unsuccessful by another. One evaluation may show that resources were applied and the target audience reached; another, that the situation of beneficiaries changed little; a third, that the positive effects were limited to the most prepared groups or disappeared when the support ended. What can we conclude?
The uncomfortable answer is that each evaluation can reveal a part of the truth. A public policy produces different results, at different times and for different groups. Some are visible; others remain hidden. Some confirm its objectives; others contradict them. Evaluating, therefore, is not simply finding a ready-made and unequivocal result, but constructing a reasoned judgment about a reality that rarely fits into a single measure.
The seduction of the method
In recent decades, the evaluation of public policies has become methodologically more sophisticated, driven by technological advances, the expansion of databases, and powerful statistical and econometric programs. Randomized experiments, matching methods, difference-in-differences, regression discontinuity, instrumental variables, and other techniques have broadened the ability to separate association from causation: can the observed outcome be attributed to the intervention, or would it have occurred even without it?
These advances helped to curb "gut feeling," a methodology widely used in the analysis of public policies. "Guess feeling" relies on personal impressions, isolated examples, or the opinions of authorities to judge successes and failures. Good methods were, therefore, a necessary advancement.
The problem is confusing, as often happens, sophistication with quality. An estimate can be technically accurate and answer a relatively irrelevant question; it can identify the average effect without revealing who won, who lost, why the result occurred, or whether it could be reproduced under other conditions. It can rigorously measure a small part of the policy and be presented as an objective judgment about the whole.
The incorporation of sophisticated procedures into policy evaluation has brought an ambition characteristic of the scientific method: controlling factors, constructing comparisons, and identifying the specific effect of an intervention. This transposition, however, has limits, because public policies are not medicines, and society is not a laboratory. They traverse institutions, agents, and territories, are interpreted and adapted, and may reach beneficiaries in a way that differs from what was anticipated. Furthermore, control groups may also be affected by the policy or by other interventions. Therefore, an effect observed in a given context does not automatically become a general rule.
Sophisticated assessments are often expensive, time-consuming, and information-intensive. Often, the costs are justified; in other cases, large teams are mobilized to certify conclusions that simpler assessments could have anticipated. In these cases, the method adds precision to plausible conclusions, but not necessarily knowledge proportional to the cost.
This does not mean dismissing evidence, replacing rigorous investigation with guesswork, or disregarding the accumulated knowledge of managers. It simply means that rigor is not synonymous with technical complexity. A good evaluation begins with the relevant question, not the available technique. In certain situations, it will be essential to construct a robust counterfactual; in others, understanding the implementation, listening to the actors, or reconstructing processes will be more useful than estimating an average effect. The method should serve the understanding of the policy, without reducing it to what it can measure.
The time for evaluation is not the time for politics.
Even when the question is relevant and the method appropriate, it is necessary to decide when to evaluate. Public policies do not produce all their effects at the same time, nor do they follow the schedules of governments or research contracts.
Many policies need time to mature; institutions, to gain trust; organizations, to learn; beneficiaries, to adapt their behavior; and investments, to produce results. An early assessment can confuse implementation difficulties with permanent incapacity and condemn an experiment when it begins to accumulate learning.
The opposite is also true. Initial results can disappear when incentives are withdrawn or funding ends. A pilot project, supported by exceptional resources, may not replicate its results when scaled up. There are also policies, such as early childhood education, disease prevention, and scientific research, whose effects appear after years. In other cases, success consists of preventing events, a result that is more difficult to observe.
The timing of the evaluation, therefore, is not a merely technical decision. A policy may appear successful in the first year and disappointing later, or face initial difficulties and produce lasting effects afterward. Photographs taken at different times offer distinct portraits of the same trajectory.
In a recent article in Estadão¹ , Pedro de Camargo Neto places evaluation at the center of a broader challenge: improving the quality of the State, recovering its capacity to invest, and addressing fiscal imbalance. Every resource lost to fraud, waste , or misappropriation fails to finance its intended purpose. Pedro also reminds us that the deviation from objectives rarely happens all at once: it takes hold gradually, through outdated registries, obsolete criteria, incentives that are not reevaluated, and capture by specific interests. Hence the need for permanent monitoring, before deviation becomes the norm. Respecting the maturation period does not mean waiting indefinitely, but monitoring the policy from the beginning, correcting implementation problems, and distinguishing immediate deliverables from lasting impacts.
Proper evaluation protects good policies and requires aligning the evaluation timeline with the policy's implementation. Demanding profound impacts from a newly initiated intervention can be as misguided as celebrating immediate results without questioning their permanence. Monitoring its trajectory also allows one to perceive when it begins to lose its purpose, even if it remains active and continues to allocate resources.
Evaluate policies, not just programs.
Many evaluations focus on programs delimited by target audience, budget, instruments, and time period. This delimitation facilitates research, but it can impoverish understanding because the results rarely depend on a single intervention.
A credit program can distribute the intended resources and have little effect if beneficiaries lack technical assistance, infrastructure, or access to markets. Evaluated in isolation, it may seem well-executed. What remains outside the field of vision is public policy: the set of instruments, institutions, and decisions that should act in a coordinated manner on a given problem.
The program itself is usually easier to measure than the transformation that justified its creation. We can know how many credits were awarded or how many people participated in a course. It is more difficult to determine whether these actions expanded opportunities, reduced vulnerabilities, or permanently changed the situation faced.
A methodologically impeccable evaluation can, therefore, produce a correct but insufficient answer. It may demonstrate that a particular instrument generated a certain effect in a specific group and period, without clarifying whether the policy, considered as a whole, fulfills its purpose. The rigor of the answer does not compensate for the narrowness of the question.
Returning to the limits of evaluation
Recognizing these limitations does not diminish the importance of evaluation. On the contrary, it makes it more useful. Evaluation should not simply mean issuing a judgment – did it work or did it not work – but understanding processes, identifying for whom the results appear, monitoring their permanence, correcting deviations, and recognizing the conditions necessary to expand an experience.
In our first conversation, we discussed the risks and incentives surrounding evaluation. In this one, we saw that sophisticated methods and rigorous estimates are indispensable, but they do not replace formulating the relevant question, understanding the policy, paying attention to the context, and judging the significance of the results.
Even rigorous evaluations produce partial answers, conditioned by the questions, the timing, and the chosen approach. Therefore, a good evaluation is not one that concludes the discussion with a single number, but one that helps the state learn from experience and correct its policies while there is still time.
In the third and final conversation, we will examine, in light of the Brazilian experience, whether the State has the capacity to evaluate policies and incorporate the results into planning, budgeting, and management.
This text does not necessarily reflect the opinion of Unicamp.
1 https://www.estadao.com.br/opiniao/espaco-aberto/integridade-das-politicas-publicas/?srsltid=AfmBOorELgWec6QGqKeki-Hff43B-h3JO9enCDz6wLpsBXxXJPYLm3OX
Cover photo :

