Mixed-motive scenarios are ubiquitous in real-world multi-agent interactions, where self-interested agents often defect for immediate rewards, overlooking the potential of altruistic cooperation to improve long-term gains and collective welfare. Peer punishment can deter defection, but as costly second-order altruism, its persistent imposition may undermine the punisher's interests. Existing approaches often struggle to effectively implement punishment to promote cooperation. To balance the efficacy and cost of punishment, we propose Adaptive Punishment for Cooperation (APC), a distributed method that determines punishment intensity based on both a dynamic punishment probability and the severity of defection. This dynamic probability substantially reduces costly and ineffective punishment while also promotes cooperation. To accurately assess defection and its severity, we use a defection awareness module, whose learning is guided by game reward. Theoretical analysis and empirical results show APC performs effectively in iterated public goods game. Empirically, APC also significantly outperforms existing baselines across sequential social dilemmas, learning rational and effective pu
People increasingly rely on AI-advice when making decisions. At times, such advice can promote selfish behavior. When individuals abide by selfishness-promoting AI advice, how are they perceived and punished? To study this question, we build on theories from social psychology and combine machine-behavior and behavioral economic approaches. In a pre-registered, financially-incentivized experiment, evaluators could punish real decision-makers who (i) received AI, human, or no advice. The advice (ii) encouraged selfish or prosocial behavior, and decision-makers (iii) behaved selfishly or, in a control condition, behaved prosocially. Evaluators further assigned responsibility to decision-makers and their advisors. Results revealed that (i) prosocial behavior was punished very little, whereas selfish behavior was punished much more. Focusing on selfish behavior, (ii) compared to receiving no advice, selfish behavior was penalized more harshly after prosocial advice and more leniently after selfish advice. Lastly, (iii) whereas selfish decision-makers were seen as more responsible when they followed AI compared to human advice, punishment between the two advice sources did not vary. Over
Cooperation in large groups and one-shot interactions is often hindered by freeloading. Punishment can enforce cooperation, but it is usually regarded as wasteful because the costs of punishing offset its benefits. Here, we analyze an evolutionary game model that integrates upstream and downstream reciprocity with costly punishment: integrated strong reciprocity (ISR). We demonstrate that ISR admits a stable mixed equilibrium of ISR and unconditional defection (ALLD), and that costly punishment can become productive: When sufficiently efficient, it raises collective welfare above the no-punishment baseline. ALLD players persist as evolutionary shields, preventing invasion by unconditional cooperation (ALLC) or alternative conditional strategies (e.g., antisocial punishment). At the same time, the mixed equilibrium of ISR and ALLD remains robust under modest complexity costs that destabilize other symmetric cooperative systems.
There are countless examples of how AI can cause harm, and increasing evidence that the public are willing to ascribe blame to the AI itself, regardless of how "illogical" this might seem. This raises the question of whether and how the public might expect AI to be punished for this harm. However, public expectations of the punishment of AI have been vastly underexplored. Understanding these expectations is vital, as the public may feel the lingering effect of harm unless their desire for punishment is satisfied. We synthesise research from psychology, human-computer and -robot interaction, philosophy and AI ethics, and law to highlight how our understanding of this issue is still lacking. We call for an interdisciplinary programme of research to establish how we can best satisfy victims of AI harm, for fear of creating a "satisfaction gap" where legal punishment of AI (or not) fails to meet public expectations.
From ant-acacia mutualism to performative conflict resolution among Inuit, dedicated punishments between distinct subsets of a population are widespread and can reshape the evolutionary trajectory of cooperation. Existing studies have focused on punishments within a homogeneous population, paying little attention to cooperative dynamics in a situation where belonging to a subset is equally important to the actual strategy represented by an actor. To fill this gap, we here study a bipartite population where cooperator agents in a public goods game penalize exclusively those defectors who belong to the alternative subset. We find that cooperation can emerge and remain stable under symmetric intergroup punishment. In particular, at low punishment intensity and at a small value of the enhancement factor of the dilemma game, intergroup punishment promotes cooperation more effectively than a uniformly applied punishment. Moreover, intergroup punishment in bipartite populations tends to be more favorable for overall social welfare. When this incentive is balanced, cooperators can collectively restrain defectors of the alternative set via aggregate interactions in a randomly formed working
While voluntary participation is a key mechanism that enables altruistic punishment to emerge, its explanatory power typically rests on the common assumption that non-participants have no impact on the public good. Yet, given the decentralized nature of voluntary participation, opting out does not necessarily preclude individuals from influencing the public good. Here, we revisit the role of voluntary participation by allowing non-participants to exert either positive or negative impacts on the public good. Using evolutionary analysis in a well-mixed finite population, we find that positive externalities from non-participants lower the synergy threshold required for altruistic punishment to dominate. In contrast, negative externalities raise this threshold, making altruistic punishment harder to sustain. Notably, when non-participants have positive impacts, altruistic punishment thrives only if non-participation is incentivized, whereas under negative impacts, it can persist even when non-participation is discouraged. Our findings reveal that efforts to promote altruistic punishment must account for the active role of non-participants, whose influence can make or break collective o
We introduce a coevolutionary framework in which punishment intensity dynamically adapts to the fraction of cooperators in the population. Unlike static models, adaptive punishment reshapes the effective payoff landscape, driving transitions among canonical games, including the Prisoner's Dilemma, Harmony, Stag Hunt, and Chicken games. Analytical results reveal rich dynamical behaviors such as coexistence, bistability, limit cycle and Hopf bifurcation. These findings highlight adaptive punishment as a robust mechanism for sustaining cooperation by the coevolutionary feedback and offer insights into institutional design, ecological interactions, and social governance.
Punishment as a mechanism for promoting cooperation has been studied extensively for more than two decades, but its effectiveness remains a matter of dispute. Here, we examine how punishment's impact varies across cooperative settings through a large-scale integrative experiment. We vary 14 parameters that characterize public goods games, sampling 360 experimental conditions and collecting 147,618 decisions from 7,100 participants. Our results reveal striking heterogeneity in punishment effectiveness: while punishment consistently increases contributions, its impact on payoffs (i.e., efficiency) ranges from dramatically enhancing welfare (up to 43% improvement) to severely undermining it (up to 44% reduction) depending on the cooperative context. To characterize these patterns, we developed models that outperformed human forecasters (laypeople and domain experts) in predicting punishment outcomes in new experiments. Communication emerged as the most predictive feature, followed by contribution framing (opt-out vs. opt-in), contribution type (variable vs. all-or-nothing), game length (number of rounds), peer outcome visibility (whether participants can see others' earnings), and the
Indirect reciprocity promotes cooperation by allowing individuals to help others based on reputation rather than direct reciprocation. Because it relies on accurate reputation information, its effectiveness can be undermined by information gaps. We examine two forms of incomplete information: incomplete observation, in which donor actions are observed only probabilistically, and reputation fading, in which recipient reputations are sometimes classified as "Unknown". Using analytical frameworks for public assessment, we show that these seemingly similar models yield qualitatively different outcomes. Under incomplete observation, the conditions for cooperation are unchanged, because less frequent updates are exactly offset by higher reputational stakes. In contrast, reputation fading hinders cooperation, requiring higher benefit-to-cost ratios as the identification probability decreases. We then evaluate costly punishment as a third action alongside cooperation and defection. Norms incorporating punishment can sustain cooperation across broader parameter ranges without reducing efficiency in the reputation fading model. This contrasts with previous work, which found punishment ineffe
We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, klDMP. Unlike existing RPRL approaches that optimize reward-seeking and punishment-related policies largely independently, KCPR enables direct interactions between companion policies by treating each as a dynamically learned prior for the other. KCSO yields coupled soft-optimal policies and KL-regularized Bellman operators, allowing reward and punishment information to jointly influence value propagation. To improve learning stability, we introduce a companion-prior softening mechanism and evaluate separate replay-buffer designs for balancing reward- and punishment-related experience. Experiments in grid-world and Gazebo robotic navigation tasks demonstrate that klDMP improves safety and learning stability while maintaining competitive task performance compared with DQN, SQL and softDMP. These results suggest that policy-level coordination provides an effective mechanism for integrating multiple behavioral objectives and may serve as a useful design princi
Indirect punishment traditionally sustains cooperation in social systems through reputation or norms, often by reducing defectors' payoffs indirectly. In this study, we redefine indirect punishment for structured populations as a spatially explicit mechanism, where individuals on a square lattice target second-order defectors--those harming their neighbors--rather than their own immediate defectors, guided by the principle: "I help you by punishing those who defect against you". Using evolutionary simulations, we compare this adapted indirect punishment to direct punishment, where individuals punish immediate defectors. Results show that within a narrow range of low punishment costs and fines, adapted indirect punishment outperforms direct punishment in promoting cooperation. However, outside this cost-fine region, outcomes vary: direct punishment may excel, both may be equally effective, or neither improves cooperation, depending on the parameter values. These findings hold even when network reciprocity alone does not support cooperation. Notably, when adapted indirect punishment outperforms direct punishment in promoting cooperation, defectors face stricter penalties without appr
Punishment is probably the most frequently used mechanism to increase cooperation in Public Goods Games (PGG); however, it is expensive. To address this problem, this paper introduces an optimal control problem that uses fractional punishment to promote cooperation. We present a series of computational experiments illustrating the effects of single and combined terms of the optimization cost function. In the findings, the optimal controller outperforms the use of constant fractional punishment and gives an insight into the period and size of the penalization to be implemented with respect to the defection in the game.
Cooperation is fundamental to human societies, and indirect reciprocity, where individuals cooperate to build a positive reputation for future benefits, plays a key role in promoting it. Previous theoretical and experimental studies have explored both the effectiveness and limitations of costly punishment in sustaining cooperation. While empirical observations show that costly punishment by third parties is common, some theoretical models suggest it may not be effective in the context of indirect reciprocity, raising doubts about its potential to enhance cooperation. In this study, we theoretically investigate the conditions under which costly punishment is effective. Building on a previous model, we introduce a new type of error in perceiving actions, where defection may be mistakenly perceived as cooperation. This extension models a realistic scenario where defectors have a strong incentive to disguise their defection as cooperation. Our analysis reveals that when defection is difficult to detect, norms involving costly punishment can emerge as the most efficient evolutionarily stable strategies. These findings demonstrate that costly punishment can play a crucial role in promoti
Punishment is a common tactic to sustain cooperation and has been extensively studied for a long time. While most of previous game-theoretic work adopt the imitation learning where players imitate the strategies who are better off, the learning logic in the real world is often much more complex. In this work, we turn to the reinforcement learning paradigm, where individuals make their decisions based upon their past experience and long-term returns. Specifically, we investigate the Prisoners' dilemma game with Q-learning algorithm, and cooperators probabilistically pose punishment on defectors in their neighborhood. Interestingly, we find that punishment could lead to either continuous or discontinuous cooperation phase transitions, and the nucleation process of cooperation clusters is reminiscent of the liquid-gas transition. The uncovered first-order phase transition indicates that great care needs to be taken when implementing the punishment compared to the continuous scenario.
We introduce the ``soft Deep MaxPain'' (softDMP) algorithm, which integrates the optimization of long-term policy entropy into reward-punishment reinforcement learning objectives. Our motivation is to facilitate a smoother variation of operators utilized in the updating of action values beyond traditional ``max'' and ``min'' operators, where the goal is enhancing sample efficiency and robustness. We also address two unresolved issues from the previous Deep MaxPain method. Firstly, we investigate how the negated (``flipped'') pain-seeking sub-policy, derived from the punishment action value, collaborates with the ``min'' operator to effectively learn the punishment module and how softDMP's smooth learning operator provides insights into the ``flipping'' trick. Secondly, we tackle the challenge of data collection for learning the punishment module to mitigate inconsistencies arising from the involvement of the ``flipped'' sub-policy (pain-avoidance sub-policy) in the unified behavior policy. We empirically explore the first issue in two discrete Markov Decision Process (MDP) environments, elucidating the crucial advancements of the DMP approach and the necessity for soft treatments o
Taxes are an essential and uniformly applied institution for maintaining modern societies. However, the levels of taxation remain an intensive debate topic among citizens. If each citizen contributes to common goals, a minimal tax would be sufficient to cover common expenses. However, this is only achievable at high cooperation level; hence, a larger tax bracket is required. A recent study demonstrated that if an appropriate tax partially covers the punishment of defectors, cooperation can be maintained above a critical level of the multiplication factor, characterizing the synergistic effect of common ventures. Motivated by real-life experiences, we revisited this model by assuming an interactive structure among competitors. All other model elements, including the key parameters characterizing the cost of punishment, fines, and tax level, remain unchanged. The aim was to determine how the spatiality of a population influences the competition of strategies when punishment is partly based on a uniform tax paid by all participants. This extension results in a more subtle system behavior in which different ways of coexistence can be observed, including dynamic pattern formation owing
Altruistic punishment, where individuals incur personal costs to punish others who have harmed third parties, presents an evolutionary conundrum as it undermines individual fitness. Resolving this puzzle is crucial for understanding the emergence and maintenance of human cooperation. This study investigates the role of an alternative strategy, the exit option, in explaining altruistic punishment. We analyze a two-stage prisoner's dilemma game in well-mixed and networked populations, considering both finite and infinite scenarios. Our findings reveal that the exit option does not significantly enhance altruistic punishment in well-mixed populations. However, in networked populations, the exit option enables the existence of altruistic punishment and gives rise to complex dynamics, including cyclic dominance and bi-stable states. This research contributes to our understanding of costly punishment and sheds light on the effectiveness of different voluntary participation strategies in addressing the conundrum of punishment.
Costly punishment has been suggested as a key mechanism for stabilizing cooperation in one-shot games. However, recent studies have revealed that the effectiveness of costly punishment can be diminished by second-order free riders (i.e., cooperators who never punish defectors) and antisocial punishers (i.e., defectors who punish cooperators). In a two-stage prisoner's dilemma game, players not only need to choose between cooperation and defection in the first stage, but also need to decide whether to punish their opponent in the second stage. Here, we extend the theory of punishment in one-shot games by introducing simple bots, who consistently choose prosocial punishment and do not change their actions over time. We find that this simple extension of the game allows prosocial punishment to dominate in well-mixed and networked populations, and that the minimum fraction of bots required for the dominance of prosocial punishment monotonically increases with increasing dilemma strength. Furthermore, if humans possess a learning bias toward a "copy the majority" rule or if bots are present at higher degree nodes in scale-free networks, the fully dominance of prosocial punishment is sti
We introduce \emph{informational punishment} to the design of mechanisms that compete with an exogenous status quo mechanism: Players can send garbled public messages with some delay, and others cannot commit to ignoring them. Optimal informational punishment ensures that full participation is without loss, even if any single player can publicly enforce the status quo mechanism. Informational punishment permits using a standard revelation principle, is independent of the mechanism designer's objective, and operates exclusively off the equilibrium path. It is robust to refinements and applies in informed-principal settings. We provide conditions that make it robust to opportunistic signal designers.