Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is significantly less sample efficient than off-policy methods in multi-agent systems. In this work, we carefully study the performance of PPO in cooperative multi-agent settings. We show that PPO-based multi-agent algorithms achieve surprisingly strong performance in four popular multi-agent testbeds: the particle-world environments, the StarCraft multi-agent challenge, Google Research Football, and the Hanabi challenge, with minimal hyperparameter tuning and without any domain-specific algorithmic modifications or architectures. Importantly, compared to competitive off-policy methods, PPO often achieves competitive or superior results in both final returns and sample efficiency. Finally, through ablation studies, we analyze implementation and hyperparameter factors that are critical to PPO's empirical performance, and give concrete practical suggestions regarding these factors. Our results show that when using these practices, simple PPO-based methods can be a strong baseline in cooperative multi-agent reinforcement learning. Source code is released at \url{https://github.com/marlbenchmark/on-policy}.
Abstract The James Webb Space Telescope is revealing a new population of dust-reddened broad-line active galactic nuclei (AGN) at redshifts z ≳ 5. Here we present deep NIRSpec/Prism spectroscopy from the Cycle 1 Treasury program Ultradeep NIRSpec and NIRCam ObserVations before the Epoch of Reionization (UNCOVER) of 15 AGN candidates selected to be compact, with red continua in the rest-frame optical but with blue slopes in the UV. From NIRCam photometry alone, they could have been dominated by dusty star formation or an AGN. Here we show that the majority of the compact red sources in UNCOVER are dust-reddened AGN: 60% show definitive evidence for broad-line H α with a FWHM > 2000 km s −1 , 20% of the current data are inconclusive, and 20% are brown dwarf stars. We propose an updated photometric criterion to select red z > 5 AGN that excludes brown dwarfs and is expected to yield >80% AGN. Remarkably, among all z phot > 5 galaxies with F277W – F444W > 1 in UNCOVER at least 33% are AGN regardless of compactness, climbing to at least 80% AGN for sources with F277W – F444W > 1.6. The confirmed AGN have black hole masses of 10 7 –10 9 M ⊙ . While their UV luminosities (−16 > M UV > −20 AB mag) are low compared to UV-selected AGN at these epochs, consistent with percent-level scattered AGN light or low levels of unobscured star formation, the inferred bolometric luminosities are typical of 10 7 –10 9 M ⊙ black holes radiating at ∼10%–40% the Eddington limit. The number densities are surprisingly high at ∼10 −5 Mpc −3 mag −1 , 100 times more common than the faintest UV-selected quasars, while accounting for ∼1% of the UV-selected galaxies. While their UV faintness suggests they may not contribute strongly to reionization, their ubiquity poses challenges to models of black hole growth.
The problem of finding a specified pattern in a time series database (i.e. query by content) has received much attention and is now a relatively mature field. In contrast, the important problem of enumerating all surprising or interesting patterns has received far less attention. This problem requires a meaningful definition of "surprise", and an efficient search technique. All previous attempts at finding surprising patterns in time series use a very limited notion of surprise, and/or do not scale to massive datasets. To overcome these limitations we introduce a novel technique that defines a pattern surprising if the frequency of its occurrence differs substantially from that expected by chance, given some previously seen data.
Assumptions regarding the importance of empathy are pervasive. Given the impact these assumptions have on research, assessment, and treatment, it is imperative to know whether they are valid. Of particular interest is a basic question: Are deficits in empathy associated with aggressive behavior? Previous attempts to review the relation between empathy and aggression yielded inconsistent results and generally included a small number of studies. To clarify these divergent findings, we comprehensively reviewed the relation of empathy to aggression in adults, including community, student, and criminal samples. A mixed effects meta-analysis of published and unpublished studies involving 106 effect sizes revealed that the relation between empathy and aggression was surprisingly weak (r = -.11). This finding was fairly consistent across specific types of aggression, including verbal aggression (r = -.20), physical aggression (r = -.12), and sexual aggression (r = -.09). Several potentially important moderators were examined, although they had little impact on the total effect size. The results of this study are particularly surprising given that empathy is a core component of many treatments for aggressive offenders and that most psychological disorders of aggression include diagnostic criteria specific to deficient empathic responding. We discuss broad conclusions, consider implications for theory, and address current limitations in the field, such as reliance on a small number of self-report measures of empathy. We highlight the need for diversity in measurement and suggest a new operationalization of empathy that may allow it to synchronize with contemporary thinking regarding its role in aggressive behavior.
Computer scientists have recently undermined our faith in the privacy-protecting power of anonymization, the name for techniques for protecting the privacy of individuals in large databases by deleting information like names and social security numbers. These scientists have demonstrated they can often 'reidentify' or 'deanonymize' individuals hidden in anonymized data with astonishing ease. By understanding this research, we will realize we have made a mistake, labored beneath a fundamental misunderstanding, which has assured us much less privacy than we have assumed. This mistake pervades nearly every information privacy law, regulation, and debate, yet regulators and legal scholars have paid it scant attention. We must respond to the surprising failure of anonymization, and this Article provides the tools to do so.
Laurence D. Robinson, Nicholas P. Jewell, Some Surprising Results about Covariate Adjustment in Logistic Regression Models, International Statistical Review / Revue Internationale de Statistique, Vol. 59, No. 2 (Aug., 1991), pp. 227-240
Primates demonstrate unparalleled ability at rapidly orienting towards important events in complex dynamic environments. During rapid guidance of attention and gaze towards potential objects of interest or threats, often there is no time for detailed visual analysis. Thus, heuristic computations are necessary to locate the most interesting events in quasi real-time. We present a new theory of sensory surprise, which provides a principled and computable shortcut to important information. We develop a model that computes instantaneous low-level surprise at every location in video streams. The algorithm significantly correlates with eye movements of two humans watching complex video clips, including television programs (17,936 frames, 2,152 saccadic gaze shifts). The system allows more sophisticated and time-consuming image analysis to be efficiently focused onto the most surprising subsets of the incoming data.
Confront and Conceal: Obama's Secret Wars and Surprising Use of American Power David E Sanger Crown Publishers, 2012 David Sanger begins his book with an anecdote that is quite telling. One midsumm...
Among the large variety of micro-organisms capable of fermentative hydrogen production, strict anaerobes such as members of the genus Clostridium are the most widely studied. They can produce hydrogen by a reversible reduction of protons accumulated during fermentation to dihydrogen, a reaction which is catalysed by hydrogenases. Sequenced genomes provide completely new insights into the diversity of clostridial hydrogenases. Building on previous reports, we found that [FeFe] hydrogenases are not a homogeneous group of enzymes, but exist in multiple forms with different modular structures and are especially abundant in members of the genus Clostridium. This unusual diversity seems to support the central role of hydrogenases in cell metabolism. In particular, the presence of multiple putative operons encoding multisubunit [FeFe] hydrogenases highlights the fact that hydrogen metabolism is very complex in this genus. In contrast with [FeFe] hydrogenases, their [NiFe] hydrogenase counterparts, widely represented in other bacteria and archaea, are found in only a few clostridial species. Surprisingly, a heteromultimeric Ech hydrogenase, known to be an energy-converting [NiFe] hydrogenase and previously described only in methanogenic archaea and some sulfur-reducing bacteria, was found to be encoded by the genomes of four cellulolytic strains: Clostridum cellulolyticum, Clostridum papyrosolvens, Clostridum thermocellum and Clostridum phytofermentans.
暂无摘要(点击查看原文获取完整内容)
暂无摘要(点击查看原文获取完整内容)
Despite the rapid rise in mothers' labor force participation, mothers' time with children has tended to be quite stable over time. In the past, nonemployed mothers' time with children was reduced by the demands of unpaid family work and domestic chores and by the use of mother substitutes for childcare, especially in large families. Today employed mothers seek ways to maximize time with children: They remain quite likely to work part-time or to exit from the labor force for some years when their children are young; they also differ from nonemployed mothers in other uses of time (housework, volunteer work, leisure). In addition, changes in children's lives (e.g., smaller families, the increase in preschool enrollment, the extended years of financial dependence on parents as more attend college) are altering the time and money investments that children require from parents. Within marriage, fathers are spending more time with their children than in the past, perhaps increasing the total time children spend with parents even as mothers work more hours away from home.
A surprising reward omission (SRO) occurs when an appetitive reinforcer is not presented (or it is reduced in magnitude or quality) even though there are signals for its impending presentation. Evidence supporting the hypothesis that SROs produce an aversive emotional reaction with physiological and behavioral consequences is reviewed. SROs are followed by pituitary–adrenal activation; changes in immune function; odor emissions in rodents; distress vocalizations in rodents and primates; and increases in locomotion, aggressive behavior, drinking, and eating. SROs can support the acquisition of new escape responses and invigorate previously acquired responses. The review identifies common aspects of these phenomena and areas in which more research is needed.
暂无摘要(点击查看原文获取完整内容)
The surprising or unexpected omission of an appetitive reinforcer has at least two effects: An allocentric effect according to which the organism updates knowledge about the environment, and an egocentric effect that allows the organism to learn about its own emotional reaction to the change. This egocentric effect (traditionally called frustration) is correlated to activation of the hypothalamic-pituitary-adrenal axis, can be modulated by treatment with anxiolytics, and is expressed in terms of behavioral changes that have an emotional component (e.g., agonistic behavior). It is hypothesized that all vertebrates share the mechanisms underlying the allocentric effect, but only mammals possess the mechanisms underlying the egocentric effect. It is further argued that frustrative mechanisms evolved in early mammals from those underlying fear conditioning.
暂无摘要(点击查看原文获取完整内容)
暂无摘要(点击查看原文获取完整内容)
暂无摘要(点击查看原文获取完整内容)
Shijie Wu, Mark Dredze. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.