Open-source reinforcement learning (RL) environments have played a crucial role in driving progress in the development of AI algorithms. In modern RL research, there is a need for simulated environments that are performant, scalable, and modular to enable their utilization in a wider range of potential real-world applications. Therefore, we present Jumanji, a suite of diverse RL environments specifically designed to be fast, flexible, and scalable. Jumanji provides a suite of environments focusing on combinatorial problems frequently encountered in industry, as well as challenging general decision-making tasks. By leveraging the efficiency of JAX and hardware accelerators like GPUs and TPUs, Jumanji enables rapid iteration of research ideas and large-scale experimentation, ultimately empowering more capable agents. Unlike existing RL environment suites, Jumanji is highly customizable, allowing users to tailor the initial state distribution and problem complexity to their needs. Furthermore, we provide actor-critic baselines for each environment, accompanied by preliminary findings on scaling and generalization scenarios. Jumanji aims to set a new standard for speed, adaptability, a
World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset. However, they are vulnerable to model exploitation where data coverage is thin. Prior work addresses this either by collecting more expert demonstrations, which is often expensive, unsafe, or unavailable, or by conservative algorithms that avoid uncertain regions, which limits generalization. We propose instead to repair exploitation directly using human preferences over imagined rollouts, leveraging the strong intuitive physics that allows humans to easily spot egregious dynamics hallucinations. We formalize this as Dynamics Learning from Human Feedback (DLHF), a Bradley-Terry preference loss over trajectory log-likelihoods under a learned dynamics model. Unfortunately, naive DLHF is sample inefficient, so we introduce RENEW, which uses epistemic uncertainty to focus finetuning where the model is most exploitable. We evaluate on several Jumanji and classic control environments and find that while naive DLHF requires an outsize preference budget, RENEW makes the framework practical by improving sample efficiency, limiting catastrophic forg
Proximal policy optimization (PPO) is a widely-used algorithm for on-policy reinforcement learning. This work offers an alternative perspective of PPO, in which it is decomposed into the inner-loop estimation of update vectors, and the outer-loop application of updates using gradient ascent with unity learning rate. Using this insight we propose outer proximal policy optimization (outer-PPO); a framework wherein these update vectors are applied using an arbitrary gradient-based optimizer. The decoupling of update estimation and update application enabled by outer-PPO highlights several implicit design choices in PPO that we challenge through empirical investigation. In particular we consider non-unity learning rates and momentum applied to the outer loop, and a momentum-bias applied to the inner estimation loop. Methods are evaluated against an aggressively tuned PPO baseline on Brax, Jumanji and MinAtar environments; non-unity learning rates and momentum both achieve statistically significant improvement on Brax and Jumanji, given the same hyperparameter tuning budget.
Scientists have created a programmable optical chip that can slow light on demand, giving engineers far greater control over how optical signals propagate through a circuit。 The technology could provide the delays, synchronization, and buffering functions needed to make light-based computing more practical。 A single chip could eventually perform se
Dying Sun-like stars may not fade away quietly。 As blobs of gas erupt unevenly from their swollen surfaces, each burst gives the star a tiny push in the opposite direction。 Thousands of these random kicks could eventually break apart distant stellar pairs or, in rare cases, drive two stars into a violent collision
10 days passed from OpenAI models exploiting JFrog Artifactory 0-day to release of a patch
Google's new API relies on parents to set age ranges in Family Link
A new AI-powered blood test could give people a remarkably early warning of serious heart and circulation problems。 Developed by researchers at the University of Hong Kong, CardiOmicScore analyzes thousands of proteins and metabolites to estimate the risk of six major cardiovascular diseases, including heart attack, stroke, heart failure, and atria
Scientists say new technologies have reopened the debate over whether Mars could someday be terraformed, turning a once impossible idea into a serious research topic。 Before anyone tries to reshape the Red Planet, though, researchers say we must understand the risks, including what might be lost if Mars already harbors its own forms of life
Young Jupiter’s powerful magnetic field may have created a safe zone where several large moons could survive。 Saturn lacked this protection, possibly explaining why Titan stands almost alone among its largest moons
Europa’s hidden ocean has made the icy moon one of the solar system’s most promising places to search for habitable conditions。 Scientists have hoped that water from this deep ocean might rise through cracks and form shallow reservoirs that future spacecraft could more easily study。 New simulations suggest that journey is unlikely because turbulent
A new theoretical study offers a possible explanation for how the Universe can grow more complex without violating the second law of thermodynamics。 Using a quantum gravity framework called Gravity from Entropy, mathematician Ginestra Bianconi found that the Universe’s total entropy may rise as space expands, even while entropy within each unit of
Ars Pro subscribers get customizable layouts along with no ads and no trackers
Researchers have created self-destructing living plastic that uses engineered bacteria to completely break itself down when activated。 The material degrades in just six days without creating microplastics, offering a potential new solution for single-use plastic waste
NASA’s Swift Observatory observed a supermassive black hole ripping apart a star more than 30,000 light-years from the center of a distant galaxy。 The extraordinary flare briefly outshone its entire host galaxy in ultraviolet light and revealed a black hole about a million times the Sun’s mass
Scientists have uncovered new evidence that Venus may still be tearing itself apart from within。 Advanced 3D simulations indicate that some of the planet's giant rift valleys formed relatively recently and could still be expanding。 The results suggest Venus has a far more active interior than researchers once believed, challenging the long-held vie