Artificial Intelligence (AI) agents capable of autonomous learning and independent decision-making hold great promise for addressing complex challenges across various critical infrastructure domains, including transportation, energy systems, and manufacturing. However, the surge in the design and deployment of AI systems, driven by various stakeholders with distinct and unaligned objectives, introduces a crucial challenge: How can uncoordinated AI systems coexist and evolve harmoniously in shared environments without creating chaos or compromising safety? To address this, we advocate for a fundamental rethinking of existing multi-agent frameworks, such as multi-agent systems and game theory, which are largely limited to predefined rules and static objective structures. We posit that AI agents should be empowered to adjust their objectives dynamically, make compromises, form coalitions, and safely compete or cooperate through evolving relationships and social feedback. Through two case studies in critical infrastructure applications, we call for a shift toward the emergent, self-organizing, and context-aware nature of these multi-agentic AI systems.
A press release from the National Institute of Standards and Technology (NIST)could potentially impede progress toward improving the analysis of forensic evidence and the presentation of forensic analysis results in courts in the United States and around the world. "NIST experts urge caution in use of courtroom evidence presentation method" was released on October 12, 2017, and was picked up by the phys.org news service. It argues that, except in exceptional cases, the results of forensic analyses should not be reported as "likelihood ratios". The press release, and the journal article by NIST researchers Steven P. Lund & Harri Iyer on which it is based, identifies some legitimate points of concern, but makes a strawman argument and reaches an unjustified conclusion that throws the baby out with the bathwater.
The simplest version of the daemon paradigm suggests the modulated 2-6-keV range events in DAMA/NaI and DAMA/LIBRA detectors are caused by the iodine ions knocked out elastically by the electrically neutral c-daemons moving with V = 30-50 km/s (c-daemon is a complex of negative daemon located in a remainder of formerly captured nucleus where the daemon decomposes nucleons one by one with ~10^-6 s mean interval). Furthermore, after the 2-6 keV event occurred, in subsequent ~10^-6 s, the c-daemon (which becomes negative during this time) recaptures new nucleus with resulting scintillations in ~10 MeV range! The last possibility was so far overlooked in the experiments as it did not stem from WIMP hypotheses. A modification of the NaI(Tl) experiments is suggested for revealing the effect described. Independently of the outcome, any obtained result will be important for refining the daemon paradigm further on.
As cellular networks are turning into a platform for ubiquitous data access, cellular operators are facing a severe data capacity crisis due to the exponential growth of traffic generated by mobile users. In this work, we investigate the benefits of sharing infrastructure and spectrum among two cellular operators. Specifically, we provide a multi-cell analytical model using stochastic geometry to identify the performance gain under different sharing strategies, which gives tractable and accurate results. To validate the performance using a realistic setting, we conduct extensive simulations for a multi-cell OFDMA system using real base station locations. Both analytical and simulation results show that even a simple cooperation strategy between two similar operators, where they share spectrum and base stations, roughly quadruples capacity as compared to the capacity of a single operator. This is equivalent to doubling the capacity per customer, providing a strong incentive for operators to cooperate, if not actually merge.
Regulating artificial intelligence (AI) has become necessary in light of its deployment in high-risk scenarios. This paper explores the proposal to extend legal personhood to AI and robots, which had not yet been examined through the lens of the general public. We present two studies (N = 3,559) to obtain people's views of electronic legal personhood vis-à-vis existing liability models. Our study reveals people's desire to punish automated agents even though these entities are not recognized any mental state. Furthermore, people did not believe automated agents' punishment would fulfill deterrence nor retribution and were unwilling to grant them legal punishment preconditions, namely physical independence and assets. Collectively, these findings suggest a conflict between the desire to punish automated agents and its perceived impracticability. We conclude by discussing how future design and legal decisions may influence how the public reacts to automated agents' wrongdoings.
https://littletech。org/https://static。com/4a/bf/9c4021d8404386b0a311dcccf0
The heterogeneity in the organization of software engineering (SE) research historically exists, i.e., funded research model and hands-on model, which makes software engineering become a thriving interdisciplinary field in the last 50 years. However, the funded research model is becoming dominant in SE research recently, indicating such heterogeneity has been seriously and systematically threatened. In this essay, we first explain why the heterogeneity is needed in the organization of SE research, then present the current trend of SE research nowadays, as well as the consequences and potential futures. The choice is at our hands, and we urge our community to seriously consider maintaining the heterogeneity in the organization of software engineering research.
This paper offers a call to action. We urge our colleagues in the research community to play a greater role in the articulation of our findings to the public. To illustrate the stakes we present a case study on the initial stages of an LLM-based machine translation application's deployment in a real-world context: a text-2-911 system advertising capabilities in 55 languages for use in emergencies in which it may be difficult to call operators directly. We identify a number of common misconceptions about technologies such as these, concluding with a set of concrete recommendations and best practices for stakeholders at every stage of the development and deployment pipeline. While the advancement of scientific research often lies in solving the "hard" problems, we argue it is often the "easy" ones -- problems for which the latest technology is often unnecessary -- that are most overlooked.
Robustness verification of neural networks, referring to formally proving that neural networks satisfy robustness properties, is of crucial importance in safety-critical applications, where model failures can result in loss of human life or million-dollar damages. However, the dependability of verification results may be questioned due to sources of randomness in machine learning, and although this has been widely investigated for accuracy, its impact on robustness verification remains unknown. In this paper, we demonstrate a concerning result: Models that differ only in random seeds during training exhibit extreme variance in their certified robustness, with a standard deviation that is statistically larger than the marginal robustness improvements reported in recent machine learning papers. In addition, we also show that certified robustness generalization to unseen data varies significantly across datasets, falling short of the dependability expectations for safety-critical tasks. Our findings are major concerns because: (i) machine learning results in certified robustness are likely unconvincing due to extreme variance in certified robustness, and (ii) a ``lucky'' model seed in
Bursts from the very early universe may lead to a detectable signal via the production of positrons, whose annihilation gives an observable X-ray signal. Using the absorption parameters for the annihilation photons of 511 keV, it is found that observable photons would originate at a red-shift around $z\approx$ 200-300, resulting in soft X-rays of energy $\sim$ 2-3 keV at present. Positrons are expected to be absent at these times or red-shifts in the standard picture of the early universe. Detection of the X-rays would thus provide dramatic support for the hypothesis of the bursts, explosive events at very early times. We urge the search for such a signal.
"Math is not a spectator sport." "Lecturing is educational malpractice." Slogans like these rally some mathematicians to teach classes that feature "active learning", where lecturing is eschewed for student participation. Yet as much as I believe that students must do math to learn math, I also find blanket statements to be more about bandwagons than considered reflection on teaching. In this column, published in the Fall 2021 AWM Newsletter, I urge us to think through the math we offer students and how we set up students to learn. Although I draw primarily from my experiences teaching proofs in abstract algebra and real analysis, the scenarios extend to other topics in first year undergraduate education and beyond.
Recent experiments on bulk and thin film bilayer nickelate high-$T_c$ superconductors urge for clarification of their pairing mechanism. Debates exist on whether the hybridization or the Hund's coupling between the nickel $d_{x^2-y^2}$ and $d_{z^2}$ orbitals plays a primary role in driving the superconductivity. Here, we study the Hund scenario and make comparisons with the hybridization scenario using the same dynamic Schwinger boson approach. Our calculations reveal several key features of the Hund-driven superconductivity, including an isotropic $s$-wave gap, a lower maximum $T_c$, and Fermi liquid normal states, that differ from the hybridization-driven mechanism. We attribute these differences to their distinct low-energy dynamics. Comparison with recent experiments suggests that the Hund scenario alone is not enough to explain the bilayer nickelate superconductivity in both bulk and thin films.
Improvements in large language models have led to increasing optimism that they can serve as reliable evaluators of natural language generation outputs. In this paper, we challenge this optimism by thoroughly re-evaluating five state-of-the-art factuality metrics on a collection of 11 datasets for summarization, retrieval-augmented generation, and question answering. We find that these evaluators are inconsistent with each other and often misestimate system-level performance, both of which can lead to a variety of pitfalls. We further show that these metrics exhibit biases against highly paraphrased outputs and outputs that draw upon faraway parts of the source documents. We urge users of these factuality metrics to proceed with caution and manually validate the reliability of these metrics in their domain of interest before proceeding.
The histopathology analysis is of great significance for the diagnosis and prognosis of cancers, however, it has great challenges due to the enormous heterogeneity of gigapixel whole slide images (WSIs) and the intricate representation of pathological features. However, recent methods have not adequately exploited geometrical representation in WSIs which is significant in disease diagnosis. Therefore, we proposed a novel weakly-supervised framework, Geometry-Aware Transformer (GOAT), in which we urge the model to pay attention to the geometric characteristics within the tumor microenvironment which often serve as potent indicators. In addition, a context-aware attention mechanism is designed to extract and enhance the morphological features within WSIs.
In a recent Comment (arXiv:2411.10522, Nat Rev Phys 7, 2 (2025)), fifteen prominent leaders in the field of condensed matter physics declare that hydride superconductivity is real and urge funding agencies to continue to support the field. I question the validity and constructiveness of their argument.
Recent progress in research on Deep Graph Networks (DGNs) has led to a maturation of the domain of learning on graphs. Despite the growth of this research field, there are still important challenges that are yet unsolved. Specifically, there is an urge of making DGNs suitable for predictive tasks on realworld systems of interconnected entities, which evolve over time. With the aim of fostering research in the domain of dynamic graphs, at first, we survey recent advantages in learning both temporal and spatial information, providing a comprehensive overview of the current state-of-the-art in the domain of representation learning for dynamic graphs. Secondly, we conduct a fair performance comparison among the most popular proposed approaches on node and edge-level tasks, leveraging rigorous model selection and assessment for all the methods, thus establishing a sound baseline for evaluating new architectures and approaches
I stress the importance of retaining a healthy classical limit while we search for an ultraviolet completion to quantum gravity. A key problem with negative-norm quantizations of higher derivative Lagrangians is that their classical limits do not correspond to real-valued metrics evolving in a real-valued spacetime. I also demonstrate that no completion based on the flat spacetime background S-matrix can suffice by providing an explicit example of a theory with unit S-matrix which still shows interesting changes in single-particle kinematics and in the evolution of its background. I discuss the implications of these considerations for the program of Asymptotic Safety. Finally, I urge that some attention be given to the possibility that quantum general relativity might make sense if only we could go beyond conventional perturbation theory.
This paper emphasizes the importance of reporting experiment details in subjective evaluations and demonstrates how such details can significantly impact evaluation results in the field of speech synthesis. Through an analysis of 80 papers presented at INTERSPEECH 2022, we find a lack of thorough reporting on critical details such as evaluator recruitment and filtering, instructions and payments, and the geographic and linguistic backgrounds of evaluators. To illustrate the effect of these details on evaluation outcomes, we conducted mean opinion score (MOS) tests on three well-known TTS systems under different evaluation settings and we obtain at least three distinct rankings of TTS models. We urge the community to report experiment details in subjective evaluations to improve the reliability and interpretability of experimental results.
We summarize evidence that multiple supernovae exploded within 100 pc of Earth in the past few Myr. These events had dramatic effects on the heliosphere, compressing it to within ~20 au. We advocate for cross-disciplinary research of nearby supernovae, including on interstellar dust and cosmic rays. We urge for support of theory work, direct exploration, and study of extrasolar astrospheres.
The Universal Basic Computing Power (UBCP) initiative ensures global, free access to a set amount of computing power specifically for AI research and development (R&D). This initiative comprises three key elements. First, UBCP must be cost free, with its usage limited to AI R&D and minimal additional conditions. Second, UBCP should continually incorporate the state of the art AI advancements, including efficiently distilled, compressed, and deployed training data, foundational models, benchmarks, and governance tools. Lastly, it's essential for UBCP to be universally accessible, ensuring convenience for all users. We urge major stakeholders in AI development large platforms, open source contributors, and policymakers to prioritize the UBCP initiative.