Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these models possess many of the required component capabilities, but the conditions that coincide in collaboration, including time pressure, information asymmetry, and imperfect communication, are usually studied in isolation. We introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent can see and manipulate the bomb but does not have the defusal instructions; the other has the instructions but cannot see or manipulate the bomb. Neither agent can succeed alone: success requires effective and efficient communication. Unlike turn-based proxies, GPTNT requires agents to act asynchronously and communicate in real time. GPTNT is designed to separate collaboration from reliance on memorized solutions: the instruction manual, the partner, or both can be withheld to isolate what a model derives in the moment from what it already knows. We show that GPTNT poses a substantial challenge for s
We argue that Bonferroni correction is a better choice for online experimentation than it is commonly given credit for. The case rests on four considerations. First, it is the simplest broadly implementable FWER-controlling method that produces unconditional simultaneous confidence intervals for every metric. Second, in a well-specified decision framework, guardrail and quality metrics use intersection-union logic and cannot inflate the false positive rate, so the Bonferroni denominator is the number of success metrics only, not the total metric count. Third, it is uniquely tractable for pre-experiment sample size calculations. Fourth, we contextualise the power cost empirically. Drawing on a simulation study and an empirical analysis of 1,296 experiments run on Spotify's experimentation platform, Confidence, we show that the power loss relative to more sophisticated FWER methods depends on both how the correction family is specified and how many metrics are truly non-null. When guardrail metrics are incorrectly included in the family, Holm and Hommel are nearly indistinguishable from Bonferroni. When the family is correctly restricted to success metrics only, they gain roughly 4--
Equations are ubiquitous in most mathematical activities. Nevertheless, in this paper it is shown how to do standard mathematics without any equation at all. More than that, it is proven there is a foundational framework for standard mathematics where equations cannot be even written, in the sense they are not formulas. The proof of those claims is very simple, almost obvious. I use this framework to suggest a way to deal with certain notions of indiscernibility between `objects', with special emphasis on some aspects of quantum mechanics. Finally I compare this approach to quasi-set theory, an unnecessarily complicated formal work designed to deal with violation of Leibniz Principle of the Identity of Indiscernibles.
We reject unjustified criticism of our published article [2209.07992] by Gill and Lambare [arXiv:2211.02481, arXiv:2208.09930]. They completely misinterpret the content and conclusions of this article. They construct a counterfactual probabilistic model in which random variables representing outcomes of four experiments performed using incompatible experimental settings are jointly distributed. Thus, CHSH inequalities trivially hold for all finite samples generated by their model. Their model defines a probabilistic coupling for our model describing only the raw data from Bell tests. The existence of this coupling does not invalidate the derivation of the contextual probabilistic model describing the final data from Bell tests. Only these final data are used to test Bell inequalities. Inequalities cannot be derived because our model violates statistical independence. Our contextual model allows to explain in a local and causal way the violation of inequalities and the apparent violation of no-signaling reported in these experiments.
I employ an optimization-based inference methodology together with an Ising model, in an intentionally ineffectual manner, to get away with murdering an obstreperous scientific collaborator. The antics of this collaborator, hereafter "Conan O'Brien," were impeding the publication of an important manuscript. With my tenure date looming, I found myself desperate. Luckily, I study inference, a computational means to find a solution to a physical problem, based on available measurements (say, a dead body) and a dynamical model assumed to give rise to those measurements (a murderer). If the measurements are insufficient and/or the model is incomplete, one obtains multiple "degenerate" solutions to the problem. Degenerate solutions are all equally valid given the information available, and thus render meaningless the notion of one "correct" solution. Typically in scientific research, degeneracy is undesirable. Here I describe the opposite situation: a quest to create degenerate solutions in which to cloak myself. Or even better: to render measurements incompatible with a solution in which I am the murderer. Moreover, I show how one may sabotage an inference procedure to commit an untrace
Kupczynski (2023) claims that Gill and Lambare (2022a, 2022b) misrepresent several of his published papers. This paper shows that the latest version of his "contextuality by default" model of a Bell experiment places no constraints whatsoever on the statistics of observed results in Bell type experiments. It thereby effectively allows arbitrary non-locality, ie direct causal effects of local measurement settings on distant measurement outcomes.
Marian Kupczynski(MK)is the author of a controversial paper published (2020) in the journal Frontiers in Physics. The work is built around a mathematical claim by MK which is actually false, and MK's logical reasoning around his claim is also incorrect. The same claim was made by him in several other recent papers published in other journals. A proof that the claimed result is false is the main content of our present "Comment". It is purely a mathematical counter-example to a mathematical claim in a number of MK's papers.
We dedicate this to the life and work of Robin Hudson -- a mathematical physicist who developed the peerless quantum stochastic calculus, but who also inspired generations of researchers with both his intellect and wit.
Among all cybersecurity and privacy workers, the Data Protection Officer (DPO) stands between those auditing a company's compliance and those acting as management advisors. A person that must be somehow versed in legal, management, and cybersecurity technical skills. We describe how this role tackles socio-technical risks in everyday scenarios.
In a recent preprint Gill and Lambare, criticize our paper published in Frontiers in Physics. Their criticism is unfounded and misleading. They define a probabilistic coupling, in which BI-CHSH hold for all finite samples. It does not mean, that BI-CHSH hold in our model, in which four incompatible experiments are described by setting dependent random variables implemented on 4 disjoint dedicated probability spaces. A joint probability distribution of these random variables does not exist and may not be used to derive inequalities. Moreover, their probabilistic coupling is useless, for a subsequent contextual model, which we construct to describe final data from Bell tests and to explain, in a locally causal way, the reported violations of inequalities and apparent violations of no-signaling. Neither quantum probabilistic model of an ideal EPRB experiment nor local realistic and stochastic hidden variable models may explain reported non-signaling Therefore; it is obvious that our model extends the set of probability distributions of possible measurements allowed in the standard hidden variable models. Gill and Lambare seem not understand , the main message of our paper, that the vi
This is an extended essay review of Tanya and Jeffrey Bub's Totally Random: Why Nobody Understands Quantum Mechanics: A serious comic on entanglement. Princeton and Oxford: Princeton University Press (2018), ISBN: 9780691176956, 272 pp., 7x10 in., 254 b/w illus., £18.99 / $22.95 (paperback). We review the philosophical aspects of the book, provide suggestions for instructors on how to use the book in a class setting, and evaluate the authors' artistic choices in the context of comics theory.
In 1981, David Mermin described a cleverly simplified version of Bell's theorem. It pointed out in a straightforward way that interpreting entanglement from a local realist point of view can be problematic. I propose here an extended version of Mermin's device that can actually be given a simple local realist interpretation through a sample selection bias, and I argue that we still have no scientific reason to believe that the moon could possibly not be there when nobody looks.
Program synthesis from natural language (NL) is practical for humans and, once technically feasible, would significantly facilitate software development and revolutionize end-user programming. We present SAPS, an end-to-end neural network capable of mapping relatively complex, multi-sentence NL specifications to snippets of executable code. The proposed architecture relies exclusively on neural components, and is trained on abstract syntax trees, combined with a pretrained word embedding and a bi-directional multi-layer LSTM for processing of word sequences. The decoder features a doubly-recurrent LSTM, for which we propose novel signal propagation schemes and soft attention mechanism. When applied to a large dataset of problems proposed in a previous study, SAPS performs on par with or better than the method proposed there, producing correct programs in over 92% of cases. In contrast to other methods, it does not require post-processing of the resulting programs, and uses a fixed-dimensional latent representation as the only interface between the NL analyzer and the source code generator.
Informationally complete measurements are a dramatic discovery of quantum information science, and the symmetric IC measurements, known as SICs, are in many ways optimal among them. Close study of three of the "sporadic SICs" reveals an illuminating relation between different ways of quantifying the extent to which quantum theory deviates from classical expectations.
Various Bell inequalities are trivial algebraic properties satisfied by each line of particular data spreadsheets.It is surprising that their violation in some experiments, allows to speculate about the existence of nonlocal influences in Nature and to doubt the existence of the objective external physical reality. Such speculations are rooted in incorrect interpretations of quantum mechanics and in a failure of local realistic hidden variable models to reproduce quantum predictions for spin polarisation correlation experiments. These hidden variable models use counterfactual joint probability distributions of only pairwise measurable random variables to prove the inequalities. In real experiments Alice and Bob, using 4 incompatible pairs of experimental settings, estimate imperfect correlations between clicks, registered by their detectors. Clicks announce detection of photons and are coded by 1 or -1. Expectations of corresponding ,only pairwise measurable, random variables are estimated and compared with quantum predictions. These estimates violate significantly the inequalities. Since all these random variables cannot be jointly measured , a joint probability distribution of th
We study the distribution of envy in random matching markets under the Deferred Acceptance (DA) algorithm. Using tools from applied probability, we compute the expected number of proposing agents whom nobody envies and those who envy nobody. We obtain an exact finite-market expression for the former, based on a connection with the coupon collector problem, and asymptotic bounds for the latter. To put these quantities into perspective, we compare them to their counterparts under Random Serial Dictatorship (RSD): while RSD assigns a constant fraction of agents to their top choice, both DA and RSD leave exactly $H_n$ proposing agents unenvied in expectation. Our results show that these clearly unimprovable proposing agents constitute a vanishing fraction of the market.
Digital systems have become simultaneously more powerful and more wasteful. Features accumulate that nobody uses. Data is collected that nobody analyzes. AI is deployed at significant energy and water costs for gains that a simpler approach could have achieved. And through all of it, the people who depend on these systems quietly absorb the consequences in cognitive load, lost time, and eroded trust. This paper introduces GreenZ, a three-layer Sustainable UX Framework for complex digital systems. Its three layers are a Philosophy Layer built around ten published principles, an Operational Frameworks Layer comprising five applied systems, and a Tools and Canvases Layer of practical audit instruments and decision models. Two contributions sit at the framework's core: a Digital Waste Taxonomy classifying eight distinct waste types, and an AI Sufficiency Decision Model that asks whether AI should exist in a given flow before any question of how to implement it. GreenZ v1 is theoretically grounded but empirically unvalidated. A practitioner expert review study is underway at the time of submission. The paper presents the framework's architecture, its conceptual foundations, its position
As $n\to \infty$, do the algebras of $n\times n$ complex matrices look alike? Nobody knows, but let's play a game and see how far we get!
Nobody knows how language works, but many theories abound. Transformers are a class of neural networks that process language automatically with more success than alternatives, both those based on neural computations and those that rely on other (e.g. more symbolic) mechanisms. Here, I highlight direct connections between the transformer architecture and certain theoretical perspectives on language. The empirical success of transformers relative to alternative models provides circumstantial evidence that the linguistic approaches that transformers embody should be, at least, evaluated with greater scrutiny by the linguistics community and, at best, considered to be the currently best available theories.
Estimation of software reliability often poses a considerable challenge, particularly for critical softwares. Several methods of estimation of reliability of software are already available in the literature. But, so far almost nobody used the concept of size of a bug for estimating software reliability. In this article we make used of the bug size or the eventual bug size which helps us to determine reliability of software more precisely. The size-biased model developed here can also be used for similar fields like hydrocarbon exploration. The model has been validated through simulation and subsequently used for a critical space application software testing data. The estimated results match the actual observations to a large extent.