共找到 20 条结果
We ask: Can focusing on likely classes of a single, in-domain sample improve model predictions? Prior work argued ``no''. We put forward a novel rationale in favor of ``yes'': Sharedness of features among classes indicates their reliability for a single sample. We aim for an affirmative answer without using hand-engineered augmentations or auxiliary tasks. We propose two novel test-time fine-tuning methods to improve uncertain model predictions. Instead of greedily selecting the most likely class, we introduce an additional step, \emph{focus on the likely classes}, to refine predictions. By applying a single gradient descent step with a large learning rate, we refine predictions when an initial forward pass indicates high uncertainty. The experimental evaluation demonstrates accuracy gains for one of our methods on average, which emphasizes shared features among likely classes. The gains are confirmed across diverse text and image domain models.
The core composition of ultramassive white dwarfs remains an open question in stellar evolution. The carbon content of white dwarf cores is critical to their role as progenitors of Type Ia supernovae. However, because the stellar photosphere only extends to the outermost layer of the star, observational probes of core compositions are limited. Here we present gravitational redshift measurements of an ultramassive white dwarf, SDSS J060851.44-005950.3, which indicate the likely presence of an oxygen-neon core. We measure the mass ($1.226_{-0.025}^{+0.024} M_\odot$) and radius ($0.491_{-0.009}^{+0.009}~R_\oplus$) of the white dwarf using gravitational redshifts from high-resolution UVES and MagE spectra paired with independent constraints from photometry. By comparing to state-of-the-art mass-radius relations for ultramassive white dwarfs, we find preference for a oxygen-neon core over a carbon-oxygen core, with a Bayes factor of $2.7$. This is a white dwarf which is likely structurally incapable of producing a Type Ia supernova, according to current understanding of supernova physics. This object provides evidence that white dwarfs which pass through the Q-branch without experiencin
A human decision-maker benefits the most from an AI assistant that corrects for their biases. For problems such as generating interpretation of a radiology report given findings, a system predicting only highly likely outcomes may be less useful, where such outcomes are already obvious to the user. To alleviate biases in human decision-making, it is worth considering a broad differential diagnosis, going beyond the most likely options. We introduce a new task, "less likely brainstorming," that asks a model to generate outputs that humans think are relevant but less likely to happen. We explore the task in two settings: a brain MRI interpretation generation setting and an everyday commonsense reasoning setting. We found that a baseline approach of training with less likely hypotheses as targets generates outputs that humans evaluate as either likely or irrelevant nearly half of the time; standard MLE training is not effective. To tackle this problem, we propose a controlled text generation method that uses a novel contrastive learning strategy to encourage models to differentiate between generating likely and less likely outputs according to humans. We compare our method with severa
An important goal of precision medicine is to personalize medical treatment by identifying individuals who are most likely to benefit from a specific treatment. The Likely Responder (LR) framework, which identifies a subpopulation where treatment response is expected to exceed a certain clinical threshold, plays a role in this effort. However, the LR framework, and more generally, data-driven subgroup analyses, often fail to account for uncertainty in the estimation of model-based data-driven subgrouping. We propose a simple two-stage approach that integrates subgroup identification with subsequent subgroup-specific inference on treatment effects. We incorporate model estimation uncertainty from the first stage into subgroup-specific treatment effect estimation in the second stage, by utilizing Bayesian posterior distributions from the first stage. We evaluate our method through simulations, demonstrating that the proposed Bayesian two-stage model produces better calibrated confidence intervals than naïve approaches. We apply our method to an international COVID-19 treatment trial, which shows substantial variation in treatment effects across data-driven subgroups.
Quantum digital signatures (QDSs) can provide information-theoretic security of messages against forgery and repudiation. Compared with previous QDS protocols that focus on signing one-bit messages, hash function-based QDS protocols can save quantum resources and are able to sign messages of arbitrary length. Using the idea of likely bit strings, we propose an efficient QDS protocol with hash functions over long distances. Our method of likely bit strings can be applied to any quantum key distribution-based QDS protocol to significantly improve the signature rate and dramatically increase the secure signature distance of QDS protocols. In order to save computing resources, we propose an improved method where Alice participates in the verification process of Bob and Charlie. This eliminates the computational complexity relating to the huge number of all likely strings. We demonstrate the advantages of our method and our improved method with the example of sending-or-not-sending QDS. Under typical parameters, both our method and our improved method can improve the signature rate by more than 100 times and increase the signature distance by about 150 km compared with hash function-bas
Self-consistency (Wang et al., 2023) suggests that the most consistent answer obtained through large language models (LLMs) is more likely to be correct. In this paper, we challenge this argument and propose a nuanced correction. Our observations indicate that consistent answers derived through more computation i.e. longer reasoning texts, rather than simply the most consistent answer across all outputs, are more likely to be correct. This is predominantly because we demonstrate that LLMs can autonomously produce chain-of-thought (CoT) style reasoning with no custom prompts merely while generating longer responses, which lead to consistent predictions that are more accurate. In the zero-shot setting, by sampling Mixtral-8x7B model multiple times and considering longer responses, we achieve 86% of its self-consistency performance obtained through zero-shot CoT prompting on the GSM8K and MultiArith datasets. Finally, we demonstrate that the probability of LLMs generating a longer response is quite low, highlighting the need for decoding strategies conditioned on output length.
The principle of rewarding a crowd for surprisingly common answers has been used in the literature for designing a number of truthful information elicitation mechanisms. A related method has also been proposed in the literature for better aggregation of crowd wisdom. Drawing a comparison between crowd based collective intelligence systems and large language models, we define the notion of 'surprisingly likely' textual response of a large language model. This notion is inspired by the surprisingly common principle, but tailored for text in a language model. Using benchmarks such as TruthfulQA and openly available LLMs: GPT-2 and LLaMA-2, we show that the surprisingly likely textual responses of large language models are more accurate in many cases compared to standard baselines. For example, we observe up to 24 percentage points aggregate improvement on TruthfulQA and up to 70 percentage points improvement on individual categories of questions in this benchmark. We also provide further analysis of the results, including the cases when surprisingly likely responses are less or not more accurate.
Standard regularized training procedures correspond to maximizing a posterior distribution over parameters, known as maximum a posteriori (MAP) estimation. However, model parameters are of interest only insomuch as they combine with the functional form of a model to provide a function that can make good predictions. Moreover, the most likely parameters under the parameter posterior do not generally correspond to the most likely function induced by the parameter posterior. In fact, we can re-parametrize a model such that any setting of parameters can maximize the parameter posterior. As an alternative, we investigate the benefits and drawbacks of directly estimating the most likely function implied by the model and the data. We show that this procedure leads to pathological solutions when using neural networks and prove conditions under which the procedure is well-behaved, as well as a scalable approximation. Under these conditions, we find that function-space MAP estimation can lead to flatter minima, better generalization, and improved robustness to overfitting.
In this paper, we study the Onsager-Machlup function and its relationship to the Freidlin-Wentzell function for measures equivalent to arbitrary infinite dimensional Gaussian measures. The Onsager-Machlup function can serve as a density on infinite dimensional spaces, where a uniform measure does not exist, and has been seen as the Lagrangian for the ``most likely element". The Freidlin-Wentzell rate function is the large deviations rate function for small-noise limits and has also been identified as a Lagrangian for the ``most likely element". This leads to a conundrum - what is the relationship between these two functions? We show both pointwise and $Γ$-convergence (which is essentially the convergence of minimizers) of the Onsager-Machlup function under the small-noise limit to the Freidlin-Wentzell function - and give an expression for both. That is, we show that the small-noise limit of the most likely element is the most likely element in the small noise limit for infinite dimensional measures that are equivalent to a Gaussian. Examples of measures include the law of solutions to path-dependent stochastic differential equations and the law of an infinite system of random alge
Many complex real world phenomena exhibit abrupt, intermittent or jumping behaviors, which are more suitable to be described by stochastic differential equations under non-Gaussian Lévy noise. Among these complex phenomena, the most likely transition paths between metastable states are important since these rare events may have a high impact in certain scenarios. Based on the large deviation principle, the most likely transition path could be treated as the minimizer of the rate function upon paths that connect two points. One of the challenges to calculate the most likely transition path for stochastic dynamical systems under non-Gaussian Lévy noise is that the associated rate function can not be explicitly expressed by paths. For this reason, we formulate an optimal control problem to obtain the optimal state as the most likely transition path. We then develop a neural network method to solve this issue. Several experiments are investigated for both Gaussian and non-Gaussian cases.
A program invariant is a property that holds for every execution of the program. Recent work suggest to infer likely-only invariants, via dynamic analysis. A likely invariant is a property that holds for some executions but is not guaranteed to hold for all executions. In this paper, we present work in progress addressing the challenging problem of automatically verifying that likely invariants are actual invariants. We propose a constraint-based reasoning approach that is able, unlike other approaches, to both prove or disprove likely invariants. In the latter case, our approach provides counter-examples. We illustrate the approach on a motivating example where automatically generated likely invariants are verified.
We investigate how people perceive ChatGPT, and, in particular, how they assign human-like attributes such as gender to the chatbot. Across five pre-registered studies (N = 1,552), we find that people are more likely to perceive ChatGPT to be male than female. Specifically, people perceive male gender identity (1) following demonstrations of ChatGPT's core abilities (e.g., providing information or summarizing text), (2) in the absence of such demonstrations, and (3) across different methods of eliciting perceived gender (using various scales and asking to name ChatGPT). Moreover, we find that this seemingly default perception of ChatGPT as male can reverse when ChatGPT's feminine-coded abilities are highlighted (e.g., providing emotional support for a user).
This paper is a submission to the Open Philanthropy AI Worldviews Contest. In it, we estimate the likelihood of transformative artificial general intelligence (AGI) by 2043 and find it to be <1%. Specifically, we argue: The bar is high: AGI as defined by the contest - something like AI that can perform nearly all valuable tasks at human cost or less - which we will call transformative AGI is a much higher bar than merely massive progress in AI, or even the unambiguous attainment of expensive superhuman AGI or cheap but uneven AGI. Many steps are needed: The probability of transformative AGI by 2043 can be decomposed as the joint probability of a number of necessary steps, which we group into categories of software, hardware, and sociopolitical factors. No step is guaranteed: For each step, we estimate a probability of success by 2043, conditional on prior steps being achieved. Many steps are quite constrained by the short timeline, and our estimates range from 16% to 95%. Therefore, the odds are low: Multiplying the cascading conditional probabilities together, we estimate that transformative AGI by 2043 is 0.4% likely. Reaching >10% seems to require probabilities that feel u
We prove a general likely intersections theorem, a counterpart to the Zilber-Pink conjectures, under the assumption that the Ax-Schanuel property and some mild additional conditions are known to hold for a given category of complex quotient spaces definable in some fixed o-minimal expansion of the ordered field of real numbers. For an instance of our general result, consider the case of subvarieties of Shimura varieties. Let $S$ be a Shimura variety. Let $π:D \to Γ\backslash D = S$ realize $S$ as a quotient of $D$, a homogeneous space for the action of a real algebraic group $G$, by the action of $Γ< G$, an arithmetic subgroup. Let $S' \subseteq S$ be a special subvariety of $S$ realized as $π(D')$ for $D' \subseteq D$ a homogeneous space for an algebraic subgroup of $G$. Let $X \subseteq S$ be an irreducible subvariety of $S$ not contained in any proper weakly special subvariety of $S$. Assume that the intersection of $X$ with $S'$ is persistently likely meaning that whenever $ζ:S_1 \to S$ and $ξ:S_1 \to S_2$ are maps of Shimura varieties (meaning regular maps of varieties induced by maps of the corresponding Shimura data) with $ζ$ finite, $\dim ξζ^{-1} X + \dim ξζ^{-1} S' \geq
We present first results from follow-up of targets in the northern hemisphere Beta Pictoris and AB Doradus moving group candidate list of Schlieder, Lepine, and Simon (2012). We obtained high-resolution, near-infrared spectra of 27 candidate members to measure their radial velocities and confirm consistent group kinematics. We identify 15 candidates with consistent predicted and measured radial velocities, perform analyses of their 6-dimensional (U,V,W,X,Y,Z) Galactic kinematics, and compare to known group member distributions. Based on these analyses, we propose that 7 Beta Pic and 8 AB Dor candidates are likely new group members. Four of the likely new Beta Pic stars are binaries; one a double lined spectroscopic system. Three of the proposed AB Dor stars are binaries. Counting all binary components, we propose 22 likely members of these young, moving groups. The majority of the proposed members are M2 to M5 dwarfs, the earliest being of type K2. We also present preliminary parameters for the two new spectroscopic binaries identified in the data, the proposed Beta Pic member and a rejected Beta Pic candidate. Our candidate selection and follow-up has thus far identified more than
Error bounds based on worst likely assignments use permutation tests to validate classifiers. Worst likely assignments can produce effective bounds even for data sets with 100 or fewer training examples. This paper introduces a statistic for use in the permutation tests of worst likely assignments that improves error bounds, especially for accurate classifiers, which are typically the classifiers of interest.
We present new photometric observations of Supernova (SN) 2003ie starting one month before discovery, obtained serendipitously while observing its host galaxy. With only a weak upper limit derived on the mass of its progenitor (<25 M_sun) from pre-explosion studies, this event could be a potential exception to the "red supergiant (RSG) problem" (the lack of high mass RSGs exploding as Type IIP supernovae). However, this is true only if SN2003ie was a Type IP event, something which has never been determined. Using recently derived core collapse SN light curve templates, as well as by comparison to other known SNe, we find that SN2003ie was indeed a likely Type IIP event. However, it is found to be a member of the faint Type IIP class. Previous members of this class have been shown to arise from relatively low mass progenitors (<12 M_sun). It therefore seems unlikely that this SN had a massive RSG progenitor. The use of core collapse SN light curve templates is shown to be helpful in classifying SNe with sparse coverage. These templates are likely to become more robust as large homogeneous samples of core collapse events are collected.
If a macroscopic (random) classical system is put into a random state in phase space, it will of course the most likely have an almost maximal entropy according to second law of thermodynamics. We will show, however, the following theorem: If it is enforced to be periodic with a given period $T$ in advance, the distribution of the entropy for the otherwise random state will be much more smoothed out, and the entropy could be very likely much smaller than the maximal one. Even quantum mechanically we can understand that such a lower than maximal entropy is likely. A corollary turns out to be that the entropy in such closed time-like loop worlds remain constant.
Finding the most likely path to a set of failure states is important to the analysis of safety-critical systems that operate over a sequence of time steps, such as aircraft collision avoidance systems and autonomous cars. In many applications such as autonomous driving, failures cannot be completely eliminated due to the complex stochastic environment in which the system operates. As a result, safety validation is not only concerned about whether a failure can occur, but also discovering which failures are most likely to occur. This article presents adaptive stress testing (AST), a framework for finding the most likely path to a failure event in simulation. We consider a general black box setting for partially observable and continuous-valued systems operating in an environment with stochastic disturbances. We formulate the problem as a Markov decision process and use reinforcement learning to optimize it. The approach is simulation-based and does not require internal knowledge of the system, making it suitable for black-box testing of large systems. We present formulations for fully observable and partially observable systems. In the latter case, we present a modified Monte Carlo
As galaxies form hierarchically, larger satellites may accrete alongside smaller companions in group infall events. This coordinated accretion is likely to have left signatures in the Milky Way's stellar halo at the present day. Our goal is to characterise the possible groups of companions that accompanied larger known accretion events of our Galaxy, and infer where their stellar material could be in physical and dynamical space at present day. We use the AURIGA simulation suite of Milky Way-like haloes to identify analogues to these large accretion events and their group infall companions, and we follow their evolution in time. We find that most of the material from larger accretion events is deposited on much more bound orbits than their companions. This implies a weak dynamical association between companions and debris, but it is strongest with the material lost first. As a result, the companions of the Milky Way's earliest building blocks are likely to contribute stars to the solar neighbourhood, whilst the companions of our last major merger are likely found in both the solar neighbourhood and the outer halo. More recently infallen groups of satellites, or those of a smaller m