共找到 20 条结果
In this paper, we present a geometric form of the Hahn-Banach extension theorem for $L^{0}-$linear functions and prove that the geometric form is equivalent to the analytic form of the Hahn-Banach extension theorem. Further, we use the geometric form to give a new proof of a known basic strict separation theorem in random locally convex modules. Finally, using the basic strict separation theorem we establish the Goldstine-Weston theorem in random normed modules under the two kinds of topologies----the $(ε,λ)-$topology and the locally $L^{0}-$convex topology, and also provide a counterexample showing that the Goldstine-Weston theorem under the locally $L^{0}-$convex topology can only hold for random normed modules with the countable concatenation property.
Recent empirical evidence suggests that the Weston-Watkins support vector machine is among the best performing multiclass extensions of the binary SVM. Current state-of-the-art solvers repeatedly solve a particular subproblem approximately using an iterative strategy. In this work, we propose an algorithm that solves the subproblem exactly using a novel reparametrization of the Weston-Watkins dual problem. For linear WW-SVMs, our solver shows significant speed-up over the state-of-the-art solver when the number of classes is large. Our exact subproblem solver also allows us to prove linear convergence of the overall solver.
Multiclass extensions of the support vector machine (SVM) have been formulated in a variety of ways. A recent empirical comparison of nine such formulations [Doǧan et al. 2016] recommends the variant proposed by Weston and Watkins (WW), despite the fact that the WW-hinge loss is not calibrated with respect to the 0-1 loss. In this work we introduce a novel discrete loss function for multiclass classification, the ordered partition loss, and prove that the WW-hinge loss is calibrated with respect to this loss. We also argue that the ordered partition loss is maximally informative among discrete losses satisfying this property. Finally, we apply our theory to justify the empirical observation made by Doǧan et al. that the WW-SVM can work well even under massive label noise, a challenging setting for multiclass SVMs.
This article generalize the classical Goldstine-Weston theorem on normed spaces to one on random normed modules: the image of a random normed module $(E,\|\cdot\|)$ under the random natural embedding $J$ is dense in its double random conjugate space $E^{**}$ with respect to the $(ε,λ)$ weak star topology; and $J(E)$ is also dense in $E^{**}$ with respect to the locally $L^{0}$-convex weak star topology if $E$ has the countable concatenation property.
We give an algorithm to determine factorization types of primes in the number fields generated by a single point of odd order on an elliptic curve. We apply this to compute coefficients of the Dedekind zeta function of the field.
Universal basic income (UBI) is a tax scheme that uniformly redistributes aggregate income amongst the entire population of an economy. We prove the existence of an equilibrium in a model that implements universal basic income. The economic agents choose the proportion of their time to work and earn wages that can be used towards consumption and investment in a financial market with a traded stock and annuity. A proportion of the earned wages is uniformly distributed amongst all agents, leading to interconnectedness of the agents' decision problems, which are already dependent on one another through the financial market. The decision problems are further entangled by Nash perceptions of labor; the agents respond to the labor choices of others and act upon their perceived income in their decision problems. The equilibrium is constructed and proven to exist using a backward stochastic differential equation (BSDE) approach for a BSDE system with a quadratic structure that decouples. We analyze the effects of a universal basic income policy on labor market participation, the stock market, and welfare. While universal basic income policies affect labor market participation and welfare m
We prove the existence of a Radner equilibrium in a model with population growth and analyze the effects on asset prices. A finite population of agents grows indefinitely at a Poisson rate, while receiving unspanned income and choosing between consumption and investing into an annuity with infinitely-lived exponential preferences. After establishing the existence of an equilibrium for a truncated number of agents, we prove that an equilibrium exists for the model with unlimited population growth. Our numerics show that increasing the birth rate reduces oscillations in the equilibrium annuity price, and when younger agents prioritize the present more than older agents, the equilibrium annuity price rises compared to a uniform demographic.
Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration between human researchers and AIs to achieve co-superintelligence. That is, specifically targeting improving AI systems' ability to work with human researchers to conduct AI research together, from ideation to experimentation, in order to both accelerate AI research and to generally endow both AIs and humans with safer superintelligence through their symbiosis. Focusing on including human research improvement in the loop will both get us there faster, and more safely.
Nontrivial $p$-polygonal equalities impose certain conditions on the geometry of a metric space $(X,d)$ and so it is of interest to be able to identify the values of $p \in [0,\infty)$ for which such equalities exist. Following work of Li and Weston, Kelleher, Miller, Osborn and Weston established that if a metric space $(X,d)$ is of $p$-negative type, then $(X,d)$ admits no nontrivial $p$-polygonal equalities if and only if it is of strict $p$-negative type. In this note we remove the underlying premise of $p$-negative type from this theorem. As an application we show that the set of all $p$ for which a finite metric space $(X,d)$ admits a nontrivial $p$-polygonal equality is always a closed interval of the form $[\wp, \infty)$, where $\wp > 0$, or the empty set. It follows that for each $q ot= 2$, the Schatten $q$-class $\mathcal{C}_{q}$ admits a nontrivial $p$-polygonal equality for each $p > 0$. Other spaces with this same property include $C[0, 1]$ and $\ell_{q}^{(3)}$ for all $q > 2$.
Large language models (LLMs) can spend extra compute during inference to generate intermediate thoughts, which helps to produce better final responses. Since Chain-of-Thought (Wei et al., 2022), many such System 2 techniques have been proposed such as Rephrase and Respond (Deng et al., 2023a), System 2 Attention (Weston and Sukhbaatar, 2023) and Branch-Solve-Merge (Saha et al., 2023). In this work we investigate self-supervised methods to ``compile'' (distill) higher quality outputs from System 2 techniques back into LLM generations without intermediate reasoning token sequences, as this reasoning has been distilled into System 1. We show that several such techniques can be successfully distilled, resulting in improved results compared to the original System 1 performance, and with less inference cost than System 2. We posit that such System 2 distillation will be an important feature of future continually learning AI systems, enabling them to focus System 2 capabilities on the reasoning tasks that they cannot yet do well.
We prove the existence of a continuous-time Radner equilibrium with multiple agents and transaction costs. The agents are incentivized to trade towards a targeted number of shares throughout the trading period and seek to maximize their expected wealth minus a penalty for deviating from their targets. Their wealth is further reduced by transaction costs that are proportional to the number of stock shares traded. The agents' targeted number of shares is publicly known, making the resulting equilibrium fully revealing. In equilibrium, each agent optimally chooses to trade for an initial time interval before stopping trade. Our equilibrium construction and analysis involves identifying the order in which the agents stop trade. The transaction cost level impacts the equilibrium stock price drift. We analyze the equilibrium outcomes and provide numerical examples.
Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To help rectify these issues, we introduce System 2 Attention (S2A), which leverages the ability of LLMs to reason in natural language and follow instructions in order to decide what to attend to. S2A regenerates the input context to only include the relevant portions, before attending to the regenerated context to elicit the final response. In experiments, S2A outperforms standard attention-based LLMs on three tasks containing opinion or irrelevant information, QA, math word problems and longform generation, where S2A increases factuality and objectivity, and decreases sycophancy.
This paper studies carbon taxes effectiveness to induce a transition to cleaner production when a firm faces different technologies and demands. To determine carbon taxes effectiveness, we propose a framework based on a strategic capacity planning under carbon taxes model, that consider proper perfomance measures. The model, which is formulated as a mixed integer linear problem (MILP), considers issues that previous work have not studied jointly, such as machine replacement, workforce planning, and maintenance. The effectiveness measures consider levels of clean production and periods to reach a technological transition. Our computational experiments, based on a real case, have shown that carbon taxes by themselves do not necessarily induce a transition to clean production, since their effectiveness depends on the available technology relationship and the demand magnitude.
A limited participation economy models the real-world phenomenon that some economic agents have access to more of the financial market than others. We prove the global existence of a Radner equilibrium with limited participation, where the agents have exponential preferences and derive utility from both running consumption and terminal wealth. Our analysis centers around the existence and uniqueness of a solution to a coupled system of quadratic backward stochastic differential equations (BSDEs). We prove that the BSDE system has a unique $\mathcal{S}^\infty\times\text{bmo}$ solution. We define a candidate equilibrium in terms of the BSDE solution and prove through a verification argument that the candidate is a Radner equilibrium with limited participation. This work generalizes the model of Basak and Cuoco (1998) to allow for a stock with a general dividend stream and agents with exponential preferences. We also provide an explicit example.
Doust and Weston introduced a new method called "enhanced negative type" for calculating a non trivial lower bound p(T) on the supremal strict p-negative type of any given finite metric tree (T,d). In the context of finite metric trees any such lower bound p(T) > 1 is deemed to be non trivial. In this paper we refine the technique of enhanced negative type and show how it may be applied more generally to any finite metric space (X,d) that is known to have strict p-negative type for some non negative p. This allows us to significantly improve the lower bounds on the supremal strict p-negative type of finite metric trees that were given by Doust and Weston and, moreover, leads in to one of our main results: The supremal p-negative type of a finite metric space cannot be strict. By way of application we are then able to exhibit large classes of finite metric spaces (such as finite isometric subspaces of Hadamard manifolds) that must have strict p-negative type for some p > 1. We also show that if a metric space (finite or otherwise) has p-negative type for some p > 0, then it must have strict q-negative type for all q in [0,p). This generalizes a well known theorem of Schoenb
We introduce a neural network with a recurrent attention model over a possibly large external memory. The architecture is a form of Memory Network (Weston et al., 2015) but unlike the model in that work, it is trained end-to-end, and hence requires significantly less supervision during training, making it more generally applicable in realistic settings. It can also be seen as an extension of RNNsearch to the case where multiple computational steps (hops) are performed per output symbol. The flexibility of the model allows us to apply it to tasks as diverse as (synthetic) question answering and to language modeling. For the former our approach is competitive with Memory Networks, but with less supervision. For the latter, on the Penn TreeBank and Text8 datasets our approach demonstrates comparable performance to RNNs and LSTMs. In both cases we show that the key concept of multiple computational hops yields improved results.
A long-term goal of machine learning research is to build an intelligent dialog agent. Most research in natural language understanding has focused on learning from fixed training sets of labeled data, with supervision either at the word level (tagging, parsing tasks) or sentence level (question answering, machine translation). This kind of supervision is not realistic of how humans learn, where language is both learned by, and used for, communication. In this work, we study dialog-based language learning, where supervision is given naturally and implicitly in the response of the dialog partner during the conversation. We study this setup in two domains: the bAbI dataset of (Weston et al., 2015) and large-scale question answering from (Dodge et al., 2015). We evaluate a set of baseline learning strategies on these tasks, and show that a novel model incorporating predictive lookahead is a promising approach for learning from a teacher's response. In particular, a surprising result is that it can learn to answer questions correctly without any reward-based supervision at all.
Researchers are applying evolutionary theory to cancer by changing treatments before tumors have time to develop resistance。 Mathematical models suggest that rapid, carefully timed switches between multiple therapies could improve cure rates
Researchers have created cosmic dust from scratch by recreating space-like conditions inside glass tubes。 The dust contains complex carbon-rich molecules built from elements essential to life and produces infrared signals similar to real material found in space。 By studying these laboratory samples, scientists can explore how organic chemistry unfo
"Here’s a more natural, flowing version of that section