Recent works have studied implicit biases in deep learning, especially the behavior of last-layer features and classifier weights. However, they usually need to simplify the intermediate dynamics under gradient flow or gradient descent due to the intractability of loss functions and model architectures. In this paper, we introduce the unhinged loss, a concise loss function, that offers more mathematical opportunities to analyze the closed-form dynamics while requiring as few simplifications or assumptions as possible. The unhinged loss allows for considering more practical techniques, such as time-vary learning rates and feature normalization. Based on the layer-peeled model that views last-layer features as free optimization variables, we conduct a thorough analysis in the unconstrained, regularized, and spherical constrained cases, as well as the case where the neural tangent kernel remains invariant. To bridge the performance of the unhinged loss to that of Cross-Entropy (CE), we investigate the scenario of fixing classifier weights with a specific structure, (e.g., a simplex equiangular tight frame). Our analysis shows that these dynamics converge exponentially fast to a soluti
Convex potential minimisation is the de facto approach to binary classification. However, Long and Servedio [2010] proved that under symmetric label noise (SLN), minimisation of any convex potential over a linear function class can result in classification performance equivalent to random guessing. This ostensibly shows that convex losses are not SLN-robust. In this paper, we propose a convex, classification-calibrated loss and prove that it is SLN-robust. The loss avoids the Long and Servedio [2010] result by virtue of being negatively unbounded. The loss is a modification of the hinge loss, where one does not clamp at zero; hence, we call it the unhinged loss. We show that the optimal unhinged solution is equivalent to that of a strongly regularised SVM, and is the limiting solution for any convex potential; this implies that strong l2 regularisation makes most standard learners SLN-robust. Experiments confirm the SLN-robustness of the unhinged loss.
Van Rooyen et al. introduced a notion of convex loss functions being robust to random classification noise, and established that the "unhinged" loss function is robust in this sense. In this note we study the accuracy of binary classifiers obtained by minimizing the unhinged loss, and observe that even for simple linearly separable data distributions, minimizing the unhinged loss may only yield a binary classifier with accuracy no better than random guessing.
Labeling a training set is often expensive and susceptible to errors, making the design of robust loss functions for label noise an important problem. The symmetry condition provides theoretical guarantees for robustness to such noise. In this work, we study a symmetrization method arising from the unique decomposition of any multi-class loss function into a symmetric component and a class-insensitive term. In particular, symmetrizing the cross-entropy loss leads to a linear multi-class extension of the unhinged loss. Unlike in the binary case, the multi-class version must have specific coefficients in order to satisfy the symmetry condition. Under suitable assumptions, we show that this multi-class unhinged loss is the unique convex multi-class symmetric loss. We also show that it has a fundamental local role: the linear approximation of any symmetric loss around score vectors with equal components is equivalent to the multi-class unhinged loss. We then introduce SGCE and alpha-MAE, two loss functions that interpolate between the multi-class unhinged loss and the Mean Absolute Error while allowing control of the beta-smoothness of the loss. Experiments on standard noisy-label benc
When should we defer to AI outputs over human expert judgment? Drawing on recent work in social epistemology, I motivate the idea that some AI systems qualify as Artificial Epistemic Authorities (AEAs) due to their demonstrated reliability and epistemic superiority. I then introduce AI Preemptionism, the view that AEA outputs should replace rather than supplement a user's independent epistemic reasons. I show that classic objections to preemptionism - such as uncritical deference, epistemic entrenchment, and unhinging epistemic bases - apply in amplified form to AEAs, given their opacity, self-reinforcing authority, and lack of epistemic failure markers. Against this, I develop a more promising alternative: a total evidence view of AI deference. According to this view, AEA outputs should function as contributory reasons rather than outright replacements for a user's independent epistemic considerations. This approach has three key advantages: (i) it mitigates expertise atrophy by keeping human users engaged, (ii) it provides an epistemic case for meaningful human oversight and control, and (iii) it explains the justified mistrust of AI when reliability conditions are unmet. While d
Understanding the role of lattice geometry in shaping topological states and their properties is of fundamental importance to condensed matter and device physics. Here we demonstrate how an anisotropic crystal lattice drives a topological hybrid nodal line in transition metal tetraphosphides $Tm$P$_4$ ($Tm$ = Transition metal). $Tm$P$_4$ constitutes a unique class of black phosphorus materials formed by intercalating transition metal ions between the phosphorus layers without destroying the characteristic anisotropic band structure of the black phosphorous. Based on the first-principles calculations and $k \cdot p$ theory, we show that $Tm$P$_4$ harbor a single hybrid nodal line formed between oppositely-oriented anisotropic $Tm~d$ and P states unhinged from the high-symmetry planes. The nodal line consists of both type-I and type-II nodal band crossings whose nature and location are determined by the effective-mass anisotropies of the intersecting bands. We further discuss a possible topological phase transition to exemplify the formation of the hybrid nodal line state in $Tm$P$_4$. Our results offer a comprehensive study for understanding the interplay between structural motifs-d
It has recently been shown that supervised learning with the popular logistic loss is equivalent to optimizing the exponential loss over sufficient statistics about the class: Rademacher observations (rados). We first show that this unexpected equivalence can actually be generalized to other example / rado losses, with necessary and sufficient conditions for the equivalence, exemplified on four losses that bear popular names in various fields: exponential (boosting), mean-variance (finance), Linear Hinge (on-line learning), ReLU (deep learning), and unhinged (statistics). Second, we show that the generalization unveils a surprising new connection to regularized learning, and in particular a sufficient condition under which regularizing the loss over examples is equivalent to regularizing the rados (with Minkowski sums) in the equivalent rado loss. This brings simple and powerful rado-based learning algorithms for sparsity-controlling regularization, that we exemplify on a boosting algorithm for the regularized exponential rado-loss, which formally boosts over four types of regularization, including the popular ridge and lasso, and the recently coined slope --- we obtain the first p
We prove that any finite collection of polygons of equal area has a common hinged dissection. That is, for any such collection of polygons there exists a chain of polygons hinged at vertices that can be folded in the plane continuously without self-intersection to form any polygon in the collection. This result settles the open problem about the existence of hinged dissections between pairs of polygons that goes back implicitly to 1864 and has been studied extensively in the past ten years. Our result generalizes and indeed builds upon the result from 1814 that polygons have common dissections (without hinges). We also extend our common dissection result to edge-hinged dissections of solid 3D polyhedra that have a common (unhinged) dissection, as determined by Dehn's 1900 solution to Hilbert's Third Problem. Our proofs are constructive, giving explicit algorithms in all cases. For a constant number of planar polygons, both the number of pieces and running time required by our construction are pseudopolynomial. This bound is the best possible, even for unhinged dissections. Hinged dissections have possible applications to reconfigurable robotics, programmable matter, and nanomanufac
Creativity is a deeply debated topic, as this concept is arguably quintessential to our humanity. Across different epochs, it has been infused with an extensive variety of meanings relevant to that era. Along these, the evolution of technology have provided a plurality of novel tools for creative purposes. Recently, the advent of Artificial Intelligence (AI), through deep learning approaches, have seen proficient successes across various applications. The use of such technologies for creativity appear in a natural continuity to the artistic trend of this century. However, the aura of a technological artefact labeled as intelligent has unleashed passionate and somewhat unhinged debates on its implication for creative endeavors. In this paper, we aim to provide a new perspective on the question of creativity at the era of AI, by blurring the frontier between social and computational sciences. To do so, we rely on reflections from social science studies of creativity to view how current AI would be considered through this lens. As creativity is a highly context-prone concept, we underline the limits and deficiencies of current AI, requiring to move towards artificial creativity. We ar
In this unhinged rant, I lay out my suspicion that a lot of visualizations are bullshit: charts that do not have even the common decency to intentionally lie but are totally unconcerned about the state of the world or any practical utility. I suspect that bullshit charts take up a large fraction of the time and attention of actual visualization producers and consumers, and yet are seemingly absent from academic research into visualization design.
We show that the chiral Dirac and Majorana hinge modes in three-dimensional higher-order topological insulators (HOTIs) and superconductors (HOTSCs) can be gapped while preserving the protecting $\mathsf{C}_{2n}\mathcal T$ symmetry upon the introduction of non-Abelian surface topological order. In both cases, the topological order on a single side surface breaks time reversal symmetry, but appears with its time-reversal conjugate on alternating sides in a $\mathsf{C}_{2n}\mathcal T$ preserving pattern. In the absence of the HOTI/HOTSC bulk, such a pattern necessarily involves gapless chiral modes on hinges between $\mathsf{C}_{2n}\mathcal T$-conjugate domains. However, using a combination of $K$-matrix and anyon condensation arguments, we show that on the boundary of a 3D HOTI/HOTSC these topological orders are fully gapped and hence `anomalous'. Our results suggest that new patterns of surface and hinge states can be engineered by selectively introducing topological order only on specific surfaces.
All knowing is material. The challenge for Information Systems (IS) research is to specify how knowing is material by drawing on theoretical characterizations of the digital. Synthetic knowing is knowing informed by theorizing digital materiality. We focus on two defining qualities: liquefaction (unhinging digital representations from physical objects, qualities, or processes) and open-endedness (extendable and generative). The Internet of Things (IoT) is crucial because sensors are vehicles of liquefaction. Their expanding scope for real-time seeing, hearing, tasting, smelling, and touching increasingly mimics phenomenologically perceived reality. Empirically, we present a longitudinal case study of IoT-rendered marine environmental monitoring by an oil and gas company operating in the politically contested Arctic. We characterize synthetic knowing into four concepts, the former three tied to liquefaction and the latter to open-endedness: (i) the objects of knowing are algorithmic phenomena; (ii) the sensors increasingly conjure up phenomenological reality; (iii) knowing is scoped (configurable); and (iv) open knowing/data is politically charged.
This work proposes the Bregman-Tweedie classification model and analyzes the domain structure of the extended exponential function, an extension of the classic generalized exponential function with additional scaling parameter, and related high-level mathematical structures, such as the Bregman-Tweedie loss function and the Bregman-Tweedie divergence. The base function of this divergence is the convex function of Legendre type induced from the extended exponential function. The Bregman-Tweedie loss function of the proposed classification model is the regular Legendre transformation of the Bregman-Tweedie divergence. This loss function is a polynomial parameterized function between unhinge loss and the logistic loss function. Actually, we have two sub-models of the Bregman-Tweedie classification model; H-Bregman with hinge-like loss function and L-Bregman with logistic-like loss function. Although the proposed classification model is nonconvex and unbounded, empirically, we have observed that the H-Bregman and L-Bregman outperform, in terms of the Friedman ranking, logistic regression and SVM and show reasonable performance in terms of the classification accuracy in the category of
Ranaspumin-2 (Rsn-2) is a surfactant protein found in the foam nests of the túngara frog. Previous experimental work has led to a proposed model of adsorption which involves an unusual clam shell-like `unhinging' of the protein at an interface. Interestingly, there is no concomitant denaturation of the secondary structural elements of Rsn-2 with the large scale transformation of its tertiary structure. In this work we use both experiment and simulation to better understand the driving forces underpinning this unusual process. We develop a modified Gō-model approach where we have included explicit representation of the side-chains in order to realistically model the interaction between the secondary structure elements of the protein and the interface. Doing so allows for the study of the underlying energy landscape which governs the mechanism of Rsn-2 interfacial adsorption. Experimentally, we study targeted mutants of Rsn-2, using the Langmuir trough, pendant drop tensiometry and circular dichroism, to demonstrate that the clam-shell model is correct. We find that Rsn-2 adsorption is in fact a two-step process: the hydrophobic N-terminal tail recruits the protein to the interface a