共找到 20 条结果
Open-world machine learning is an emerging technique in artificial intelligence, where conventional machine learning models often follow closed-world assumptions, which can hinder their ability to retain previously learned knowledge for future tasks. However, automated intelligence systems must learn about novel classes and previously known tasks. The proposed model offers novel learning classes in an open and continuous learning environment. It consists of two different but connected tasks. First, it discovers unknown classes in the data and creates novel classes; next, it learns how to perform class incrementally for each new class. Together, they enable continual learning, allowing the system to expand its understanding of the data and improve over time. The proposed model also outperformed existing approaches in open-world learning. Furthermore, it demonstrated strong performance in continuous learning, achieving a highest average accuracy of 82.54% over four iterations and a minimum accuracy of 65.87%.
Unit-generated orders of a quadratic field are orders of the form $\mathcal{O} = \mathbb{Z}[\varepsilon]$, where $\varepsilon$ is a unit in the quadratic field. If the order $\mathcal{O}$ is a maximal order of a real quadratic field, then the quadratic number field is necessarily of a restricted form, being of narrow Richaud--Degert type. However, every real quadratic field contains infinitely many distinct unit-generated orders. They are parametrized as $\mathcal{O} = \mathcal{O}_{n}^{\pm}$ having quadratic discriminants $Δ(\mathcal{O}) = Δ_{n}^{+} = n^2 - 4$ (for $n \geq 3$) and $Δ(\mathcal{O}) = Δ_{n}^{-} = n^2 + 4$ (for $n \geq 1$). We show the (wide or narrow) class numbers of unit-generated orders satisfy $\log \left|{\rm Cl}(\mathcal{O})\right| \sim \log \frac{1}{2}\left|Δ(\mathcal{O})\right|$ as $\left|Δ(\mathcal{O})\right| \to \infty$, using a result of L.-K. Hua. We deduce that there are finitely many unit-generated quadratic orders of class number one and finitely many unit-generated quadratic orders whose class group is $2$-torsion. We classify all unit-generated real quadratic orders having class number one. We provide numerical lists of quadratic unit-generated orders
Selecting the best classifier among the available ones is a difficult task, especially when only instances of one class exist. In this work we examine the notion of combining one-class classifiers as an alternative for selecting the best classifier. In particular, we propose two new one-class classification performance measures to weigh classifiers and show that a simple ensemble that implements these measures can outperform the most popular one-class ensembles. Furthermore, we propose a new one-class ensemble scheme, TUPSO, which uses meta-learning to combine one-class classifiers. Our experiments demonstrate the superiority of TUPSO over all other tested ensembles and show that the TUPSO performance is statistically indistinguishable from that of the hypothetical best classifier.
We deal with a class of semilinear SPDEs driven by space-time white noise that includes the one dimensional stochastic Burgers equation. Such equations can have nonlocal and quadratic nonlinearities. We consider the problem of estimation of the diffusivity parameter in front of the second-order spatial derivative. Based on local observations in space, we study the estimator derived in [Altmeyer, Reiß, Ann. Appl. Probab.(2021)] for linear stochastic heat equation that has also been used in [Altmeyer, Cialenco, Pasemann, Bernoulli (2023)] to cover certain class of semilinear SPDEs including stochastic Burgers equations driven by trace class noise. The space-time white noise case we consider has also relevant physical motivations. After we establish new regularity results for the solution, we are able to show that our proposed estimator is strongly consistent and asymptotically normal.
This article is devoted to the study of the Schatten class membership of commutators involving singular integral operators. We utilize martingale paraproducts and Hytönen's dyadic martingale technique to obtain sufficient conditions on the weak-type and strong-type Schatten class membership of commutators in terms of Sobolev spaces and Besov spaces respectively. We also establish the complex median method, which is applicable to complex-valued functions. We apply it to get the optimal necessary conditions on the weak-type and strong-type Schatten class membership of commutators associated with non-degenerate kernels. This resolves the problem on the characterization of the weak-type and strong-type Schatten class membership of commutators. Our new approach is built on Hytönen's dyadic martingale technique and the complex median method. Compared with all the previous ones, this new one is more powerful in several aspects: $(a)$ it permits us to deal with more general singular integral operators with little smoothness; $(b)$ it allows us to deal with commutators with complex-valued kernels; $(c)$ it turns out to be powerful enough to deal with the weak-type and strong-type Schatten c
The principle of open class determinacy is preserved by pre-tame class forcing, and after such forcing, every new class well-order is isomorphic to a ground-model class well-order. Similarly, the principle of elementary transfinite recursion $\text{ETR}_Γ$ for a fixed class well-order $Γ$ is preserved by pre-tame class forcing. The full principle ETR itself is preserved by countably strategically closed pre-tame class forcing, and after such forcing, every new class well-order is isomorphic to a ground-model class well-order. Meanwhile, it remains open whether ETR is preserved by all forcing, including the forcing merely to add a Cohen real.
We show that Kelley-Morse set theory does not prove the class Fodor principle, the assertion that every regressive class function $F:S\to\text{Ord}$ defined on a stationary class $S$ is constant on a stationary subclass. Indeed, it is relatively consistent with KM for any infinite $λ$ with $ω\leqλ\leq\text{Ord}$ that there is a class function $F:\text{Ord}\toλ$ that is not constant on any stationary class. Strikingly, it is consistent with KM that there is a class $A\subseteqω\times\text{Ord}$, such that each section $A_n=\{α\mid (n,α)\in A\}$ contains a class club, but $\bigcap_n A_n$ is empty. Consequently, it is relatively consistent with KM that the class club filter is not $σ$-closed.
Despite progress in adversarial training (AT), there is a substantial gap between the top-performing and worst-performing classes in many datasets. For example, on CIFAR10, the accuracies for the best and worst classes are 74% and 23%, respectively. We argue that this gap can be reduced by explicitly optimizing for the worst-performing class, resulting in a min-max-max optimization formulation. Our method, called class focused online learning (CFOL), includes high probability convergence guarantees for the worst class loss and can be easily integrated into existing training setups with minimal computational overhead. We demonstrate an improvement to 32% in the worst class accuracy on CIFAR10, and we observe consistent behavior across CIFAR100 and STL10. Our study highlights the importance of moving beyond average accuracy, which is particularly important in safety-critical applications.
In light of the celebrated theorem of Vopěnka (1972), proving in ZFC that every set is generic over HOD, it is natural to inquire whether the set-theoretic universe $V$ must be a class-forcing extension of HOD by some possibly proper-class forcing notion in HOD. We show, negatively, that if ZFC is consistent, then there is a model of ZFC that is not a class-forcing extension of its HOD for any class forcing notion definable in HOD and with definable forcing relations there (allowing parameters). Meanwhile, S. Friedman (2012) showed, positively, that if one augments HOD with a certain ZFC-amenable class $A$, definable in $V$, then the set-theoretic universe $V$ is a class-forcing extension of the expanded structure $\langle\text{HOD},\in,A\rangle$. Our result shows that this augmentation process can be necessary. The same example shows that $V$ is not necessarily a class-forcing extension of the mantle, and the method provides counterexamples to the intermediate model property, namely, a class-forcing extension $V\subseteq W\subseteq V[G]$ with an intermediate transitive inner model $W$ that is neither a class-forcing extension of $V$ nor a ground model of $V[G]$ by any definable cl
The principle of open determinacy for class games---two-player games of perfect information with plays of length $ω$, where the moves are chosen from a possibly proper class, such as games on the ordinals---is not provable in Zermelo-Fraenkel set theory ZFC or Gödel-Bernays set theory GBC, if these theories are consistent, because provably in ZFC there is a definable open proper class game with no definable winning strategy. In fact, the principle of open determinacy and even merely clopen determinacy for class games implies Con(ZFC) and iterated instances Con(Con(ZFC)) and more, because it implies that there is a satisfaction class for first-order truth and indeed a transfinite tower of truth predicates for iterated truth-about-truth, relative to any class parameter. This is perhaps explained, in light of the Tarskian recursive definition of truth, by the more general fact that the principle of clopen determinacy is exactly equivalent over GBC to the principle of elementary transfinite recursion ETR over well-founded class relations. Meanwhile, the principle of open determinacy for class games is provable in the stronger theory GBC+$Π^1_1$-comprehension, a proper fragment of Kelle
The class forcing theorem, which asserts that every class forcing notion $\mathbb{P}$ admits a forcing relation $\Vdash_{\mathbb{P}}$, that is, a relation satisfying the forcing relation recursion -- it follows that statements true in the corresponding forcing extensions are forced and forced statements are true -- is equivalent over Gödel-Bernays set theory GBC to the principle of elementary transfinite recursion $\text{ETR}_{\text{Ord}}$ for class recursions of length $\text{Ord}$. It is also equivalent to the existence of truth predicates for the infinitary languages $\mathcal{L}_{\text{Ord},ω}(\in,A)$, allowing any class parameter $A$; to the existence of truth predicates for the language $\mathcal{L}_{\text{Ord},\text{Ord}}(\in,A)$; to the existence of $\text{Ord}$-iterated truth predicates for first-order set theory $\mathcal{L}_{ω,ω}(\in,A)$; to the assertion that every separative class partial order $\mathbb{P}$ has a set-complete class Boolean completion; to a class-join separation principle; and to the principle of determinacy for clopen class games of rank at most $\text{Ord}+1$. Unlike set forcing, if every class forcing notion $\mathbb{P}$ has a forcing relation merely
In this work, we define cost-free learning (CFL) formally in comparison with cost-sensitive learning (CSL). The main difference between them is that a CFL approach seeks optimal classification results without requiring any cost information, even in the class imbalance problem. In fact, several CFL approaches exist in the related studies, such as sampling and some criteria-based pproaches. However, to our best knowledge, none of the existing CFL and CSL approaches are able to process the abstaining classifications properly when no information is given about errors and rejects. Based on information theory, we propose a novel CFL which seeks to maximize normalized mutual information of the targets and the decision outputs of classifiers. Using the strategy, we can deal with binary/multi-class classifications with/without abstaining. Significant features are observed from the new strategy. While the degree of class imbalance is changing, the proposed strategy is able to balance the errors and rejects accordingly and automatically. Another advantage of the strategy is its ability of deriving optimal rejection thresholds for abstaining classifications and the "equivalent" costs in binary
Class-Incremental Learning (CIL) aims to learn a classification model with the number of classes increasing phase-by-phase. An inherent problem in CIL is the stability-plasticity dilemma between the learning of old and new classes, i.e., high-plasticity models easily forget old classes, but high-stability models are weak to learn new classes. We alleviate this issue by proposing a novel network architecture called Adaptive Aggregation Networks (AANets), in which we explicitly build two types of residual blocks at each residual level (taking ResNet as the baseline architecture): a stable block and a plastic block. We aggregate the output feature maps from these two blocks and then feed the results to the next-level blocks. We adapt the aggregation weights in order to balance these two types of blocks, i.e., to balance stability and plasticity, dynamically. We conduct extensive experiments on three CIL benchmarks: CIFAR-100, ImageNet-Subset, and ImageNet, and show that many existing CIL methods can be straightforwardly incorporated into the architecture of AANets to boost their performances.
We introduce and study a class of non-Hermitian Hamiltonians which have velocity dependent potentials. Since stability can not be advocated directly from the classical potential, we show that the energy spectra are real and bounded from below which proves the stability of the spectra of all members in the class. We find that the introduced class of non-Hermitian Hamiltonians do have a corresponding superpartner class of non-Hermitian Hamiltonians. We were able to introduce supercharges which in conjunction with the corresponding super Hamiltonians constitute a closed super algebra. Among the introduced Hamiltonians, we show that non-$\mathcal{PT }$-symmetric Hamiltonians can be transformed into their corresponding superpartner Hamiltonians via a specific canonical transformation while the $\mathcal{PT }$-symmetric ones failed to be mapped to their corresponding superpartner Hamiltonians via the same canonical transformation. Since canonical transformations preserve the spectrum, we conclude that non-$\mathcal{PT }$-symmetric Hamiltonians out of the introduced class of Hamiltonians have the same spectrum as the corresponding superpartner Hamiltonians and thus Susy is broken for such
With the rapid increase of large-scale, real-world datasets, it becomes critical to address the problem of long-tailed data distribution (i.e., a few classes account for most of the data, while most classes are under-represented). Existing solutions typically adopt class re-balancing strategies such as re-sampling and re-weighting based on the number of observations for each class. In this work, we argue that as the number of samples increases, the additional benefit of a newly added data point will diminish. We introduce a novel theoretical framework to measure data overlap by associating with each sample a small neighboring region rather than a single point. The effective number of samples is defined as the volume of samples and can be calculated by a simple formula $(1-β^{n})/(1-β)$, where $n$ is the number of samples and $β\in [0,1)$ is a hyperparameter. We design a re-weighting scheme that uses the effective number of samples for each class to re-balance the loss, thereby yielding a class-balanced loss. Comprehensive experiments are conducted on artificially induced long-tailed CIFAR datasets and large-scale datasets including ImageNet and iNaturalist. Our results show that wh
Deep neural network classifiers partition input space into high confidence regions for each class. The geometry of these class manifolds (CMs) is widely studied and intimately related to model performance; for example, the margin depends on CM boundaries. We exploit the notions of Gaussian width and Gordon's escape theorem to tractably estimate the effective dimension of CMs and their boundaries through tomographic intersections with random affine subspaces of varying dimension. We show several connections between the dimension of CMs, generalization, and robustness. In particular we investigate how CM dimension depends on 1) the dataset, 2) architecture (including ResNet, WideResNet \& Vision Transformer), 3) initialization, 4) stage of training, 5) class, 6) network width, 7) ensemble size, 8) label randomization, 9) training set size, and 10) robustness to data corruption. Together a picture emerges that higher performing and more robust models have higher dimensional CMs. Moreover, we offer a new perspective on ensembling via intersections of CMs. Our code is at https://github.com/stanislavfort/slice-dice-optimize/
Graph-modification problems, where we modify a graph by adding or deleting vertices or edges or contracting edges to obtain a graph in a {\it simpler} class, is a well-studied optimization problem in all algorithmic paradigms including classical, approximation and parameterized complexity. Specifically, graph-deletion problems, where one needs to delete a small number of vertices to make the resulting graph to belong to a given non-trivial hereditary graph class, captures several well-studied problems including {\sc Vertex Cover}, {\sc Feedback Vertex Set}, {\sc Odd Cycle Transveral}, {\sc Cluster Vertex Deletion}, and {\sc Perfect Deletion}. Investigation into these problems in parameterized complexity has given rise to powerful tools and techniques. We initiate a study of a natural variation of the problem of deletion to {\it scattered graph classes}. We want to delete at most $k$ vertices so that in the resulting graph, each connected component belongs to one of a constant number of graph classes. As our main result, we show that this problem is fixed-parameter tractable (FPT) when the deletion problem corresponding to each of the finite number of graph classes is known to be FP
In this paper we prove that finite index subgroups of genus 3 mapping class and Torelli groups that contain the group generated by Dehn twists on bounding simple closed curves are not Kahler. These results are deduced from explicit presentations of the unipotent (aka, Malcev) completion of genus 3 Torelli groups and of the relative completions of genus 3 mapping class groups. The main results follow from the fact that these presentations are not quadratic. To complete the picture, we compute presentations of completed Torelli and mapping class in genera > 3; they are quadratic. We also show that groups commensurable with hyperelliptic mapping class groups and mapping class groups in genera < 3 are not Kahler.
In this note we present a new self-contained approach to the class field theory of arithmetic schemes in the sense of Wiesend. Along the way we prove new results on space filling curves on arithmetic schemes and on the class field theory of local rings. We show how one can deduce the more classical version of higher global class field theory due to Kato and Saito from Wiesend's version. One of our new results says that the connected component of the identity element in Wiesend's class group is divisible if some obstruction is absent.
Online class imbalance learning constitutes a new problem and an emerging research topic that focusses on the challenges of online learning under class imbalance and concept drift. Class imbalance deals with data streams that have very skewed distributions while concept drift deals with changes in the class imbalance status. Little work exists that addresses these challenges and in this paper we introduce queue-based resampling, a novel algorithm that successfully addresses the co-existence of class imbalance and concept drift. The central idea of the proposed resampling algorithm is to selectively include in the training set a subset of the examples that appeared in the past. Results on two popular benchmark datasets demonstrate the effectiveness of queue-based resampling over state-of-the-art methods in terms of learning speed and quality.