共找到 20 条结果
The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as how best to include under-represented groups in data-driven policy and decision making, or the health benefits of assistive technologies such as wearables. We provide a conceptual overview, in particular contrasting small data with big data, and identify common themes from exemplary case studies and application areas. Potential solutions are described in a more detailed technical overview of current data analysis and modelling techniques, highlighting contributions from different disciplines, such as knowledge-driven modelling from statistics and data-driven modelling from computer science. By linking application settings, conceptual contributions and specific techniques, we highlight what is already feasible and suggest what an agenda for fully leveraging small data might look like.
We consider the possibility of designing an election method that eliminates the incentives for a voter to rank any other candidate equal to or ahead of his or her sincere favorite. We refer to these methods as satisfying the ``Strong Favorite Betrayal Criterion" (SFBC). Methods satisfying our strategic criteria can be classified into four categories, according to their geometrical properties. We prove that two categories of methods are highly restricted and closely related to positional methods (point systems) that give equal points to a voter's first and second choices. The third category is tightly restricted, but if criteria are relaxed slightly a variety of interesting methods can be identified. Finally, we show that methods in the fourth category are largely irrelevant to public elections. Interestingly, most of these methods for satisfying the SFBC do so only ``weakly," in that these methods make no meaningful distinction between the first and second place on the ballot. However, when we relax our conditions and allow (but do not require) equal rankings for first place, a wider range of voting methods are possible, and these methods do indeed make meaningful distinctions betw
Large Language Models (LLMs) have made significant progress in advancing artificial general intelligence (AGI), leading to the development of increasingly large models such as GPT-4 and LLaMA-405B. However, scaling up model sizes results in exponentially higher computational costs and energy consumption, making these models impractical for academic researchers and businesses with limited resources. At the same time, Small Models (SMs) are frequently used in practical settings, although their significance is currently underestimated. This raises important questions about the role of small models in the era of LLMs, a topic that has received limited attention in prior research. In this work, we systematically examine the relationship between LLMs and SMs from two key perspectives: Collaboration and Competition. We hope this survey provides valuable insights for practitioners, fostering a deeper understanding of the contribution of small models and promoting more efficient use of computational resources. The code is available at https://github.com/tigerchen52/role_of_small_models
From flocking birds to schooling fish, organisms interact to form collective dynamics across the natural world. Self-organization is present at smaller scales as well: cells interact and move during development to produce patterns in fish skin, and wound healing relies on cell migration. Across these examples, scientists are interested in shedding light on the individual behaviors informing spatial group dynamics and in predicting the patterns that will emerge under altered agent interactions. One challenge to these goals is that images of self-organization -- whether empirical or generated by models -- are qualitative. To get around this, there are many methods for transforming qualitative pattern data into quantitative information. In this tutorial chapter, I survey some methods for quantifying self-organization, including order parameters, pair correlation functions, and techniques from topological data analysis. I also discuss some places that I see as especially promising for quantitative data, modeling, and data-driven approaches to continue meeting in the future.
Histogram-based empirical Bayes methods developed for analyzing data for large numbers of genes, SNPs, or other biological features tend to have large biases when applied to data with a smaller number of features such as genes with expression measured conventionally, proteins, and metabolites. To analyze such small-scale and medium-scale data in an empirical Bayes framework, we introduce corrections of maximum likelihood estimators (MLE) of the local false discovery rate (LFDR). In this context, the MLE estimates the LFDR, which is a posterior probability of null hypothesis truth, by estimating the prior distribution. The corrections lie in excluding each feature when estimating one or more parameters on which the prior depends. An application of the new estimators and previous estimators to protein abundance data illustrates how different estimators lead to very different conclusions about which proteins are affected by cancer. The estimators are compared using simulated data of two different numbers of features, two different detectability levels, and all possible numbers of affected features. The simulations show that some of the corrected MLEs substantially reduce a negative bi
Based on a nonsmooth coherence condition, we construct and prove the convergence of a forward-backward splitting method that alternates between steps on a fine and a coarse grid. Our focus is a total variation regularised inverse imaging problems, specifically, their dual problems, for which we develop in detail the relevant coarse-grid problems. We demonstrate the performance of our method on total variation denoising and magnetic resonance imaging.
Geometrical methods in quantum information are very promising for both providing technical tools and intuition into difficult control or optimization problems. Moreover, they are of fundamental importance in connecting pure geometrical theories, like GR, to quantum mechanics, like in the AdS/CFT correspondence. In this paper, we first make a survey of the most important settings in which geometrical methods have proven useful to quantum information theory. Then, we lay down a general framework for an action principle for quantum resources like entanglement, coherence, and anti-flatness. We discuss the case of a two-qubit system.
We investigate the role of compressibility in the modified quasi-linear viscoelastic (MQLV) constitutive framework for soft solids at finite strain, where shear and bulk responses are governed by distinct relaxation functions. Analytical and semi-analytical results are derived for simple shear and torsion, under incompressible and slightly compressible assumptions. We show that compressibility affects the response only when volume changes occur: under isochoric deformations, the bulk contribution vanishes, while even small deviations from isochoricity significantly alter the normal response. Shear stress and torque are largely insensitive to compressibility, whereas normal stress and axial force exhibit pronounced sensitivity due to the coupling between shear and bulk relaxation. We further demonstrate that volumetric effects interact with the Poynting effect: in simple shear they oppose each other, reducing relaxation, while in torsion they reinforce each other, enhancing it. These trends agree with brain tissue experiments but reveal limitations of the slightly compressible model for highly compressible materials, such as agarose gels. Overall, the results emphasise the importanc
The finite difference time domain method is one of the simplest and most popular methods in computational electromagnetics. This work considers two possible ways of generalising it to a meshless setting by employing local radial basis function interpolation. The resulting methods remain fully explicit and are convergent if properly chosen hyperviscosity terms are added to the update equations. We demonstrate that increasing the stencil size of the approximation has a desirable effect on numerical dispersion. Furthermore, our proposed methods can exhibit a decreased dispersion anisotropy compared to the finite difference time domain method.
Models accounting for imperfect detection are important. Single-visit methods have been proposed as an alternative to multiple-visits methods to relax the assumption of closed population. Knape and Korner-Nievergelt (2015) showed that under certain models of probability of detection single-visit methods are statistically non-identifiable leading to biased population estimates. There is a close relationship between estimation of the resource selection probability function (RSPF) using weighted distributions and single-visit methods for occupancy and abundance estimation. We explain the precise mathematical conditions needed for RSPF estimation as stated in Lele and Keim (2006). The identical conditions, that remained unstated in our papers on single-visit methodology, are needed for single-visit methodology to work. We show that the class of admissible models is quite broad and does not excessively restrict the application of the RSPF or the single-visit methodology. To complement the work by Knape and Korner-Nievergelt, we study the performance of multiple-visit methods under the scaled logistic detection function and a much wider set of situations. In general, under the scaled log
This paper examines changes in occupational crowding of immigrant women in frontline industries in the United States during the onset of COVID-19, and we contextualize their experiences against the backdrop of broader race-based and gender-based occupational crowding. Building on the occupational crowding hypothesis, which suggests that marginalized workers are crowded in a small number of occupations to prop up wages of socially-privileged workers, we hypothesize that immigrant, Black, and Hispanic workers were shunted into frontline work to prop up the health of others during the pandemic. Our analysis of American Community Survey microdata indicates that immigrant workers, particularly immigrant women, were increasingly crowded in frontline work during the onset of the pandemic. We also find that US-born Black and Hispanic workers disproportionately faced COVID-19 exposure in their work, but were not increasingly crowded into frontline occupations following the onset of the pandemic. The paper also provides a rationale for considering the occupational crowding hypothesis along the dimensions of both wages and occupational health.
A constellation of remote sensing small satellite system has been developed for infrastructure monitoring in India by using SAR Payload. The LEO constellation of the small satellites is designed in a way, which can cover the entire footprint of India. Since India lies a little above the equatorial region, the orbital parameters are adjusted in a way that inclination of 36 degrees and RAAN varies from 70-130 degrees at a height of 600 km has been considered. A total number of 4 orbital planes are designed in which each orbital plane consisting 3 small satellites with 120-degrees true anomaly separation. Each satellite is capable of taking multiple look images with the minimum resolution of 1 meter per pixel and swath width of 10 km approx. The multiple look images captured by the SAR payload help in continuous infrastructure monitoring of our interested footprint area in India. Each small satellite is equipped with a communication payload that uses X-band and VHF antenna, whereas the TT&C will use a high data-rate S-band transmitter. The paper presents only a coverage metrics analysis method of our designed constellation for our India footprint by considering the important metri
Chest X-ray is a commonly used tool during triage, diagnosis and management of respiratory diseases. In resource-constricted settings, optimizing this resource can lead to valuable cost savings for the health care system and the patients as well as to and improvement in consult time. We used prospectively-collected data from 137 patients referred for chest X-ray at the Christian Medical Center and Hospital (CMCH) in Purnia, Bihar, India. Each patient provided at least five coughs while awaiting radiography. Collected cough sounds were analyzed using acoustic AI methods. Cross-validation was done on temporal and spectral features on the cough sounds of each patient. Features were summarized using standard statistical approaches. Three models were developed, tested and compared in their capacity to predict an abnormal result in the chest X-ray. All three methods yielded models that could discriminate to some extent between normal and abnormal with the logistic regression performing best with an area under the receiver operating characteristic curves ranging from 0.7 to 0.78. Despite limitations and its relatively small sample size, this study shows that AI-enabled algorithms can use
In this paper, we consider gradient methods for minimizing smooth convex functions, which employ the information obtained at the previous iterations in order to accelerate the convergence towards the optimal solution. This information is used in the form of a piece-wise linear model of the objective function, which provides us with much better prediction abilities as compared with the standard linear model. To the best of our knowledge, this approach was never really applied in Convex Minimization to differentiable functions in view of the high complexity of the corresponding auxiliary problems. However, we show that all necessary computations can be done very efficiently. Consequently, we get new optimization methods, which are better than the usual Gradient Methods both in the number of oracle calls and in the computational time. Our theoretical conclusions are confirmed by preliminary computational experiments.
Current methods for investigation of receptor - ligand interactions in drug discovery are based on three-dimensional complementarity of receptor and ligand surfaces, and they include pharmacophore modelling, QSAR, molecular docking etc. Those methods only consider short-range molecular interactions (distances <5A), and not include long-range interactions (distances >5A) which are essential for kinetic of biochemical reactions because they influence the number of productive collisions between interacting molecules. Previously was shown that the electron-ion interaction potential (EIIP) represents the physical property which determines the long-range properties of biological molecules. This molecular descriptor served as a base for development of the informational spectrum method (ISM), a virtual spectroscopy method for investigation of protein-protein interactions. In this paper, we proposed a new approach to treat small molecules as linear entities, allowing study of the small molecule - protein interaction by ISM. We analyzed here 21 sets of KEGG drug-protein interactions and showed that this new approach allows an efficient discrimination between biologically active and ina
A profinite group is called small if it has only finitely many open subgroups of index n for each positive integer n. We show that every Frattini cover of a small profinite group is small. A profinite group is called strongly complete if every subgroup of finite index is open. We show that two profinite groups that are elementarily equivalent, in the first-order language of groups, are isomorphic if one of them is strongly complete, extending a result of Moshe Jarden and Alexander Lubotzky which treats the case of finitely generated profinite groups.
The study of topological phases of light suggests novel opportunities for creating robust optical structures and on-chip photonic devices which are immune against scattering losses and structural disorder. However, many recent demonstrations of topological effects in optics employ structures with relatively large scales. Here we discuss the physics and realisation of topological photonics on small scales, with the dimensions often smaller or comparable with the wavelength of light. We highlight the recent experimental demonstrations of small-scale topological states based on arrays of resonant nanoparticles and discuss a novel photonic platform employing higher-order topological effects for creating subwavelength highly efficient topologically protected optical cavities. We pay a special attention to the recent progress on topological polaritonic structures and summarize with our vision on the future directions of nanoscale topological photonics and its impact on other fields.
This editorial from the PASP Special Focus Issue "Techniques and Methods for Astrophysical Data Visualization" summarizes contributions from authors, their software and tutorials, video abstracts, and 3D content. PASP and IOP have made this Focus Issue an ongoing project and will continue accepting submissions throughout 2017. For more information and to view the video abstract visit: http://iopscience.iop.org/journal/1538-3873/page/Techniques-and-Methods-for-Astrophysical-Data-Visualization
Random effects meta-analysis is a widely applied methodology to synthetize research findings of studies in a specific scientific question. Besides estimating the mean effect, an important aim of the meta-analysis is to summarize the heterogeneity, i.e. the variation in the underlying effects caused by the differences in study circumstances. The prediction interval is frequently used for this purpose: a 95% prediction interval contains the true effect of a similar new study in 95% of the cases when it is constructed, or in other words, it covers 95% of the true effects distribution on average. In this article, after providing a clear mathematical background, we present an extensive simulation investigating the performance of all frequentist prediction interval methods published to date. The work focuses on the distribution of the coverage probabilities and how these distributions change depending on the amount of heterogeneity and the number of involved studies. Although the single requirement that a prediction interval has to fulfill is to keep a nominal coverage probability on average, we demonstrate why the distribution of coverages cannot be disregarded, and that for small numbe
A third workshop on small-x physics, within the Small-x Collaboration, was held in Hamburg in May 2004 with the aim of overviewing recent theoretical progress in this area and summarizing the experimental status.