We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format. Thus we can approach self-evaluation on open-ended sampling tasks by asking models to first propose answers, and then to evaluate the probability "P(True)" that their answers are correct. We find encouraging performance, calibration, and scaling for P(True) on a diverse array of tasks. Performance at self-evaluation further improves when we allow models to consider many of their own samples before predicting the validity of one specific possibility. Next, we investigate whether models can be trained to predict "P(IK)", the probability that "I know" the answer to a question, without reference to any particular proposed answer. Models perform well at predicting P(IK) and partially generalize across tasks, though they struggle with calibration of P(IK) on new tasks. The predicted P(IK) probabilities also increase appropriately in the presence of relevant source materials in the context, and in the presence of hints towards the solution of mathematical word problems. We hope these observations lay the groundwork for training more honest models, and for investigating how honesty generalizes to cases where models are trained on objectives other than the imitation of human writing.
Researchers have begun to explore and identify various gradations in sexual orientation identity, paying attention to alternative sexual identity categories and attempting to clarify potential subtypes of same-sex sexuality, particularly among women. This study utilizes both quantitative and qualitative data to explore the behavioral experiences and identity development processes among women of a particular sexual identity subtype, "mostly straight." Participants were 349 female college students whose primary sexual identities included exclusively straight, mostly straight, bisexual, and lesbian. Results indicated that, on most behavioral variables, mostly straight women fell directly between and were significantly different from exclusively straight and bisexual/lesbian women. Mostly straight women were also distinct from exclusively straight women but were similar to bisexual women and lesbians on several quantitative measures of identity. Narratives about sexual identity development for mostly straight women revealed the complexities of sexual identity exploration, uncertainty, and commitment within this population. As a whole, this study encourages researchers to begin to recognize and examine mostly straight as a distinct sexual identity subtype in young women.
Intrinsically disordered proteins and regions carry out varied and vital cellular functions. Proteins with disordered regions are especially common in eukaryotic cells, with a subset of these proteins being mostly disordered, e.g., with more disordered than ordered residues. Two distinct methods have been previously described for using amino acid sequences to predict which proteins are likely to be mostly disordered. These methods are based on the net charge-hydropathy distribution and disorder prediction score distribution. Each of these methods is reexamined, and the prediction results are compared herein. A new prediction method based on consensus is described. Application of the consensus method to whole genomes reveals that approximately 4.5% of Yersinia pestis, 5% of Escherichia coli K12, 6% of Archaeoglobus fulgidus, 8% of Methanobacterium thermoautotrophicum, 23% of Arabidopsis thaliana, and 28% of Mus musculus proteins are mostly disordered. The unexpectedly high frequency of mostly disordered proteins in eukaryotes has important implications both for large-scale, high-throughput projects and also for focused experiments aimed at determination of protein structure and function.
The core methods in today's econometric toolkit are linear regression for statistical control, instrumental variables methods for the analysis of natural experiments, and differences-in-differences methods that exploit policy changes. In the modern experimentalist paradigm, these techniques address clear causal questions such as: Do smaller classes increase learning? Should wife batterers be arrested? How much does education raise wages? Mostly Harmless Econometrics shows how the basic tools of applied econometrics allow the data to speak.In addition to econometric essentials, Mostly Harmless Econometrics covers important new extensions--regression-discontinuity designs and quantile regression--as well as how to get standard errors right. Joshua Angrist and Jorn-Steffen Pischke explain why fancier econometric techniques are typically unnecessary and even dangerous. The applied econometric methods emphasized in this book are easy to use and relevant for many areas of contemporary social science. An irreverent review of econometric essentials A focus on tools that applied researchers use most Chapters on regression-discontinuity designs, quantile regression, and standard errors Many empirical examples A clear and concise resource with wide applications
Research Article| August 01, 1974 Segregation of Magma from a Mostly Crystalline Mush NORMAN H. SLEEP NORMAN H. SLEEP 1Department of Earth and Planetary Sciences, Massachusetts Institute of Technology, Cambridge, Massachusetts 02139 Search for other works by this author on: GSW Google Scholar GSA Bulletin (1974) 85 (8): 1225–1232. https://doi.org/10.1130/0016-7606(1974)85<1225:SOMFAM>2.0.CO;2 Article history first online: 01 Jun 2017 Cite View This Citation Add to Citation Manager Share Icon Share Facebook Twitter LinkedIn MailTo Tools Icon Tools Get Permissions Search Site Citation NORMAN H. SLEEP; Segregation of Magma from a Mostly Crystalline Mush. GSA Bulletin 1974;; 85 (8): 1225–1232. doi: https://doi.org/10.1130/0016-7606(1974)85<1225:SOMFAM>2.0.CO;2 Download citation file: Ris (Zotero) Refmanager EasyBib Bookends Mendeley Papers EndNote RefWorks BibTex toolbar search Search Dropdown Menu toolbar search search input Search input auto suggest filter your search All ContentBy SocietyGSA Bulletin Search Advanced Search Abstract A model based on the fluid dynamics of intermixed fluids indicated that favorable conditions for segregation of melt from a mostly crystalline mush upwelling from the asthenosphere included a high concentration of melt, a narrow conduit containing the mush, a large grain size, and a large ratio of the viscosity of grains to the viscosity of melt. The greater depth, inferred geochemically, of the source regions of seamount lavas, compared with mid-oceanic ridge lavas, is attributable to the dependence of segregation on conduit width. For thermal reasons, conduit width is smaller at greater depths for small sources of material such as seamounts. Observed systematic enrichment of seamount and island-arc lavas in radiogenic isotopics (which, for chemical reasons, would be most abundant in domains containing high fractions of water) relative to mid-oceanic ridge lavas may be due to disproportionate representation of ubiquitous, small-scale, water-rich domains in the source region of seamount lavas. The fraction of melt is strongly dependent on water content in the deep source regions of seamount lavas but only weakly dependent in the shallow source regions of mid-oceanic ridge lavas. This also complicates identification of subducted sediments in island-arc lavas. The viscosity of partial melt in the asthenosphere, inferred from seismic attenuation studies, is too high for that melt to segregate efficiently. First Page Preview Close Modal You do not have access to this content, please speak to your institutional administrator if you feel you should have access.
暂无摘要(点击查看原文获取完整内容)
BACKGROUND: Interference competition occurs when access to resources is negatively affected by the presence of other individuals. Within a species or population, this is known as mutual interference, and it is often modelled with a scaling exponent, m, on the number of predators. Originally, mutual interference was thought to vary along a continuum from prey dependence (no interference; m = 0) to ratio dependence (m = -1), but a debate in the 1990's and early 2000's focused on whether prey or ratio dependence was the better simplification. Some have argued more recently that mutual interference is likely to be mostly intermediate (that is, between prey and ratio dependence), but this possibility has not been evaluated empirically. RESULTS: We gathered estimates of mutual interference from the literature, analyzed additional data, and created the largest compilation of unbiased estimates of mutual interference yet produced. In this data set, both the alternatives of prey dependence and ratio dependence were observed, but only one data set was consistent with prey dependence. There was a tendency toward ratio dependence reflected by a median m of -0.7 and a mean m of -0.8. CONCLUSIONS: Overall, the data support the hypothesis that interference is mostly intermediate in magnitude. The data also indicate that interference competition is common, at least in the systems studied to date. Significant questions remain regarding how different factors influence interference, and whether interference can be viewed as a characteristic of a particular population or whether it generally shifts from low to high levels as populations increase in density.
A growing number of young men today say they are “mostly straight” and yet feel a slight but enduring desire for men. Ritch Savin-Williams explores the stories of 40 mostly straight young men to help us understand the biological, psychological, and cultural forces that are loosening the sexual bind many boys and young men experience.
暂无摘要(点击查看原文获取完整内容)
Complement receptor 2-negative (CR2/CD21(-)) B cells have been found enriched in patients with autoimmune diseases and in common variable immunodeficiency (CVID) patients who are prone to autoimmunity. However, the physiology of CD21(-/lo) B cells remains poorly characterized. We found that some rheumatoid arthritis (RA) patients also display an increased frequency of CD21(-/lo) B cells in their blood. A majority of CD21(-/lo) B cells from RA and CVID patients expressed germline autoreactive antibodies, which recognized nuclear and cytoplasmic structures. In addition, these B cells were unable to induce calcium flux, become activated, or proliferate in response to B-cell receptor and/or CD40 triggering, suggesting that these autoreactive B cells may be anergic. Moreover, gene array analyses of CD21(-/lo) B cells revealed molecules specifically expressed in these B cells and that are likely to induce their unresponsive stage. Thus, CD21(-/lo) B cells contain mostly autoreactive unresponsive clones, which express a specific set of molecules that may represent new biomarkers to identify anergic B cells in humans.
Abstract High-resolution observations and regional climate model simulations reveal that precipitation over the Maritime Continent is mostly concentrated over islands. Analysis of the diurnal cycles of precipitation and winds indicates that this is predominantly caused by sea-breeze convergence over islands, reinforced by mountain–valley winds and further amplified by the cumulus merger processes. Comparison of a regional climate model control simulation to a flat-island run and an all-ocean run demonstrates that the underrepresentation of islands and terrain in the Maritime Continent weakens the atmospheric disturbance associated with the diurnal cycle, and hence underestimates precipitation. The implication of these regional modeling results is that systematic errors in coarse-resolution global circulation models probably result from insufficient representation of land–sea breezes associated with the complex topography in the Maritime Continent. It is found that precipitation in the Maritime Continent, simulated by a global model, is indeed smaller than observed. The simulated upper-atmospheric velocity potential, which represents large-scale tropospheric heating, was substantially displaced eastward compared to observations. Possible approaches toward solving this problem are suggested.
暂无摘要(点击查看原文获取完整内容)
This theoretical article discusses the emerging concept of awareness of age-related change (AARC). We propose that a focus on AARC extends the research traditions on subjective age experiences and age identity and that examination of this concept can serve a stimulating role in social gerontology. After defining and contrasting AARC against similar concepts, several reasons for the relevance of this mostly unexplored construct are provided. The sample domains of health and physical functioning, cognitive functioning, and interpersonal relations are used to illustrate the relevance of AARC. Based on this review, we then provide a heuristic framework that describes antecedents, processes, and outcomes related to AARC. Overall, we argue that research on AARC should become an integral part of social gerontological research.
Male-typed leadership schemas have been widely acknowledged as barriers to women’s success in leadership roles. We explore how local organizational agents and contexts enable women leaders to overcome these barriers and achieve success at the highest levels in firms. Specifically, we focus on chief executive officer (CEO) succession events and study how several facets of predecessor CEOs and the succession context combine to influence incoming women’s post-succession performance. We conduct a qualitative comparative case study of all CEO successions that involved female successors between 1989 and 2009 across the largest corporations in the United States. Our findings suggest that women’s success occurred when a confluence of local firm-level factors and attributes of the (mostly) male predecessors promoted gender-inclusive gatekeeping during succession. Our qualitative comparative analysis approach reveal three recipes for female success: “handing over the legacy,” “partnering the legacy,” and “turning around the legacy.” Moreover, a comparison to a matched-sample of men CEO succession events shows that these three recipes for success are unique to women. Based upon our findings, we propose that male predecessors’ gender-inclusive gatekeeping facilitates female leaders’ success and occurs when local enabling conditions and the embedded context enact agentic and structural mechanisms to alter leadership schemas.
The advent of Web 2.0 has lead to the proliferation of client-side code that is typically written in JavaScript. This code is often combined — or mashed-up — with other code and content from disparate, mutually untrusting parties, leading to undesirable security and reliability consequences. This paper proposes GATEKEEPER, a mostly static approach for soundly enforcing security and reliability policies for JavaScript programs. GATEKEEPER is a highly extensible system with a rich, expressive policy language, allowing the hosting site administrator to formulate their policies as succinct Datalog queries. The primary application of GATEKEEPER this paper explores is in reasoning about JavaScript widgets such as those hosted by widget portals Live.com and Google/IG. Widgets submitted to these sites can be either malicious or just buggy and poorly written, and the hosting site has the authority to reject the submission of widgets that do not meet the site’s security policies. To show the practicality of our approach, we describe nine representative security and reliability policies. Statically checking these policies results in 1,341 verified warnings in 684 widgets, no false negatives, due to the soundness of our analysis, and false positives affecting only two widgets. 1
We present a method for adapting garbage collectors designed to run sequentially with the client, so that they may run concurrently with it. We rely on virtual memory hardware to provide information about pages that have been updated or "dirtied" during a given period of time. This method has been used to construct a mostly parallel trace-and-sweep collector that exhibits very short pause times. Performance measurements are given.
This paper presents a reconfigurable continuous-time delta-sigma modulator for analog-to-digital conversion that consists mostly of digital circuitry. It is a voltage-controlled ring oscillator based design with new digital background calibration and self-cancelling dither techniques applied to enhance performance. Unlike conventional delta-sigma modulators, it does not contain analog integrators, feedback DACs, comparators, or reference voltages, and does not require a low-jitter clock. Therefore, it uses less area than comparable conventional delta-sigma modulators, and the architecture is well-suited to IC processes optimized for fast digital circuitry. The prototype IC is implemented in 65 nm LP CMOS technology with power dissipation, output sample-rate, bandwidth, and peak SNDR ranges of 8-17 mW, 0.5-1.15 GHz, 3.9-18 MHz, and 67-78 dB, respectively, and an active area of 0.07.
Carbon fluxes in subduction zones can be better constrained by including new estimates of carbon concentration in subducting mantle peridotites, consideration of carbonate solubility in aqueous fluid along subduction geotherms, and diapirism of carbon-bearing metasediments. Whereas previous studies concluded that about half the subducting carbon is returned to the convecting mantle, we find that relatively little carbon may be recycled. If so, input from subduction zones into the overlying plate is larger than output from arc volcanoes plus diffuse venting, and substantial quantities of carbon are stored in the mantle lithosphere and crust. Also, if the subduction zone carbon cycle is nearly closed on time scales of 5-10 Ma, then the carbon content of the mantle lithosphere + crust + ocean + atmosphere must be increasing. Such an increase is consistent with inferences from noble gas data. Carbon in diamonds, which may have been recycled into the convecting mantle, is a small fraction of the global carbon inventory.
The ketocarotenoid astaxanthin can be found in the microalgae Haematococcus pluvialis, Chlorella zofingiensis, and Chlorococcum sp., and the red yeast Phaffia rhodozyma. The microalga H. pluvialis has the highest capacity to accumulate astaxanthin up to 4-5% of cell dry weight. Astaxanthin has been attributed with extraordinary potential for protecting the organism against a wide range of diseases, and has considerable potential and promising applications in human health. Numerous studies have shown that astaxanthin has potential health-promoting effects in the prevention and treatment of various diseases, such as cancers, chronic inflammatory diseases, metabolic syndrome, diabetes, diabetic nephropathy, cardiovascular diseases, gastrointestinal diseases, liver diseases, neurodegenerative diseases, eye diseases, skin diseases, exercise-induced fatigue, male infertility, and HgCl₂-induced acute renal failure. In this article, the currently available scientific literature regarding the most significant activities of astaxanthin is reviewed.
Sensor networks are often desired to last many times longer than the active lifetime of individual sensors. This is usually achieved by putting sensors to sleep for most of their lifetime. On the other hand, surveillance kind of applications require guaranteed k-coverage of the protected region at all times. As a result, determining the appropriate number of sensors to deploy that achieves both goals simultaneously becomes a challenging problem. In this paper, we consider three kinds of deployments for a sensor network on a unit square - a √n x √n grid, random uniform (for all n points), and Poisson (with density n). In all three deployments, each sensor is active with probability p, independently from the others. Then, we claim that the critical value of the function npπr2/log(np) is 1 for the event of k-coverage of every point. We also provide an upper bound on the window of this phase transition. Although the conditions for the three deployments are similar, we obtain sharper bounds for the random deployments than the grid deployment, which occurs due to the boundary condition. In this paper, we also provide corrections to previously published results for the grid deployment model. Finally, we use simulation to show the usefulness of our analysis in real deployment scenarios.