Using the Scopus dataset (1996-2007) a grand matrix of aggregated journal-journal citations was constructed. This matrix can be compared in terms of the network structures with the matrix contained in the Journal Citation Reports (JCR) of the Institute of Scientific Information (ISI). Since the Scopus database contains a larger number of journals and covers also the humanities, one would expect richer maps. However, the matrix is in this case sparser than in the case of the ISI data. This is due to (i) the larger number of journals covered by Scopus and (ii) the historical record of citations older than ten years contained in the ISI database. When the data is highly structured, as in the case of large journals, the maps are comparable, although one may have to vary a threshold (because of the differences in densities). In the case of interdisciplinary journals and journals in the social sciences and humanities, the new database does not add a lot to what is possible with the ISI databases.
Data science has become increasingly essential for the production of official statistics, as it enables the automated collection, processing, and analysis of large amounts of data. With such data science practices in place, it enables more timely, more insightful and more flexible reporting. However, the quality and integrity of data-science-driven statistics rely on the accuracy and reliability of the data sources and the machine learning techniques that support them. In particular, changes in data sources are inevitable to occur and pose significant risks that are crucial to address in the context of machine learning for official statistics. This paper gives an overview of the main risks, liabilities, and uncertainties associated with changing data sources in the context of machine learning for official statistics. We provide a checklist of the most prevalent origins and causes of changing data sources; not only on a technical level but also regarding ownership, ethics, regulation, and public perception. Next, we highlight the repercussions of changing data sources on statistical reporting. These include technical effects such as concept drift, bias, availability, validity, accur
Gynandromorphs are creatures where at least two different body sections are a different sex. Bilateral gynandromorphs are half male and half female. Here we develop a theory of gynandromorph ontogeny based on developmental control networks. The theory explains the embryogenesis of all known variations of gynandromorphs found in multicellular organisms. The theory also predicts a large variety of more subtle gynandromorphic morphologies yet to be discovered. The network theory of gynandromorph development has direct relevance to understanding sexual dimorphism (differences in morphology between male and female organisms of the same species) and medical pathologies such as hemihyperplasia (asymmetric development of normally symmetric body parts in a unisexual individual). The network theory of gynandromorphs brings up fundamental open questions about developmental control in ontogeny. This in turn suggests a new theory of the origin and evolution of species that is based on cooperative interactions and conflicts between developmental control networks in the haploid genomes and epigenomes of potential sexual partners for reproduction. This network-based theory of the origin of species
We compare the network of aggregated journal-journal citation relations provided by the Journal Citation Reports (JCR) 2012 of the Science and Social Science Citation Indexes (SCI and SSCI) with similar data based on Scopus 2012. First, global maps were developed for the two sets separately; sets of documents can then be compared using overlays to both maps. Using fuzzy-string matching and ISSN numbers, we were able to match 10,524 journal names between the two sets; that is, 96.4% of the 10,936 journals contained in JCR or 51.2% of the 20,554 journals covered by Scopus. Network analysis was then pursued on the set of journals shared between the two databases and the two sets of unique journals. Citations among the shared journals are more comprehensively covered in JCR than Scopus, so the network in JCR is denser and more connected than in Scopus. The ranking of shared journals in terms of indegree (that is, numbers of citing journals) or total citations is similar in both databases overall (Spearman's \r{ho} > 0.97), but some individual journals rank very differently. Journals that are unique to Scopus seem to be less important--they are citing shared journals rather than bein
Using three years of the Journal Citation Reports (2011, 2012, and 2013), indicators of transitions in 2012 (between 2011 and 2013) are studied using methodologies based on entropy statistics. Changes can be indicated at the level of journals using the margin totals of entropy production along the row or column vectors, but also at the level of links among journals by importing the transition matrices into network analysis and visualization programs (and using community-finding algorithms). Seventy-four journals are flagged in terms of discontinuous changes in their citations; but 3,114 journals are involved in "hot" links. Most of these links are embedded in a main component; 78 clusters (containing 172 journals) are flagged as potential "hot spots" emerging at the network level. An additional finding is that PLoS ONE introduced a new communication dynamics into the database. The limitations of the methodology are elaborated using an example. The results of the study indicate where developments in the citation dynamics can be considered as significantly unexpected. This can be used as heuristic information; but what a "hot spot" in terms of the entropy statistics of aggregated cit
Rankings of scholarly journals based on citation data are often met with skepticism by the scientific community. Part of the skepticism is due to disparity between the common perception of journals' prestige and their ranking based on citation counts. A more serious concern is the inappropriate use of journal rankings to evaluate the scientific influence of authors. This paper focuses on analysis of the table of cross-citations among a selection of Statistics journals. Data are collected from the Web of Science database published by Thomson Reuters. Our results suggest that modelling the exchange of citations between journals is useful to highlight the most prestigious journals, but also that journal citation data are characterized by considerable heterogeneity, which needs to be properly summarized. Inferential conclusions require care in order to avoid potential over-interpretation of insignificant differences between journal ratings. Comparison with published ratings of institutions from the UK's Research Assessment Exercise shows strong correlation at aggregate level between assessed research quality and journal citation `export scores' within the discipline of Statistics.
Using "Analyze Results" at the Web of Science, one can directly generate overlays onto global journal maps of science. The maps are based on the 10,000+ journals contained in the Journal Citation Reports (JCR) of the Science and Social Science Citation Indices (2011). The disciplinary diversity of the retrieval is measured in terms of Rao-Stirling's "quadratic entropy." Since this indicator of interdisciplinarity is normalized between zero and one, the interdisciplinarity can be compared among document sets and across years, cited or citing. The colors used for the overlays are based on Blondel et al.'s (2008) community-finding algorithms operating on the relations journals included in JCRs. The results can be exported from VOSViewer with different options such as proportional labels, heat maps, or cluster density maps. The maps can also be web-started and/or animated (e.g., using PowerPoint). The "citing" dimension of the aggregated journal-journal citation matrix was found to provide a more comprehensive description than the matrix based on the cited archive. The relations between local and global maps and their different functions in studying the sciences in terms of journal lit
After the repeal of Roe vs. Wade in June 2022, women face long-distance travel across state lines to access abortion care. For women who also face socioeconomic hardship, travel for abortion care is a significant burden. To ease this burden, abortion access nonprofits are funding and/or supplying transportation to abortion clinics. However, due to the uneven distribution of demand and supply for abortions, these nonprofits do not have efficient logistical operations. As a result, low-income, underserved women may not have access to adequate reproductive healthcare, thus widening healthcare inequity gaps. Nonprofits may also risk not serving the needs of vulnerable women without access to adequate reproductive healthcare, and in doing so, waste resources, money, and volunteer hours. To address these challenges, we create an interactive, web-based planning tool, the Reproductive Healthcare Equity Algorithm (RHEA), to guide nonprofits in strategically allocating resources and serving demand. RHEA leverages an optimization model to determine the maximum flow and minimum transportation cost to route women across a network of counties and abortion clinics, subject to transportation suppl
Background: In 2015, the Zika arbovirus (ZIKV) began circulating in the Americas, rapidly expanding its global geographic range in explosive outbreaks. Unusual among mosquito-borne diseases, ZIKV has been shown to also be sexually transmitted, although sustained autochthonous transmission due to sexual transmission alone has not been observed, indicating the reproduction number (R0) for sexual transmission alone is less than 1. Critical to the assessment of outbreak risk, estimation of the potential attack rates, and assessment of control measures, are estimates of the basic reproduction number, R0. Methods: We estimated the R0 of the 2015 ZIKV outbreak in Barranquilla, Colombia, through an analysis of the exponential rise in clinically identified ZIKV cases (n = 359 to the end of November, 2015). Findings: The rate of exponential rise in cases was rho=0.076 days-1, with 95 percent CI [0.066,0.087] days-1. We used a vector-borne disease model with additional direct transmission to estimate the R0; assuming the R0 of sexual transmission alone is less than 1, we estimated the total R0 = 3.8 [2.4,5.6], and that the fraction of cases due to sexual transmission was 0.23 [0.01,0.47] with
Government actions, such as the Medina v. Planned Parenthood South Atlantic Supreme Court ruling and the passage of the Big Beautiful Bill Act, have aimed to restrict or prohibit Medicaid funding for Planned Parenthood Healthcare Centers (PPHCs) at both the state and national levels. These funding cuts are particularly harmful in states like California, which has a large population of Medicaid users. This analysis focuses on the distribution of Planned Parenthood clinics and Federally Qualified Health Centers (FQHCs), which offer essential reproductive healthcare services including, but not limited to, abortions, birth control, HIV services, pregnancy testing and planning, STD testing and treatment, and cancer screenings. While expanded funding for FQHCs has been proposed as a solution, it fails to address the locational accessibility of Medicaid-funded health centers that provide sexual and reproductive care. To assess this issue, we analyze the proximity of data points representing California's PPHC and FQHC locations. Topological Data Analysis (TDA)-an approach that examines the shape and structure of data -- is used to detect disparities in reproductive and sexual healthcare co
The question as to why most higher organisms reproduce sexually has remained open despite extensive research, and has been called "the queen of problems in evolutionary biology". Theories dating back to Weismann have suggested that the key must lie in the creation of increased variability in offspring, causing enhanced response to selection. Rigorously quantifying the effects of assorted mechanisms which might lead to such increased variability, and establishing that these beneficial effects outweigh the immediate costs of sexual reproduction has, however, proved problematic. Here we introduce an approach which does not focus on particular mechanisms influencing factors such as the fixation of beneficial mutants or the ability of populations to deal with deleterious mutations, but rather tracks the entire distribution of a population of genotypes as it moves across vast fitness landscapes. In this setting simulations now show sex robustly outperforming asex across a broad spectrum of finite or infinite population models. Concentrating on the additive infinite populations model, we are able to give a rigorous mathematical proof establishing that sexual reproduction acts as a more ef
A number of journal classification systems have been developed in bibliometrics since the launch of the Citation Indices by the Institute of Scientific Information (ISI) in the 1960s. These systems are used to normalize citation counts with respect to field-specific citation patterns. The best known system is the so-called "Web-of-Science Subject Categories" (WCs). In other systems papers are classified by algorithmic solutions. Using the Journal Citation Reports 2014 of the Science Citation Index and the Social Science Citation Index (n of journals = 11,149), we examine options for developing a new system based on journal classifications into subject categories using aggregated journal-journal citation data. Combining routines in VOSviewer and Pajek, a tree-like classification is developed. At each level one can generate a map of science for all the journals subsumed under a category. Nine major fields are distinguished at the top level. Further decomposition of the social sciences is pursued for the sake of example with a focus on journals in information science (LIS) and science studies (STS). The new classification system improves on alternative options by avoiding the problem
A previous study of symmetric collisions of massive nuclei has shown that current models of multi-nucleon transfer (MNT) reactions do not adequately describe the transfer product yields. To gain further insight into this problem, we have measured the yields of MNT products in the interaction of 977 (E/A = 4.79 MeV) and 1143 MeV (E/A = 5.60 MeV) $^{204}$Hg with $^{208}$Pb. We find that the yield of multi-nucleon transfer products are similar in these two reactions and are substantially lower than those observed in the reaction of 1257 MeV (E/A = 6.16 MeV) $^{204}$Hg + $^{198}$Pt. We compare our measurements with the predictions of the GRAZING-F, di-nuclear systems (DNS) and improved quantum molecular dynamics (ImQMD) models. For the observed isotopes of the elements Au, Hg, Tl, Pb and Bi, the measured values of the MNT cross sections are orders of magnitude larger than the predicted values. Furthermore, the various models predict the formation of nuclides near the N=126 shell, which are not observed.
National statistical institutes currently investigate how to improve the output quality of official statistics based on machine learning algorithms. A key obstacle is concept drift, i.e., when the joint distribution of independent variables and a dependent (categorical) variable changes over time. Under concept drift, a statistical model requires regular updating to prevent it from becoming biased. However, updating a model asks for additional data, which are not always available. In the literature, we find a variety of bias correction methods as a promising solution. In the paper, we will compare two popular correction methods: the misclassification estimator and the calibration estimator. For prior probability shift (a specific type of concept drift), we investigate the two correction methods theoretically as well as experimentally. Our theoretical results are expressions for the bias and variance of both methods. As experimental result, we present a decision boundary (as a function of (a) model accuracy, (b) class distribution and (c) test set size) for the relative performance of the two methods. Close inspection of the results will provide a deep insight into the effect of pri
Phosphorus (P) is considered to be one of the key elements for life, making it an important element to look for in the abundance analysis of spectra of stellar systems. Yet, there exists only a handful of spectroscopic studies to estimate the P abundances and investigate its trend across a range of metallicities. We have observed full HK band spectra at a spectral resolving power of R=45,000 with IGRINS instrument. Abundances are determined using SME in combination with 1D MARCS stellar atmosphere models. The investigated sample of stars have reliable stellar parameters estimated using optical FIES spectra (GILD; Jönsson et al. in prep.). In order to determine the P abundances from the 16482.92 Angstrom P line, we take special care of the CO($ν=7-4$) blend. We determine the C, N, O abundances from atomic carbon and a range of non-blended molecular lines (CO, CN, OH) which are aplenty in the H band region of K giant stars, assuring an appropriate modelling of the blending CO($ν=7-4$) line. We present [P/Fe] vs [Fe/H] trend for 38 K giant stars in the metallicity range of -1.2 dex $<$ [Fe/H] $<$ 0.4 dex. We find that our trend matches well with the compiled literature sample of
Publication patterns of 79 forest scientists awarded major international forestry prizes during 1990-2010 were compared with the journal classification and ranking promoted as part of the 'Excellence in Research for Australia' (ERA) by the Australian Research Council. The data revealed that these scientists exhibited an elite publication performance during the decade before and two decades following their first major award. An analysis of their 1703 articles in 431 journals revealed substantial differences between the journal choices of these elite scientists and the ERA classification and ranking of journals. Implications from these findings are that additional cross-classifications should be added for many journals, and there should be an adjustment to the ranking of several journals relevant to the ERA Field of Research classified as 0705 Forestry Sciences.
Mobile phone data are an interesting new data source for official statistics. However, multiple problems and uncertainties need to be solved before these data can inform, support or even become an integral part of statistical production processes. In this paper, we focus on arguably the most important problem hindering the application of mobile phone data in official statistics: detecting home locations. We argue that current efforts to detect home locations suffer from a blind deployment of criteria to define a place of residence and from limited validation possibilities. We support our argument by analysing the performance of five home detection algorithms (HDAs) that have been applied to a large, French, Call Detailed Record (CDR) dataset (~18 million users, 5 months). Our results show that criteria choice in HDAs influences the detection of home locations for up to about 40% of users, that HDAs perform poorly when compared with a validation dataset (the 35°-gap), and that their performance is sensitive to the time period and the duration of observation. Based on our findings and experiences, we offer several recommendations for official statistics. If adopted, our recommendatio
From medical charts to national census, healthcare has traditionally operated under a paper-based paradigm. However, the past decade has marked a long and arduous transformation bringing healthcare into the digital age. Ranging from electronic health records, to digitized imaging and laboratory reports, to public health datasets, today, healthcare now generates an incredible amount of digital information. Such a wealth of data presents an exciting opportunity for integrated machine learning solutions to address problems across multiple facets of healthcare practice and administration. Unfortunately, the ability to derive accurate and informative insights requires more than the ability to execute machine learning models. Rather, a deeper understanding of the data on which the models are run is imperative for their success. While a significant effort has been undertaken to develop models able to process the volume of data obtained during the analysis of millions of digitalized patient records, it is important to remember that volume represents only one aspect of the data. In fact, drawing on data from an increasingly diverse set of sources, healthcare data presents an incredibly comp
This paper develops mathematical models describing the evolutionary dynamics of both asexually and sexually reproducing populations of diploid unicellular organisms. We consider two forms of genome organization. In one case, we assume that the genome consists of two multi-gene chromosomes, while in the second case we assume that each gene defines a separate chromosome. If the organism has $ l $ homologous pairs that lack a functional copy of the given gene, then the fitness of the organism is $ κ_l $. The $ κ_l $ are assumed to be monotonically decreasing, so that $ κ_0 = 1 > κ_1 > κ_2 > ... > κ_{\infty} = 0 $. For nearly all of the reproduction strategies we consider, we find, in the limit of large $ N $, that the mean fitness at mutation-selection balance is $ \max\{2 e^{-μ} - 1, 0\} $, where $ N $ is the number of genes in the haploid set of the genome, $ ε$ is the probability that a given DNA template strand of a given gene produces a mutated daughter during replication, and $ μ= N ε$. The only exception is the sexual reproduction pathway for the multi-chromosomed genome. Assuming a multiplicative fitness landscape where $ κ_l = α^{l} $ for $ α\in (0, 1) $, this str
Healthcare is one of the most promising areas for machine learning models to make a positive impact. However, successful adoption of AI-based systems in healthcare depends on engaging and educating stakeholders from diverse backgrounds about the development process of AI models. We present a broadly accessible overview of the development life cycle of clinical AI models that is general enough to be adapted to most machine learning projects, and then give an in-depth case study of the development process of a deep learning based system to detect aortic aneurysms in Computed Tomography (CT) exams. We hope other healthcare institutions and clinical practitioners find the insights we share about the development process useful in informing their own model development efforts and to increase the likelihood of successful deployment and integration of AI in healthcare.