Sparse Mixtures of Experts (MoEs) are typically trained to operate at a fixed sparsity level, e.g. $k$ in a top-$k$ gating function. This global sparsity level determines an operating point on the accuracy/latency curve; currently, meeting multiple efficiency targets means training and maintaining multiple models. This practice complicates serving, increases training and maintenance costs, and limits flexibility in meeting diverse latency, efficiency, and energy requirements. We show that pretrained MoEs are more robust to runtime sparsity shifts than commonly assumed, and introduce MoE-PHDS ({\bf P}ost {\bf H}oc {\bf D}eclared {\bf S}parsity), a lightweight SFT method that turns a single checkpoint into a global sparsity control surface. PHDS mixes training across sparsity levels and anchors with a short curriculum at high sparsity, requiring no architectural changes. The result is predictable accuracy/latency tradeoffs from one model: practitioners can ``dial $k$'' at inference time without swapping checkpoints, changing architecture, or relying on token-level heuristics. Experiments on OLMoE-1B-7B-0125, Qwen1.5-MoE-A2.7B, and proprietary models fit on multiple operating points s
To sustain innovation and safeguard national security, the U.S. must strengthen domestic pathways to computing PhDs by engaging talented undergraduates early - before they are committed to industry - with research experiences, mentorship, and financial support for graduate studies.
We have recently compiled a database with all doctoral dissertations (PhDs) completed in modern Greece (1837-2014), in the general area of astronomy and astrophysics, as well as in space and ionospheric physics. A preliminary statistical analysis of the data is presented, along with a discussion of the general trends observed.
Statistical and probabilistic reasoning enlightens our judgments about uncertainty and the chance or beliefs on the occurrence of random events in everyday life. Therefore, there are scientists working with Probability and Statistics in various fields of knowledge, what favors the formation of scientific network collaborations of researchers with different backgrounds. Here, we propose to describe the Brazilian PhDs who work with probability and statistics. In particular, we analyze national and states collaboration networks of such researchers by calculating different metrics. We show that there is a greater concentration of nodes in and around the cites which host Probability and Statistics graduate programs. Moreover, the states that host P & S Doctoral programs are the most central. We also observe a disparity in the size of the states networks. The clustering coefficient of the national network suggests that this network and regional differences especially with respect to states from South-east and North is not cohesive and, probably, it is in a maturing stage
This contribution presents a poll undertaken at the beginning of 2012, and addressed to every doctor in astronomy who obtained his/her degree in France. Its goal is to motivate the French astronomical community to think and discuss about what should be the training of PhDs, and what should be its objective. Further discussions and reactions can be posted e.g. on http://docastro.blogspot.fr/. A worrying results from the poll is that the majority of the participants would not encourage a young student to start a thesis in astronomy. The main reasons for this fact may be the high pressure on astronomy positions and the little interest a doctorate has for other careers in France. I suggest we either have to modify our formations or reduce the number of thesis starting each year in astronomy.
In this invited Editorial for Software and Computing for Big Science, we describe the SMARTHEP Innovative Training Network funded via the Marie Skłodowska-Curie Actions between 2021 and 2025. SMARTHEP trained 12 PhD students to advance machine learning and real-time analysis in high-energy physics experiments and industrial applications. We present the perspective of students, supervisors, and external observers of the network, concerning the work done within the network, the added value compared to ``typical'' PhD positions, and the emerging themes and directions from our experiences in the past four years.
The malicious usage of large language models (LLMs) has motivated the detection of LLM-generated texts. Previous work in topological data analysis shows that the persistent homology dimension (PHD) of text embeddings can serve as a more robust and promising score than other zero-shot methods. However, effectively detecting short LLM-generated texts remains a challenge. This paper presents Short-PHD, a zero-shot LLM-generated text detection method tailored for short texts. Short-PHD stabilizes the estimation of the previous PHD method for short texts by inserting off-topic content before the given input text and identifies LLM-generated text based on an established detection threshold. Experimental results on both public and generated datasets demonstrate that Short-PHD outperforms existing zero-shot methods in short LLM-generated text detection. Implementation codes are available online.
Computer science attracts few women, and their proportion decreases through advancing career stages. Few women progress to PhD studies in CS after completing master's studies. Empowering women at this stage in their careers is essential to unlock untapped potential for society, industry and academia. This paper identifies students' career assumptions and information related to PhD studies focused on gender-based differences. We propose a Women Career Lunch program to inform female master students about PhD studies that explains the process, clarifies misconceptions, and alleviates concerns. An extensive survey was conducted to identify factors that encourage and discourage students from undertaking PhD studies. We identified statistically significant differences between those who undertook PhD studies and those who didn't, as well as gender differences. A catalogue of questions to initiate discussions with potential PhD students which allowed them to explore these factors was developed and translated to 8 languages. Encouraging factors toward PhD study include interest and confidence in research arising from a research involvement during earlier studies; enthusiasm for and self-con
Visualization is a heterogeneous field, and this aspect is often reflected by the organizational structures at higher education institutions that academic researchers in visualization and related fields including computer graphics, human-computer interaction, and media design are typically affiliated with. It may thus be a challenge for new PhD students to grasp the fragmented structure of their new workplace, form collegial relations across the institution, and to build a coherent picture of the discipline as a whole. We report an attempt to address this challenge, in the form of an introductory course on the subject of Visualization Technology and Methodology for PhD students at the Division for Media and Information Technology, Linköping University, Sweden. We discuss the course design, including interactions with other doctoral education activities and field trips to multiple research groups and units within the division (ranging from scientific visualization and computer graphics to media design and visual communication). Lessons learned from the course preparation work as well as the first instance of the course offered during autumn term 2023 can be helpful to researchers an
We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-ef
Spatial transcriptomics (ST) measures gene expression at a set of spatial locations in a tissue. Communities of nearby cells that express similar genes form \textit{spatial domains}. Specialized ST clustering algorithms have been developed to identify these spatial domains. These methods often identify spatial domains at a single morphological scale, and interactions across multiple scales are often overlooked. For example, large cellular communities often contain smaller substructures, and heterogeneous frontier regions often lie between homogeneous domains. Topological data analysis (TDA) is an emerging mathematical toolkit that studies the underlying features of data at various geometric scales. It is especially useful for analyzing complex biological datasets with multiscale characteristics. Using TDA, we develop Persistent Homology for Domains at Multiple Scales (PHD-MS) to locate tissue structures that persist across morphological scales. We apply PHD-MS to highlight multiscale spatial domains in several tissue types and ST technologies. We also compare PHD-MS domains against ground-truth domains in expert-annotated tissues, where PHD-MS outperforms traditional clustering app
The FAIR (Findable, Accessible, Interoperable, and Reusable) data principles have gained significant attention as a means to enhance data sharing, collaboration, and reuse across various domains. Here, we explore the potential benefits of implementing FAIR data practices within engineering projects, with a monetary focus in the German context, but by considering aspects which are relatively universal. By examining the FAIR-data aspect of a Materials Science and Engineering PhD project, it becomes evident that substantial cost savings can be achieved. The estimated savings are 2,600 Euros per year from the PhD project considered. This study underscores the importance of implementing FAIR data practices in engineering projects and highlights some significant economic benefits that can be derived from such initiatives. By embracing FAIR principles, organizations in the engineering sector can unlock the full potential of their data, optimize resource allocation, and drive innovation in a cost-effective manner.
The probability hypothesis density (PHD) and Poisson multi-Bernoulli (PMB) filters are two popular set-type multi-object filters. Motivated by the fact that the multi-object filtering density after each update step in the PHD filter is a PMB without approximation, in this paper we present a multi-object smoother involving PHD forward filtering and PMB backward smoothing. This is achieved by first running the PHD filtering recursion in the forward pass and extracting the PMB filtering densities after each update step before the Poisson Point Process approximation, which is inherent in the PHD filter update. Then in the backward pass we apply backward simulation for sets of trajectories to the extracted PMB filtering densities. We call the resulting multi-object smoother hybrid PHD-PMB trajectory smoother. Notably, the hybrid PHD-PMB trajectory smoother can provide smoothed trajectory estimates for the PHD filter without labeling or tagging, which is not possible for existing PHD smoothers. Also, compared to the trajectory PHD filter, which can only estimate alive trajectories, the hybrid PHD-PMB trajectory smoother enables the estimation of the set of all trajectories. Simulation re
Studying the factors that influence the quality of physics PhD students' doctoral experiences, especially those that motivate them to stay or leave their programs, is critical for providing them with more holistic and equitable support. Prior literature on doctoral attrition has found that students with clear research interests who establish an advisor-advisee relationship early in their graduate careers are most likely to persist. However, these trends have not been investigated in the context of physics, and the underlying reasons for why these characteristics are associated with leaving remain unstudied. Using semi-structured interviews with 40 first and second year physics PhD students, we construct a model describing the characteristic pathways that physics PhD students take while evaluating interest congruence of prospective research groups. We show how access to undergraduate research and other formative experiences helped some students narrow their interests and look for research groups before arriving to graduate school. In turn, these students reported fewer difficulties finding a group than students whose search for an advisor took place during the first year of their Ph
Multimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE). This paper contributes a ChatGPT-Prompted visual hallucination evaluation Dataset (PhD) for objective VHE at a large scale. The essence of VHE is to ask an MLLM questions about specific images to assess its susceptibility to hallucination. Depending on what to ask (objects, attributes, sentiment, etc.) and how the questions are asked, we structure PhD along two dimensions, i.e. task and mode. Five visual recognition tasks, ranging from low-level (object / attribute recognition) to middle-level (sentiment / position recognition and counting), are considered. Besides a normal visual QA mode, which we term PhD-base, PhD also asks questions with specious context (PhD-sec) or with incorrect context ({PhD-icc), or with AI-generated counter common sense images (PhD-ccs). We construct PhD by a ChatGPT-assisted semi-automated pipeline, encompassing four pivotal modules: task-specific hallucinatory item (hitem) selection, hitem-embedded question generation, specious / incorrect context generation, and counter-common-sense (CCS) image generation. With over 14k daily i
We provide an example of the application of quantitative techniques, tools, and topics from mathematics and data science to analyze the mathematics community itself in order to quantify and document inequity in our discipline. This work is a contribution to the new and growing interdisciplinary field recently termed "mathematics of Mathematics," or "MetaMath." Using data about PhD-granting institutions in the United States and publicly available funding data from the National Science Foundation, we highlight inequalities in departments at U.S. institutions of higher education that produce PhDs in the mathematical sciences. Specifically, we determine that a small fraction of mathematical sciences departments receive a large majority of federal funding awarded to support mathematics in the United States. Additionally, we identify the extent to which women faculty members are underrepresented in mathematical sciences PhD-granting institutions in the United States. We also show that this underrepresentation of women faculty is even more pronounced in departments that received more federal grant funding.
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain.
Joining a research group is one of the most important events on a graduate student's path to earning a PhD, but the ways students go about searching for a group remain largely unstudied. It is therefore crucial to investigate whether departments are equitably supporting students as they look for an advisor, especially as students today enter graduate school with more diverse backgrounds than ever before. To better understand the phenomenon of finding a research group, we use a comparative case study approach to contrast important aspects of two physics PhD students' experiences. Semi-structured interviews with the students chronicled their interactions with departments, faculty, and the graduate student community, and described the resources they found most and least helpful. Our results reveal significant disparities in students' perceptions of how to find an advisor, as well as inequities in resources that negatively influenced one student's search. We also uncover substantial variation regarding when in their academic careers the students began searching for a graduate advisor, indicating the importance of providing students with consistent advising throughout their undergraduat
The Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter is an almost exact closed-form approximation to the Bayes-optimal multi-target tracking algorithm. Due to its optimality guarantees and ease of implementation, it has been studied extensively in the literature. However, the challenges involved in implementing the GM-PHD filter efficiently in a distributed (multi-sensor) setting have received little attention. The existing solutions for distributed PHD filtering either have a high computational and communication cost, making them infeasible for resource-constrained applications, or are unable to guarantee the asymptotic convergence of the distributed PHD algorithm to an optimal solution. In this paper, we develop a distributed GM-PHD filtering recursion that uses a probabilistic communication rule to limit the communication bandwidth of the algorithm, while ensuring asymptotic optimality of the algorithm. We derive the convergence properties of this recursion, which uses weighted average consensus of Gaussian mixtures (GMs) to lower (and asymptotically minimize) the Cauchy-Schwarz divergence between the sensors' local estimates. In addition, the proposed method is a