Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist centres, where timely inference, limited hardware, and local handling of patient images are vital for practical, privacy-preserving clinical prescreening. Here we present Pocket-Dentist, an efficiency-aware benchmark for dental multimodal question answering that brings together three datasets spanning approximately 1,159 patients from BRAR and MetaDent, five task types and seven metrics. Across 14 typical VLMs, our results reveal an interesting observation: compact VLMs, such as 2B-parameter models, become competitive with much larger VLMs on most metrics after lightweight adaptation while requiring substantially lower computational costs in dental image understanding. Deployed locally on an iPhone 17 Pro, our finetuned compact VLM Pocket-Dentist-2B processed each sample in 4.31 s, reducing latency by 4.9x and memory use by 2.3x compared with a 7B baseline. Our project page is available at https://2026-icml.github.io/pocket-dentist-icml.
Decarbonizing electricity generation, heating, and transportation simultaneously requires integrated planning tools that can coordinate multiple energy production sources and demand points while remaining reliable under uncertainty. This paper develops a carbon-aware and uncertainty-resilient optimization framework for a grid-connected multiple source hub that co-optimizes electricity, heating, cooling, and transport energy services with an explicit hydrogen sub-hub. The proposed model is formulated as an MILP over a 25-year planning horizon (2025-2050). The hub integrates renewable electricity (Photovoltaic and wind), dispatchable resources (including natural-gas-based conversion), storage systems, demand response, and a hydrogen subsystem comprising an electrolyzer and hydrogen storage to supply hydrogen-vehicle demand and provide temporal flexibility. Two policy archetypes are examined: a Carbon Tax (price instrument), and a Net-Zero pathway (quantity instrument). To hedge feasibility-critical operational uncertainty, the deterministic model is extended using budgeted robust optimization and a tunable uncertainty budget. The developed scheme is applied to the province of Ontario
Self-represented tenants, landlords, and help-desk staff need to be pointed at the provision of law that actually governs a question, with a correct statutory citation. We study this task on the Ontario Residential Tenancies Act, 2006 (RTA) and its core regulation, asking the operator's question empirically: is fine-tuning enough, or is hybrid retrieval needed? We run a four-arm head-to-head on Qwen2.5-7B-Instruct (base zero-shot, LoRA SFT-only, RAG-only, and an SFT+RAG hybrid), scored on citation exact-match (section+subsection) over a small, human-verification-pending real eval set. The base model cannot cite the RTA and SFT-only mis-recalls sections; retrieval is essential and drives hallucination to zero by construction; and the SFT+RAG hybrid scores highest at 0.481 exact-match with zero hallucinated citations. Its edge comes from SFT making provision selection more robust to the higher-recall candidate sets that hurt zero-shot RAG. Notably, this cheap bge-small hybrid matches or beats a pipeline built on bigger, specialized retrieval models (a larger embedder and a cross-encoder reranker), and a larger/improved training set does not help either: strong statutory-citation perf
Model interpretability is crucial for establishing AI safety and clinician trust in medical applications for example, in survival modelling with competing risks. Recent deep learning models have attained very good predictive performance but their limited transparency, being black-box models, hinders their integration into clinical practice. To address this gap, we propose an intrinsically interpretable survival model called CRISPNAM-FG. Leveraging the structure of Neural Additive Models (NAMs) with separate projection vectors for each risk, our approach predicts the Cumulative Incidence Function using the Fine-Gray formulation, achieving high predictive power with intrinsically transparent and auditable predictions. We validated the model on several benchmark datasets and applied our model to predict future foot complications in diabetic patients across 29 Ontario hospitals (2016-2023). Our method achieves competitive performance compared to other deep survival models while providing transparency through shape functions and feature importance plots.
Existing benchmarks evaluating biases in large language models (LLMs) primarily rely on explicit cues, declaring protected attributes like religion, race, gender by name. However, real-world interactions often contain implicit biases, inferred subtly through names, cultural cues, or traits. This critical oversight creates a significant blind spot in fairness evaluation. We introduce ImplicitBBQ, a benchmark extending the Bias Benchmark for QA (BBQ) with implicitly cued protected attributes across 6 categories. Our evaluation of GPT-4o on ImplicitBBQ illustrates troubling performance disparity from explicit BBQ prompts, with accuracy declining up to 7% in the "sexual orientation" subcategory and consistent decline located across most other categories. This indicates that current LLMs contain implicit biases undetected by explicit benchmarks. ImplicitBBQ offers a crucial tool for nuanced fairness evaluation in NLP.
Mixed Reality (MR) is proven in the literature to support precise spatial dental drill positioning by superimposing 3D widgets. Despite this, the related knowledge about widget's visual design and interactive user feedback is still limited. Therefore, this study is contributed to by co-designed MR drill tool positioning widgets with two expert dentists and three MR experts. The results of co-design are two static widgets (SWs): a simple entry point, a target axis, and two dynamic widgets (DWs), variants of dynamic error visualization with and without a target axis (DWTA and DWEP). We evaluated the co-designed widgets in a virtual reality simulation supported by a realistic setup with a tracked phantom patient, a virtual magnifying loupe, and a dentist's foot pedal. The user study involved 35 dentists with various backgrounds and years of experience. The findings demonstrated significant results; DWs outperform SWs in positional and rotational precision, especially with younger generations and subjects with gaming experiences. The user preference remains for DWs (19) instead of SWs (16). However, findings indicated that the precision positively correlates with the time trade-off. Th
The effect of changing greenhouse gas concentrations on climate was examined. Calculations of the climate sensitivity, the warming of the Earth due to a doubling of atmospheric CO2, are discussed. Ontario was responsible for 0.35% of the world's CO2 emissions in 2019 and this amount was 20% lower than in 2005. The predictions of Global Climate Models (GCM) were compared to observations. Records since 1880 show an overall warming of about 1 C. The GCMs do not account for observed decadal temperature fluctuations and consistently overestimate the warming. Ontario's contribution to global warming is only 9.2 x 10^{-5} C/year using the Intergovernmental Panel on Climate Change recommended climate sensitivity value. Measurements of the polar ice caps reveal a decrease in the minimum September Arctic sea ice extent during 1979-2022 but the trend levelled off after 2007; while the average Antarctic sea ice extent slightly increased. Sea level increased slightly throughout the 20th century. Ontario's contribution to anthropogenic sea level rise is about 0.005 mm/year. Sea level along Ontario's Hudson Bay coast is decreasing due to isostatic rebound of the land following the last Ice Age. T
A natural way of quantifying the ``amount of information'' in decision problems yields a globally concave value for information. Another (in contrast, adversarial) way almost never does.
Diagnosing and managing oral diseases necessitate advanced visual interpretation across diverse imaging modalities and integrated information synthesis. While current AI models excel at isolated tasks, they often fall short in addressing the complex, multimodal requirements of comprehensive clinical dental practice. Here we introduce DentVLM, a multimodal vision-language model engineered for expert-level oral disease diagnosis. DentVLM was developed using a comprehensive, large-scale, bilingual dataset of 110,447 images and 2.46 million visual question-answering (VQA) pairs. The model is capable of interpreting seven 2D oral imaging modalities across 36 diagnostic tasks, significantly outperforming leading proprietary and open-source models by 19.6% higher accuracy for oral diseases and 27.9% for malocclusions. In a clinical study involving 25 dentists, evaluating 1,946 patients and encompassing 3,105 QA pairs, DentVLM surpassed the diagnostic performance of 13 junior dentists on 21 of 36 tasks and exceeded that of 12 senior dentists on 12 of 36 tasks. When integrated into a collaborative workflow, DentVLM elevated junior dentists' performance to senior levels and reduced diagnosti
Objective: To develop machine learning models that can predict the number of COVID-19 cases per day given the last 14 days of environmental and mobility data. Approach: COVID-19 data from four counties around Toronto, Ontario, were used. Data were prepared into daily records containing the number of new COVID case counts, patient demographic data, outdoor weather variables, indoor environment factors, and human movement based on cell mobility and public health restrictions. This data was analyzed to determine the most important variables and their interactions. Predictive models were developed using CNN and LSTM deep neural network approaches. A 5-fold chronological cross-validation approach used these methods to develop predictive models using data from Mar 1 to Oct 14 2020, and test them on data covering Oct 15 to Dec 24 2020. Results: The best LSTM models forecasted tomorrow's daily COVID case counts with 90.7% accuracy, and the 7-day rolling average COVID case counts with 98.1% accuracy using independent test data. The best models to forecast the next 7 days of daily COVID case counts did so with 79.4% accuracy over all days. Models forecasting the 7-day rolling average case co
Severe windstorms pose threats to people, human-made structures, and the environment. An investigation of insured losses caused by windstorms is a multipurpose study that serves to advance the resilience and sustainability of modern communities. The present study proposes a systematic analysis of insured losses imposed by different types of windstorms in two Canadian provinces, Ontario (ON) and Quebec (QC), during the period 2008-2021. Actual wind damage data from the Canadian insurance market were considered in this study. Our calculations show that ON and QC received half of all wind catastrophes across Canada, and nearly three-quarters of all types of catastrophes in ON and QC were wind-related ones. The total windstorm loss of over CA$5.2 billion was not evenly distributed between QC and ON, but rather had a QC:ON ratio of 1:3.1. We attributed this discrepancy in the inflicted damage between two provinces to the predominantly eastward and northeastward storm trajectories and the higher density of wealth and population in ON. Convective storms were the most devastating wind type comprising nearly 65% and 67% of the total number of events and associated damage, respectively. Fina
The paper describes a system and experimental procedure that use integrating passive detectors, such as thermoluminescent dosimeters (TLDs), for the measurement of ultra-low-level ambient dose equivalent rate values at the underground SNOLAB facility located in Sudbury, Ontario, Canada. Because these detectors are passive and can be exposed for relatively long periods of time, they can provide better sensitivity for measuring ultra-low activity levels. The final characterization of ultra-low-level ambient dose around water shielding for ongoing direct dark matter search experiments in Cube Hall at SNOLAB underground laboratory is given. The conclusion is that TLDs provide reliable results in the measurement of the ultra-low-level environmental radiation background.
A healthy smile plays a significant role in functional as well as esthetic considerations, improving confidence. It is difficult for dental professionals to strike a balance between esthetic requirements and functional requirements. Traditional smile design has had heavy reliance on dentist expertise and used plaster models and hand drawings, raising questions about the outcome for patients. Digital technology, led by Dr. Christian Coachman in 2007, allows photographic and videographic assessments, enabling improved intercommunication among specialists and patients. Advances in artificial intelligence (AI) and big data have supported analysis of facial features and development of personalized smile designs in the last few years. Outputs are, however, susceptible to practitioner bias or limitations of training data, and may be suboptimal for individual users. The study presented here suggests a comprehensive system integrating AI, big data, and recognition technologies to automate the smile design process so that both experienced and inexperienced dentists can generate pleasing aesthetics with ease. The system has a Facial Feature Extraction Module and an Image Generation Module, se
The number of confirmed COVID-19 cases reached over 1.3 million in Ontario, Canada by June 4, 2022. The continued spread of the virus underlying COVID-19 has been spurred by the emergence of variants since the initial outbreak in December, 2019. Much attention has thus been devoted to tracking and modelling the transmission of COVID-19. Compartmental models are commonly used to mimic epidemic transmission mechanisms and are easy to understand. Their performance in real-world settings, however, needs to be more thoroughly assessed. In this comparative study, we examine five compartmental models -- four existing ones and an extended model that we propose -- and analyze their ability to describe COVID-19 transmission in Ontario from January 2022 to June 2022.
Purpose: To develop a virtual reality simulator for high dose rate prostate brachytherapy and to test whether participation is associated with immediate gains in self-reported confidence across predefined procedural domains in two cohorts. Methods: Two modules were developed and implemented using Unreal Engine: patient preparation and template guided needle insertion. Oncology staff and trainees completed pre and post surveys that assessed confidence for recalling steps, explaining steps, identifying equipment, and explaining equipment function. Studies were conducted at the Hands On Brachytherapy Workshop (HOWBT) in London, Ontario, and at Sunnybrook Odette Cancer Centre in Toronto, Ontario. Paired Wilcoxon signed rank tests with two-sided p values compared before and after scores within each module. Results: Patient preparation (N=11) confidence increased for recalling steps (W=65, p=0.002), explaining steps (W=51, p = 0.023), identifying equipment (W=65, p=0.002), and explaining equipment function (W=60, p=0.0078). Needle insertion (N=27) confidence increased for recalling steps (W=292, p<0.001), explaining steps (W=347, p<0.001), identifying equipment (W=355, p<0.001),
Hallucination is a common problem for Large Vision-Language Models (LVLMs) with long generations which is difficult to eradicate. The generation with hallucinations is partially inconsistent with the image content. To mitigate hallucination, current studies either focus on the process of model inference or the results of model generation, but the solutions they design sometimes do not deal appropriately with various types of queries and the hallucinations of the generations about these queries. To accurately deal with various hallucinations, we present a unified framework, Dentist, for hallucination mitigation. The core step is to first classify the queries, then perform different processes of hallucination mitigation based on the classification result, just like a dentist first observes the teeth and then makes a plan. In a simple deployment, Dentist can classify queries as perception or reasoning and easily mitigate potential hallucinations in answers which has been demonstrated in our experiments. On MMbench, we achieve a 13.44%/10.2%/15.8% improvement in accuracy on Image Quality, a Coarse Perception visual question answering (VQA) task, over the baseline InstructBLIP/LLaVA/Vis
Surgical guide plate is an important tool for the dental implant surgery. However, the design process heavily relies on the dentist to manually simulate the implant angle and depth. When deep neural networks have been applied to assist the dentist quickly locates the implant position, most of them are not able to determine the implant depth. Inspired by the video grounding task which localizes the starting and ending time of the target video segment, in this paper, we simplify the implant depth prediction as video grounding and develop a Texture Perceive Implant Depth Prediction Network (TPNet), which enables us to directly output the implant depth without complex measurements of oral bone. TPNet consists of an implant region detector (IRD) and an implant depth prediction network (IDPNet). IRD is an object detector designed to crop the candidate implant volume from the CBCT, which greatly saves the computation resource. IDPNet takes the cropped CBCT data to predict the implant depth. A Texture Perceive Loss (TPL) is devised to enable the encoder of IDPNet to perceive the texture variation among slices. Extensive experiments on a large dental implant dataset demonstrated that the pr
Panoramic X-ray is a simple and effective tool for diagnosing dental diseases in clinical practice. When deep learning models are developed to assist dentist in interpreting panoramic X-rays, most of their performance suffers from the limited annotated data, which requires dentist's expertise and a lot of time cost. Although self-supervised learning (SSL) has been proposed to address this challenge, the two-stage process of pretraining and fine-tuning requires even more training time and computational resources. In this paper, we present a self-supervised auxiliary detection (SSAD) framework, which is plug-and-play and compatible with any detectors. It consists of a reconstruction branch and a detection branch. Both branches are trained simultaneously, sharing the same encoder, without the need for finetuning. The reconstruction branch learns to restore the tooth texture of healthy or diseased teeth, while the detection branch utilizes these learned features for diagnosis. To enhance the encoder's ability to capture fine-grained features, we incorporate the image encoder of SAM to construct a texture consistency (TC) loss, which extracts image embedding from the input and output of
This article describes the clinical validation study setup, statistical analysis and results for a deep learning algorithm which detects dental anomalies in intraoral radiographic images, more specifically caries, apical lesions, root canal treatment defects, marginal defects at crown restorations, periodontal bone loss and calculus. The study compares the detection performance of dentists using the deep learning algorithm to the prior performance of these dentists evaluating the images without algorithmic assistance. Calculating the marginal profit and loss of performance from the annotated paired image data allows for a quantification of the hypothesized change in sensitivity and specificity. The statistical significance of these results is extensively proven using both McNemar's test and the binomial hypothesis test. The average sensitivity increases from $60.7\%$ to $85.9\%$, while the average specificity slightly decreases from $94.5\%$ to $92.7\%$. We prove that the increase of the area under the localization ROC curve (AUC) is significant (from $0.60$ to $0.86$ on average), while the average AUC is bounded by the $95\%$ confidence intervals ${[}0.54, 0.65{]}$ and ${[}0.82, 0
Reliable angler activity data inform fisheries management. Traditionally, such data are gathered through surveys, but an innovative cost-effective approach involves utilizing online platforms and smartphone applications. These citizen-sourced data were reported to correlate with conventional survey information. However, the nature of this correlation--whether direct or mediated by intermediate variables--remains unclear. We applied BNs to data from conventional surveys, the Angler's Atlas website, the MyCatch smartphone application, and environmental data across Alberta and Ontario, Canada, to detect probabilistic dependencies. Using Bayesian model averaging, we quantified the strength of connections between variables. Waterbody webpage views were directly related to daily and weekly-aggregated boat counts in Ontario (51\% and 100\% probability) and to weekly-aggregated creel survey-reported fishing duration in Alberta (100\%). This highlights the value of citizen-sourced data in providing unique insights beyond meteorological factors, with online interest serving as a potentially reliable proxy for angler pressure and effort.