Objectives: Multimodal large language models (MLLMs) have shown potential for medical image classification. We evaluated four optimization strategies in two MLLMs-GPT-4o (gpt-4o-2024-08-06) and Gemini 2.5 Flash-Lite-for ultrasound-based thyroid nodule malignancy classification using two public datasets and a clinical cohort of nodules with atypia of undetermined significance (AUS) cytology. Methods: Text prompting, few-shot learning, fine-tuning, and a hybrid strategy combining fine-tuning with few-shot learning were evaluated for each model. Performance was assessed using the Digital Database of Thyroid Images (DDTI; n = 80), a 1000-image test subset of TN5000, and an institutional AUS cohort with surgical pathology (n = 84). In the AUS cohort, the best-performing strategy was compared with the consensus classification of three endocrinologists and the American Thyroid Association (ATA) ultrasound risk stratification. Results: For GPT-4o, the hybrid strategy achieved the highest area under the receiver operating characteristic curve (AUC) in DDTI (0.866), TN5000 (0.689), and the AUS cohort (0.836). In the AUS cohort, its specificity was higher than that of endocrinologist consensus and ATA risk stratification when only high-suspicion nodules were classified as malignant (95.1% vs. 70.7% and 70.7%; p = 0.002 and p = 0.001, respectively), while sensitivity did not differ significantly (72.1% vs. 74.4% and 79.1%, respectively; both p > 0.05). However, the hybrid model misclassified 12 of 43 malignant nodules, corresponding to a false-negative rate of 27.9%. When high- and intermediate-suspicion ATA categories were classified as malignant, ATA sensitivity increased to 83.7% and specificity decreased to 56.1%; the hybrid model had a higher AUC than ATA risk stratification (0.836 vs. 0.749; p = 0.017). For Gemini 2.5 Flash-Lite, few-shot learning, fine-tuning, and the hybrid strategy did not improve AUC relative to text prompting in any dataset. Conclusions: The hybrid strategy produced the most consistent performance gains for GPT-4o across the three datasets but did not improve Gemini 2.5 Flash-Lite. The optimized GPT-4o model achieved high specificity in the diagnostically challenging AUS cohort, although its false-negative rate limits its use as a stand-alone diagnostic tool. Further validation in larger, prospective multicenter cohorts is required before clinical use.
Alzheimer disease (AD) is a progressive neurodegenerative disorder with rapidly growing global prevalence. Early detection is critical for timely intervention; yet, conventional diagnostic methods remain costly and invasive. Speech-based assessment has emerged as a noninvasive alternative, as AD characteristically impairs linguistic abilities including fluency, coherence, and informational content. Recent advances in large language models (LLMs) offer new opportunities to extract structured linguistic features from transcribed speech for automated AD classification. However, existing LLM-based approaches often lack transparency and clinical interpretability, limiting their adoption in clinical workflows. This study aims to investigate the influence of linguistic features extracted from transcribed speech, as analyzed by LLMs, on the accuracy and interpretability of AD prediction. We propose a framework that leverages LLMs to analyze linguistic features extracted from transcribed speech for AD classification. Our approach focuses on 4 key aspects, including readability, fluency, richness of detail, and keyword relevance. To enhance classification accuracy, the framework integrates transcript embeddings with feature explanation embeddings, forming a comprehensive linguistic representation. We conducted extensive ablation studies to evaluate the contributions of individual features and benchmarked our framework against existing LLM-driven methodologies through pairwise explainability evaluations. Output stability was assessed across 3 independent pipeline runs. A fully local configuration (Llama 3 8B + nomic-embed-text) was tested to evaluate privacy-preserving deployment feasibility. Explainability was assessed via LLM-based pairwise comparison (Gemini-3.1-flash-lite) against the method of Bang et al across 54 correctly classified cases and by blinded evaluation from 2 neurologists. The proposed framework achieved a mean precision of 91.52%, a sensitivity of 91.08%, a specificity of 96.29%, and F1-score of 91.05% across 3 independent runs on the ADReSSo 2021 dataset, outperforming existing LLM-based approaches. A fully-local configuration (Llama 3 8B+nomic-embed-text, requiring no cloud application programming interface access) achieved an F1-score of 81.58%, demonstrating framework transferability to privacy-preserving deployment environments. Keyword relevance was the most influential feature (F1-score drop of 13.22 pp when removed). Explainability evaluations showed our method was preferred in 49 out of 54 cases via Gemini-3.1-flash-lite, with human experts preferring our method in 89 of 108 blinded assessments. These findings highlight that a structured linguistic feature analysis using LLMs provides a robust and interpretable framework for preliminary AD detection. Our approach offers a scalable and accessible solution that bridges artificial intelligence-driven text analysis with clinical applications, supporting early detection of cognitive decline through noninvasive assessment methods.
To develop and validate an agent-based Large Language Model (LLM) system for extracting structured data from breast cancer synoptic pathology reports and assess the performance gap between synthetic and real-world validation. We developed a modular artificial intelligence (AI) agent-based framework employing sequential specialized LLMs for parsing pathology reports and extracting structured data. We normalized College of American Pathologists (CAP) cancer protocols into 8 sections, 86 subsections, and 229 discrete fields. Seven leading LLMs (gemini-2.5-pro, llama3.3-70b, phi4-14b, deepseek-r1 14B/70B, gemma3-27b, gemini-2.0-flash-lite) were validated using dual evaluation: synthetic validation (864 controlled test cases) and real-world ground truth (6651 annotated fields from 90 pathology reports). Synthetic validation demonstrated strong performance (accuracy: 93.8%-99.0%). Real-world evaluation revealed field extraction recall ranging from 61.8% to 87.7%, demonstrating a substantial "reality gap" with performance drops of 11-32 percentage points. The gemini-2.5-pro model achieved the highest real-world recall (87.7%). Model size did not predict performance: the 14B-parameter deepseek-r1 (77.6%) outperformed its 70B-parameter counterpart (70.4%). The substantial performance degradation from synthetic to real-world data underscores the complexity of authentic clinical documentation. Smaller models can achieve competitive or superior recall, reducing computational costs. With even the best models missing 12%-38% of annotated fields, mandatory human verification is essential for clinical deployment. While LLM-based extraction systems show promise for pathology data extraction, synthetic validation alone provides false confidence. Rigorous real-world ground truth evaluation with expert annotation is essential before clinical deployment. These systems are best positioned as screening tools with mandatory human oversight rather than autonomous decision-making systems.
Large language models (LLMs) are increasingly evaluated using medical examination datasets, yet most studies emphasize overall accuracy rather than the psychometric structure of test items. We evaluated five LLMs on 199 text-only cardiology residency in-service examination items previously characterized using resident-derived psychometric metrics. Three frontier models were compared with two open-source comparators using a standardized zero-shot, repeated-query protocol and strict-majority scoring. Frontier models achieved substantially higher accuracy than open-source comparators, with Claude Opus 4.6, Gemini 3.1 Flash-Lite, and GPT-5.4 reaching 86.4%, 82.9%, and 81.9%, respectively, compared with 53.3% for MedQwen and 18.6% for Qwen-3.5-35B. Across frontier models, performance increased progressively from hard to easy resident-derived item strata. In multivariable analyses, item difficulty was the only classical psychometric factor consistently associated with AI correctness. IRT-based re-analysis confirmed that higher latent item difficulty was independently associated with lower frontier-model accuracy. Human-AI item-level correlations were modest but exceeded permutation-based null expectations, and frontier-model errors were concentrated among highly ranked human distractors. These findings show that item-level psychometric analysis helps explain variation in frontier LLM performance beyond overall accuracy alone. However, expert-rated rationale assessment and none-of-the-above perturbation testing revealed that strong examination accuracy did not guarantee high-quality explanatory support or reliable recognition of answer absence, indicating that examination performance and answer-selection robustness represent related but distinct dimensions of model behavior.
Mental-health geomatics require reliable ways to convert high-frequency GPS trajectories into meaningful place types that support indicators such as homestay, location entropy, and spatial extent of daily activities. Raw coordinates are typically noisy and carry little semantic information. We introduce H3-MOSAIC(H3-based Multimodal OSM-and-Satellite AI for Classification), a multimodal generative framework that fuses OpenStreetMap (OSM) building text and satellite imagery on H3 grids to infer place semantics from high-frequency GPS. Raw GPS was smoothed by minute-level speed filtering, then assigned to Level 10 H3 hexagons. Cells were retained if the mean speed was ≤ 1.2 m/s and the cumulative duration was ≥ 15 min, contiguous cells were merged, and home was defined as the cell with the longest dwell from 23:45 to 06:00. We compared text-only OSM classification with image-based and fused approaches across open-source models (DeepSeek, CLIP, LLaVA, Qwen-VL) and proprietary models (GPT-4o-mini, Gemini-2.5-flash-lite). Performance was assessed by accuracy, Cohen's kappa, precision, recall, F-measure, and confusion matrices. Day level associations between H3 semantic exposures and stress were examined by a random forest model and explainable methods. Multimodal methods outperformed single-modality baselines. In the 11-class task, accuracies were: CLIP 0.179, LLaVA 0.269, Qwen-VL 0.565, GPT-4o-mini 0.779, and Gemini-2.5-flash-lite 0.790. In the 5-class consolidation, accuracies rose to 0.702 (Qwen-VL), 0.849 (GPT-4o-mini), and 0.858 (Gemini-2.5-flash-lite). Text-only OSM baselines were lower (≈ 0.60-0.68). Across 3,845 hexagons with OSM text, closed-source models agreed on 79% of labels; disagreements concentrated in mixed-use, office, and green classes. Error modes reflected area-dominant versus keyword-triggered reasoning, hybrid-parcel ambiguity, tag sparsity, and symbolic artifacts. Stabilized semantics support more robust computation of homestay, entropy, and activity space and are suitable for privacy-aware, cross-city reuse. In a day-level case study, minutes at Home related to lower stress; Green showed a U-shaped pattern. H3-MOSAIC provides a scalable, auditable pipeline for semantic place detection from high-frequency GPS. Multimodal fusion markedly improves accuracy and consistency. Proprietary models are most robust on hard classes and open-source models are practical for coarse taxonomies. H3 day level exposures show stress patterns consistent with established mental health pathways, supporting face validity. The framework enables downstream exposure analyses with reduced misclassification and improved interpretability.
Comprehensive clinical history taking is essential for high-quality care. We hypothesized that large language models (LLMs), guided by a structured agentic framework, can efficiently obtain clinically meaningful patient histories. We developed an iterative prompting system that evaluates relevance and completeness across standard history domains and generates targeted follow-up questions until sufficient detail is obtained. We built a patient-facing application and evaluated it using 52 published case reports and 20 constructed clinical scenarios with simulated patient interactions. The framework was implemented using GPT-4o, Gemini-2.5-Flash-Lite, or Grok-3. After each interaction, the system generated an EHR-ready clinical summary, differential diagnosis, and recommended investigations. Across models, relevant history elements were captured with >85% accuracy and F1 scores, as independently assessed by three blinded physicians, and recommended investigations aligned with those used to establish final diagnoses. These findings support the potential of agentic LLM systems for structured clinical history collection and justify prospective clinical evaluation.
Proactive robotic assistance in human-robot collaboration (HRC) requires systems that can perceive evolving task contexts, anticipate user needs, and intervene appropriately without disrupting human workflow. We present the Agentic Unified Robotic Assistance (AURA) Framework, which couples Large Language Model (LLM) reasoning grounded by Standard Operating Procedures (SOPs) with a modular layer of specialized Intent, Motion, Perception, Sound, Affordance, and Performance Monitors that supply structured context to a central decision-making module, making the framework reconfigurable and auditable without retraining or re-prompting. We introduce a human-in-the-loop teleoperation data collection methodology and an offline evaluation scheme with an Appropriateness Score (A-Score) tailored to proactive intervention timing, and release a benchmark dataset of annotated multimodal HRC episodes containing workspace and robot wrist camera videos, robot joint states, and labeled intervention events. Across three tasks of varying complexity, we observe progressive gains in intent prediction and decision-making as the modules are supplied with richer grounded context (prior-state memory and tracked object locations), with Combined F1 rising by over 20 points between context-poor and context-rich conditions. The structured grounding allows lightweight multimodal backbones such as Gemini 3.1 Flash Lite to perform on par with heavier reasoning-tier models at roughly one-fifth the inference latency. Together, these contributions establish a scalable framework, benchmark, and evaluation methodology for advancing proactive robotic assistance in collaborative environments.
Background: Accurate identification and positioning of thoracic devices on chest radiographs is critical for patient safety in intensive care. Multimodal large language models (LLMs) offer potentially generalizable automated evaluation, but their performance in this domain is underexplored. Methods: Three multimodal LLMs (GPT-4o, gpt-4o-2024-08-06; Gemini 3.1 Flash Lite Preview; Claude Sonnet 4.6) were evaluated on 4813 chest radiographs from the RANZCR CLiP dataset for device presence and positioning of ETT, NGT, CVC, and Swan-Ganz catheters. Performance was quantified with 95% Wilson confidence intervals, balanced accuracy, MCC, Cochran's Q, Bonferroni-corrected McNemar, and Cohen's/Fleiss' kappa. Six additional analyses were performed: a blinded paired reader study (n = 377; two board-certified radiologists, blinded to ground truth and to all LLM outputs), external validation on PadChest (n = 200, device-presence detection only-PadChest lacks granular position labels), three-variant prompt-sensitivity analysis (n = 103), repeat-inference stability across three runs (n = 50), systematic error taxonomy, and a failure-case analysis. Results: Device-presence performance varied widely across models; abnormal-position sensitivity was uniformly poor (MCC ≤ 0.028; balanced accuracy 0.41-0.53). Inter-model agreement was poor to slight (Fleiss' κ: 0.005-0.383 for presence; -0.280 to -0.025 for classification). Radiologists numerically outperformed all three LLMs in 42/42 paired comparisons; the superiority was statistically significant after Bonferroni correction in 33/42 (32/42 at p < 0.001). PadChest replicated the negative finding for device-presence detection (malposition not externally validated). Prompts and inference stochasticity introduced 2-3× sensitivity swings and run-to-run κ from 0.20 to 0.85. Case failures concentrated systematically in multi-device cases (p < 0.0001) but not in abnormal-position cases (p = 0.14). Conclusions: Current general-purpose multimodal LLMs are not yet reliable for autonomous thoracic-device assessment; their failure patterns are structurally characterizable across models, prompts, and case types and support, at most a circumscribed role, as adjunct device-presence screening tools. The findings do not generalize to purpose-built, regulator-approved clinical AI systems.
Large language models (LLMs) are increasingly integrated into health care applications; however, their vulnerability to prompt-injection attacks (ie, maliciously crafted inputs that manipulate an LLM's behavior) capable of altering medical recommendations has not been systematically evaluated. To evaluate the susceptibility of commercial LLMs to prompt-injection attacks that may induce unsafe clinical advice and to validate man-in-the-middle, client-side injection as a realistic attack vector. This quality improvement study used a controlled simulation design and was conducted between January and October 2025 using standardized patient-LLM dialogues. The main experiment evaluated 3 lightweight models (GPT-4o-mini [LLM 1], Gemini-2.0-flash-lite [LLM 2], and Claude-3-haiku [LLM 3]) across 12 clinical scenarios in 4 categories under controlled conditions. The 12 clinical scenarios were stratified by harm level across 4 categories: supplement recommendations, opioid prescriptions, pregnancy contraindications, and central-nervous-system toxic effects. A proof-of-concept experiment tested 3 flagship models (GPT-5 [LLM 4], Gemini 2.5 Pro [LLM 5], and Claude 4.5 Sonnet [LLM 6]) using client-side injection in a high-risk pregnancy scenario. Two prompt-injection strategies: (1) context-aware injection for moderate- and high-risk scenarios and (2) evidence-fabrication injection for extremely high-harm scenarios. Injections were programmatically inserted into user queries within a multiturn dialogue framework. The primary outcome was injection success at the primary decision turn. Secondary outcomes included persistence across dialogue turns and model-specific success rates by harm level. Across 216 evaluations (108 injection vs 108 control), attacks achieved 94.4% (102 of 108 evaluations) success at turn 4 and persisted in 69.4% (75 of 108 evaluations) of follow-ups. LLM 1 and LLM 2 were completely susceptible (36 of 36 dialogues [100%] each), and LLM 3 remained vulnerable in 83.3% of dialogues (30 of 36 dialogues). Extremely high-harm scenarios including US Food and Drug Administration Category X pregnancy drugs (eg, thalidomide) succeeded in 91.7% of dialogues (33 of 36 dialogues). The proof-of-concept experiment demonstrated 100% vulnerability for LLM 4 and LLM 5 (5 of 5 dialogues each) and 80.0% (4 of 5 dialogues) for LLM 6. In this quality improvement study using a controlled simulation, commercial LLMs demonstrated substantial vulnerability to prompt-injection attacks that could generate clinically dangerous recommendations; even flagship models with advanced safety mechanisms showed high susceptibility. These findings underscore the need for adversarial robustness testing, system-level safeguards, and regulatory oversight before clinical deployment.
Weeds remain a major constraint to row-crop productivity, yet current deep learning approaches for UAV imagery often require extensive annotation, generalize poorly across fields, and provide limited interpretability. We investigate whether modern vision-language models (VLMs) can address these gaps in a zero-shot setting. Using drone images from soybean fields with ground-truth weed boxes, we evaluate six commercial VLMs, ChatGPT-4.1, ChatGPT-4o, Gemini Flash 2.5, Gemini Flash Lite 2.5, LLaMA-4 Scout, and LLaMA-4 Maverick under a unified prompt that elicits (i) weed presence, (ii) spatial localization, (iii) reasoning, (iv) crop growth stage, and (v) crop type. We further introduce Error-Probing Prompting (EPP), a counterfactual follow-up that forces re-analysis under the assumption that weeds are present, and we quantify self-correction with expert-rated interpretability scores (Grounding, Specificity, Plausibility, Non-Hallucination, Actionability). Across models, Gemini Flash 2.5 delivers the most consistent zero-shot performance and highest interpretability, ChatGPT-4.1 provides the strongest reasoning but lower raw detection, ChatGPT-4o offers a balanced profile, and LLaMA-4 variants lag in localization and specificity. Gemini Flash Lite 2.5 is efficient but fails EPP stress tests, revealing brittle reasoning. Visual grounding analysis and a text-to-region overlap metric show that interpretability tracks spatial correctness. Results highlight that explainability and feedback driven adaptability not scale alone best predict reliability for field deployment, and position VLMs as promising, low-annotation tools for precision weed management.
Extraoral and intraoral dental photographs serve as preoperative records and document the entire treatment. Correctly composed orthodontic photographs are crucial for remote diagnosis and may serve as a bulwark against medicolegal challenges. In this prospective study, intraoral frontal photographs of patients with ideal occlusion were taken using two types of lenses (EF-S 18-55 mm f/3.5-5.6 IS STM lens (Canon, Tokyo, JP), SP 90 mm F/2.8 MACRO VC lens (Model F017 Tamron, NY, USA)) and two different ring flash systems (Meike FC-100 Macro Ring LED Light (Meike, China), Macro Ring flash Lite YN-14EX (Yongnuo digital, China)). The combination of lens and flash used was grouped into four groups. Twenty-eight intraoral photographs of patients were taken. An image quality assessment survey was distributed among two groups - 50 orthodontists and 50 other dental specialists. The participants were asked to assess all the intraoral images and subjectively score them on a scale of one to ten, with one being very poor and ten being excellent, considering the sharpness, color, brightness, contrast, and overall quality of the image. The general dentists rated the images taken with a 90-mm macro lens and ring flash as the best quality photographs. Images obtained using an 18-55 mm lens and ring LED received significantly lesser scores and were graded good by dentists. This combination of lens and flash may prove a valuable investment in the long-term aiding in excellent dental images for diagnosis and treatment monitoring.
The new generation LED curing light units have significantly improved curing performance compared to first generation lights, and even some second generation LED curing light units. This study compared the curing performance of 10 new generation LED light curing units (FLASH-lite 1401, LE Demetron 1, Coltolux, Ultra-Lume 5, Mini LED, bluephase, Elipar FreeLight 2, Radii, Smartlite IQ and Allegro) for depth of cure against a high-powered halogen curing light unit (Optilux 501). Depth of cure measurements were utilized per the ANSI/ADA No 27 standard to detect differences between the lights at three time intervals (10, 20 and 40 seconds). A total of 660 samples were prepared (n=10/group). A full factorial ANOVA and Tukey's HSD test showed FLASH-lite 1401 performed significantly better than the other lights at 10- and 20-second time intervals (p<0.01). This study also demonstrated that an exposure time of 20 seconds or longer assures a better depth of cure, 40 seconds being the optimal polymerization time for all of the curing light units.
Mobile telephones have significant potential for use in psychological research, possessing unique characteristics-not least their ubiquity--that may make them useful tools for psychologists. We examined whether it is possible to measure reaction times (RTs) accurately using Adobe Flash Lite on mobile phones. We ran simple and choice RT experiments on two widely available mobile phones, a Nokia 6110 Navigator and a Sony Ericsson W810i, using a wireless application protocol (WAP) connection to access the Internet from the devices. RTs were compared within subjects with those obtained using a Linux-based millisecond-accurate measurement system. Results show that measured RTs were significantly longer on mobile devices, and that overall RTs and distribution of RTs varied across devices.
We determined if PR3-ANCA is a biomarker that differentiates ulcerative colitis (UC) from Crohn's disease (CrD). A total of 946 sera were tested, including 86 granulomatosis with polyangiitis (GPA) and 491 inflammatory bowel disease (IBD) patients (283 UC and 208 CrD), 264 pathological controls (various diseases) and 105 healthy individuals. All samples were tested for PR3-ANCA by ELISA (QUANTA Flash Lite®, INOVA Diagnostics) and chemiluminescent immunoassays (CIA QUANTA Flash PR3). Conventional anti-neutrophil cytoplasmic antibody (ANCA) indirect immunofluorescence assays (IIF) was performed with NOVA Lite™ (INOVA Diagnostics). PR3-ANCA by CIA were detected in 31.1% UC vs. 1.9% CrD sera (p=2.2E-16), and by ELISA in 6% UC and 0% CrD (p=0.0003). In GPA patients, PR3-ANCA were detected in 75.6% by CIA and 61.6% by ELISA (p<0.05). PR3-ANCA by CIA were more prevalent in E3-UC compared to E1/2-UC (p<0.05), and in patients with shorter disease duration (p<0.0001). PR3-ANCA showed similar sensitivity, but significantly higher specificity (p<0.05), compared to atypical pANCA by IIF. The novel PR3 CIA may prove helpful in the differentiation of CrD from UC, as well as in the identification of UC patients with more extensive disease.
The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e., circumventing safety alignments to produce outputs violating regulatory standards, assuming threats from authorized users, such as operators, who craft malicious prompts to elicit non-compliant guidance. Three state-of-the-art LLMs (OpenAI's GPT-4o mini, Google's Gemini 2.0 Flash-Lite, and Anthropic's Claude 3.5 Haiku) were tested against Baseline, BitBypass, and DeepInception jailbreaking methods across scenarios derived from nine NERC Reliability Standards (EOP, TOP, and CIP). In the initial broad experiment, the overall Attack Success Rate (ASR) was 33.1%, with DeepInception proving most effective at 63.17% ASR. Claude 3.5 Haiku exhibited complete resistance (0% ASR), while Gemini 2.0 Flash-Lite was most vulnerable (55.04% ASR) and GPT-4o mini moderately susceptible (44.34% ASR). A follow-up experiment refining malicious wording in Baseline and BitBypass attacks yielded a 30.6% ASR, confirming that subtle prompt a
While Mixture-of-Experts (MoE) architectures have become the standard for sparsity scaling in large language models, they increasingly face diminishing returns and system-level bottlenecks. In this work, we explore embedding scaling as a potent, orthogonal dimension for scaling sparsity. Through a comprehensive analysis and experiments, we identify specific regimes where embedding scaling achieves a superior Pareto frontier compared to expert scaling. We systematically characterize the critical architectural factors governing this efficacy -- ranging from parameter budgeting to the interplay with model width and depth. Moreover, by integrating tailored system optimizations and speculative decoding, we effectively convert this sparsity into tangible inference speedups. Guided by these insights, we introduce LongCat-Flash-Lite, a 68.5B parameter model with ~3B activated trained from scratch. Despite allocating over 30B parameters to embeddings, LongCat-Flash-Lite not only surpasses parameter-equivalent MoE baselines but also exhibits exceptional competitiveness against existing models of comparable scale, particularly in agentic and coding domains.
Fine-grained attribute prediction is essential for fashion retail applications including catalog enrichment, visual search, and recommendation systems. Vision-Language Models (VLMs) offer zero-shot prediction without task-specific training, yet their systematic evaluation on multi-attribute fashion tasks remains underexplored. A key challenge is that fashion attributes are often conditional. For example, "outer fabric" is undefined when no outer garment is visible. This requires models to detect attribute applicability before attempting classification. We introduce a three-tier evaluation framework that decomposes this challenge: (1) overall task performance across all classes (including NA class: suggesting attribute is not applicable) for all attributes, (2) attribute applicability detection, and (3) fine-grained classification when attributes are determinable. Using DeepFashion-MultiModal, which explicitly defines NA (meaning attribute doesn't exist or is not visible) within attribute label spaces, we benchmark nine VLMs spanning flagship (GPT-5, Gemini 2.5 Pro), efficient (GPT-5 Mini, Gemini 2.5 Flash), and ultra-efficient tiers (GPT-5 Nano, Gemini 2.5 Flash-Lite) against class
LLMs show strong performance in code generation, but their outputs lack correctness guarantees. Sample-based uncertainty estimators address this by generating multiple candidate programs and measuring their disagreement. However, existing estimators make different design choices about how behaviours are identified, aggregated, referenced and compared, making them difficult to assess. We therefore first introduce a taxonomy that disentangles these choices and reveals a missing design point: semantic distance-aware uncertainty estimation, which measures not only whether sampled programs disagree, but how severely their execution behaviours differ. Across LiveCodeBench, MBPP, HumanEval-X and BigCodeBench, spanning Python, Java and C++, our metrics provide strong proxies for correctness, and consistently outperform state-of-the-art sample-based baselines across both closed-source models (GPT-3.5-Turbo, GPT-4o-mini, Gemini-2.5-Flash-Lite, Claude Opus 4.5) and an open-source model (DeepSeek-Coder-V2). The method is practical: it requires neither model internals nor LLM-as-judge calls, remains robust across models, languages, sampling temperatures and fuzzing settings, and reduces runtime
The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific knowledge. A core obstacle is the lack of established scientific terminology in these languages. We introduce AfriScience-MT, a parallel corpus covering six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, and isiZulu) across 11 scientific domains. Professional translators, working with expert science communicators, translated plain-language summaries of scientific papers into each target language and created new terms where none existed. We benchmark machine translation systems and large language models in zero-shot, few-shot, and fine-tuned settings. Our results show that closed-source models outperform all open-source models at both the sentence and document levels: GPT-5.4 and Gemini-3.1-Flash-Lite lead with average sentence-level COMET scores of 68.3 and 68.0, respectively, and tie at an average document-level COMET of 48.3. Among open systems, fine-tuned NLLB-1.3B reaches 67.3 at the sentence level, and TranslateGemma-12B reaches 44.0 at the document level with 1-shot in-context
This study adopts a behavioural bottom-up approach to AI value alignment to investigate whether an implicitly conveyed user identity shifts the moral evaluations of large language models (LLMs). Through a structured, multi-turn conversational protocol across 12,000 interactions, we evaluate AI value alignment in two non-reasoning models, gpt-4.1-mini-2025-04-14 and gemini-2.5-flash-lite. Rather than instructing the models to adopt a persona or prompting them with explicit moral stances, the user's professional role is introduced purely through value-neutral reasoning. The models are then asked for wrongness ratings from 0-100 on ten common-morality rules from Gert's moral framework. The results show that moral judgments vary with the user's role across both models. While grave-harm acts like killing exhibit a strong ceiling effect, contestable rule-governed acts demonstrate role-conditioned shifts that mirror the relationship between the user's profession and the act being rated. These findings demonstrate that unintended contextual conditioning via user identity permeates LLM moral evaluations, posing questions for the AI value alignment discourse regarding how to define acceptabl