The systematic integration of robotics into health service delivery systems requires periodic assessment of robotic readiness in terms of digital-health maturity regimes across countries. The current study aims to cluster 169 countries into maturity regimes and classify and predict cluster membership accuracy based on digital-health maturity dimensions determining the system's perception and interoperability, coordination, and workforce-regulatory reliability readiness. These country-level proxy concepts are applied due to a dearth of cross-country robotic readiness measures at the global level. The study also proposed an adaptive readiness framework for robotic deployment decision-making. This study applies a multilayered robotic systems readiness proxy framework based on diffusion of innovations and multi-agent systems theories, which hypothesise health-system readiness as adaptive and learning-driven instead of static. Using 31 global digital-health maturity indicators from the Global Digital Health Monitor 2023, three latent proxy dimensions are constructed to measure perception/interoperability readiness, governance/coordination readiness, and workforce-regulatory reliability readiness. The data represent 169 countries at different stages of digital-health maturity, starting from phase 1 to phase 5, where 67 countries reflect evidence-based digital-health maturity status and 102 countries had no data reported and are below phase 1. The sample included these 102 countries in the study, imputing the lowest values for each indicator to prioritise them and eliminate participation bias in global health policy decisions. To address concerns regarding missing data and its handling, the analysis was repeated across different robust scenarios. The latent proxy dimensions are created using principal component analysis, reducing 31 dimensions to 3 dimensions. Their associations were explored using ordinary least squares regression. Unsupervised clustering techniques (K-means, agglomerative hierarchical clustering, and Gaussian mixture modelling) were applied to identify latent readiness regimes, and comparative model assessment brought a four-cluster solution. External validation was conducted against the World Bank classification of countries. The separability of the latent regimes was assessed through repeated cross-validations and post-clustering recoverability analysis using conservative supervised learning models, where five latent clusters were considered in line with the World Health Organization's five-phase maturity categories. The three proxy indices have shown internally coherent loading structures with 85%-95% scale reliability. Significant positive structural associations are evident where the perception/interoperability (β = 0.343, p < 0.01; observed-only scenario) and governance/coordination readiness (β = 0.838, p < 0.01; baseline scenario) layers reflect stronger partial associations with workforce-regulatory reliability readiness (R2 = 0.892, p < 0.001; baseline scenario) for intelligent robot implementation in healthcare. K-means, Gaussian mixture modelling (GMM), and agglomerative (hierarchical) unsupervised learning models reveal good cluster distinctiveness (silhouette = 0.633 [k = 4]; 0.617 [k = 5]) varying between the scenarios, ranging from 0.357 to 0.514 under four robust cluster regimes, indicating distinct maturity regimes in the "RPI-MRCI-ACRI readiness" space. Robustness analyses indicated that the readiness structure remained considerably identifiable across three alternative missing data scenarios, although some boundary cases were overlapping. External validation using the World Bank and WHO's country classification indicated profound alignment with the final cluster regimes. Post-clustering recoverability analysis depicted that the latent clusters were significantly separable within the readiness space. Cluster membership in moderate to very good system readiness requires ∼30% contribution of each of the readiness indices to classify the country at the advanced to transformed phase of intelligent robotics implementation readiness in most of the scenarios under K-means and GMM. Machine learning (ML) models are run on a five-cluster specification given the main data categorisation requirement of phase "1" to phase "5." Among four supervised machine learning models, random forest and XGBoost achieved the highest (accuracy up to 0.980; area under the receiver operating characteristic curve (AUROC) ≈ 1.00) and most reliable performance (cross-validated accuracy ≈ 0.91), supporting robust separability and recoverability of the latent readiness regimes. SHAP analysis identified macro-level governance and coordination readiness index (0.111) as the most influential predictor of cluster classification under K-means and perception readiness (0.106) as the most influential under GMM. Findings uncover heterogeneous readiness clusters of countries and exhibit that digital health governance coherence, standards and interoperability, and institutional digital literacy are stronger predictors of robotics integration potential in healthcare than technological dimensions alone. The study does not provide a direct estimate of robotics deployment but offers a thoughtful, system-level framework for comparing structural enablers relevant to robotics ingestion for health-system strengthening. An adaptive monitoring framework is proposed as a conceptual future-work direction rather than an empirically validated component of the present study.
Skin-to-skin tactile stimulation plays a critical role in the neurodevelopment of preterm infants, contributing to improved physiological stability, sensory integration, and caregiver bonding. However, the delivery of consistent tactile therapy in neonatal intensive care units is often limited by the availability of trained personnel and the inherent variability of manual application. This creates a need for assistive technologies capable of reproducing clinically relevant tactile stimuli in a controlled and repeatable manner. This work presents a soft pneumatic robotic system designed to replicate clinically inspired tactile stimulation for neonatal therapy. The proposed approach integrates experimental characterization of manual tactile interactions, pneumatic system modeling, and closed-loop control. Force measurements obtained from a neonatology specialist were used to define clinically grounded stimulation levels, with mean values of 0.594 N for light stimulation and 1.267 N for moderate stimulation. These values were mapped to actuator pressure through experimentally identified linear relationships between force, pressure, and electrical current, enabling the definition of therapeutic pressure references. A 3 × 3 matrix of textile pneumatic actuators was integrated with an electro-pneumatic actuation system, pressure sensing, and current monitoring to implement the proposed therapy platform. Two control strategies were evaluated: a baseline pressure deadband controller and a refined deadband controller incorporating actuation tuning and current-based protection. Experimental results demonstrated that the refined controller increased the time within the therapeutic pressure band from 30.8% to 76.1%, while reducing mean pressure error from 0.68 kPa to 0.15 kPa and RMS error from 0.92 kPa to 0.38 kPa. Robustness tests under varying mechanical interface conditions showed stable and consistent pressure regulation performance. These results demonstrate the feasibility of translating clinically derived tactile stimulation into controlled pneumatic actuation using a soft robotic platform. By combining clinically grounded references, pressure-based closed-loop control, and safety-oriented current monitoring, the proposed system provides a reproducible and safe preclinical platform for neonatal tactile therapy and supports the development of assistive soft robotic technologies for clinically representative neonatal care environments.
The increasing adoption of service robots in hospitality has transformed service delivery and created new opportunities for enhancing operational efficiency and sustainable service practices. However, limited research has examined how different types of service robots influence customer emotions and how these emotional responses shape perceptions of sustainability and post-consumption behaviour. Drawing on the Stimulus-Organism-Response (S-O-R) framework and Human-Robot Interaction (HRI) theory, this study investigates the relationships among robot type, emotional responses, perceived sustainability, customer satisfaction and revisit intention in hotel settings. A quantitative cross-sectional survey was conducted among 238 hotel guests in Malaysia who had directly interacted with either high-interaction robots (e.g. front-desk or concierge robots) or low-interaction robots (e.g. delivery robots). Data were analysed using Partial Least Squares Structural Equation Modelling (PLS-SEM). The findings reveal that high-interaction robots significantly increase both positive and negative emotions. Positive emotions positively influence perceived sustainability and customer satisfaction, whereas negative emotions weaken these evaluations. Perceived sustainability significantly enhances customer satisfaction and revisit intention, while customer satisfaction emerges as the strongest predictor of revisit intention. Emotional responses also mediate the relationship between robot type and customer outcomes. This study extends HRI and sustainable hospitality literature by demonstrating that customer evaluations of robot-enabled services are shaped by emotional appraisal and sustainability perception rather than technological functionality alone. The findings provide practical insights for hospitality managers seeking to design emotionally engaging and sustainability-oriented robot-enabled service experiences.
In this study, we present a replay-based framework for uncertainty-aware persistent tracking of multiple advected surface patches using an autonomous marine vehicle operating in spatiotemporal-varying currents. The method combines three components: local flow estimation, covariance-aware patch-boundary propagation with intermittent boundary fusion, and mission-level scheduling over multiple patches. Each patch is represented by a polygonal boundary, whose vertices are propagated through the estimated flow field while carrying per-vertex covariance, thereby quantifying uncertainty growth during advection. A flow-aware gain-scheduled linear quadratic regulator (LQR) was designed to shape the desired surge speed to take advantage of favorable currents. When the vehicle services a patch, boundary detections are fused to reduce the active patch uncertainty, and optional local map-covariance refinement is used to reduce subsequent uncertainty regrowth in the surrounding flow field. A boundedness analysis shows that if each patch is revisited within a prescribed maximum interval, then the corresponding patch uncertainty remains uniformly bounded; a companion feasibility condition relates the allowable revisit interval to vehicle speed, service time, and tour length over the patch set. To validate the result, a data replay simulation using HF-radar currents from the San Francisco Bay region was used to demonstrate the expected bounded sawtooth uncertainty behavior under feasible revisit conditions. In addition, our proposed duration-weighted predictive scheduler outperforms nearest-patch and round-robin baselines and, in spatially separated patch configurations, achieves lower mean patch uncertainty and lower control-effort proxy than a highest-J baseline. These results indicate that combining uncertainty-aware propagation with cost-aware scheduling is a viable strategy for persistent monitoring of evolving marine surface phenomena.
The use of social robots in education has increasingly focused on how different role framings, such as tutor, peer, or novice, shape children's engagement and learning. However, most studies employ a single robot whose role is manipulated behaviourally (e.g., the same robot is used to act either as a peer or as a tutor), leaving open the question of how children respond when multiple, physically distinct robots adopt complementary roles within the same session. In this exploratory feasibility study, we introduce a blue multi-robot learning environment in which one robot acts as a tutor and the other as a novice peer. The robots supported an early word-learning task in which children learned novel object names and were later asked to name or identify them. In the proposed scenario, participants were exposed to novel words assigned to unusual or unconventional objects (e.g., dog toys and shapes made with interlocking disks), alongside familiar objects. After an initial training phase, children were asked to locate or name the novel items. Sixteen children aged 4-5 years interacted with both robots while their attention, affect, and engagement behaviours were recorded. The study characterised children's distribution of attention during a word-learning task involving a tutor and a novice peer robot, including attention to task-relevant objects and robots and its relation to attentional complexity, affect, and engagement, as well as the relation between individual characteristics (language ability, age, and media exposure) and task performance, measured as correct naming or pointing during recall. Children consistently attended to task-relevant objects throughout the session, with attention distributed across objects and robots in structured patterns. Attentional complexity co-occurred with sustained engagement and positive affect. Task performance and engagement showed little variation with age or media exposure, while baseline language ability was negatively associated with recall performance. Overall, the findings indicate that a multi-robot learning configuration is feasible and capable of supporting sustained engagement and structured attentional behaviour in young children. While the results are exploratory and limited by sample size, they provide initial evidence that complementary robot roles can be meaningfully integrated within early learning activities and motivate further systematic investigation.
Postural balance is essential for both humans and robots, as failures increase fall risk and limit robotic performance in real-world settings. Although humans and robots share fundamental balancing mechanics, biological complexity limits the isolation of individual muscle functions and direct principle transfer to robots. In this study, we use EPA-Walker, a bio-inspired robot actuated by electric motors and pneumatic artificial muscles (PAMs), as a physical platform to investigate perturbed standing. Here, we focus on the PAM-driven actuation to systematically examine the roles of muscle morphology and control, and to validate biomechanical findings in a robotic setting. To enable a clear upper body perturbation, a Control Moment Gyroscope (CMG) was integrated. We evaluated two stabilization paradigms: passive standing, in which joint compliance was tuned through static PAM pressurization, and active balancing, in which a bio-inspired ground reaction force (GRF) feedback controller generated muscle reflexes. The results showed that biarticular thigh muscles, particularly the hamstrings and rectus femoris, played the most prominent role in enhancing robustness through both morphology in the passive experiments and reflex control in the active experiments, consistent with findings from human perturbation studies. While ankle muscles, such as the soleus, were essential for stable standing mainly through their passive morphological contribution, their reflex-based action could also improve robustness in a more specific manner through center-of-pressure regulation. Activating a single muscle could significantly improve robustness beyond morphology, enabling recovery from 3   N m perturbations. Furthermore, synergistic reflex of biarticular muscles, especially hamstrings and gastrocnemius, extends the robustness to larger perturbations ( 5   N m ) . Our contribution highlights the synchronization of control strategies with the underlying morphological design through a universal sensory feedback signal, namely, GRF. These findings demonstrate the value of bio-inspired robots as testbeds to understand the potential principles underlying human motor control and support the transfer of such principles to legged robots and assistive systems.
Autonomous surface vehicles (ASVs) enable efficient in-situ data collection for large-scale ecological monitoring; however, effective environmental mapping requires planning strategies that account for not only informative measurements, but also vehicle motion constraints and limited mission resources. Existing approaches often rely on stationary environmental models or loosely coupled planning frameworks that do not fully exploit model uncertainty when generating feasible trajectories. To address these limitations, we propose a closed-loop informative path planning (IPP) framework that tightly integrates environmental modeling and trajectory generation for autonomous sampling. The proposed approach combines a nonstationary uncertainty representation using Gaussian Process Regression with Attentive Kernels (AK-GPR), uncertainty- guided adaptive sampling, and a Dubins-constrained RRT* planner to generate dynamically feasible and information-rich trajectories. The proposed framework is evaluated through staged experiments, including simulation and field validation, across representative ecological monitoring scenarios such as algal plume tracking, bathymetric mapping, and seagrass probability estimation. The results demonstrate that the proposed planner shows improved performance relative to the baseline planners in different initialization grid densities, with particularly strong performance in scenarios with limited prior information. In all environments, the proposed IPP framework achieved an average reduction of approximately 24% in mean absolute error (MAE) and 32% in predictive uncertainty compared to baseline planners, with improvements reaching up to 33% in MAE and 42%, respectively, in sparse initialization settings. These results demonstrate the benefits of a tightly coupled framework that balances uncertainty reduction and spatial coverage, enabling more efficient environmental exploration under realistic vehicle constraints.
Gait analysis is essential in clinical evaluation, biomechanics, and rehabilitation, yet conventional approaches often rely on complex marker-based systems that limit accessibility and real-time feedback. The SANE (eaSy gAit aNalysis systEm) platform was launched in 2020 and has previously been evaluated in multi-subject laboratory and clinical studies, where it demonstrated agreement with gold-standard marker-based systems for core spatiotemporal parameters and achieved high test-retest, inter-rater, and intra-rater reliability. Building on this validated foundation, the present work provides a system characterisation and robustness analysis of the latest dual-depth, real-time SANE implementation. The updated SANE system enables real-time extraction and visualisation of an expanded set of spatiotemporal and angular gait parameters-including gait speed, step and stride length, cadence, step and stride time, step width, foot angles, double support, and gait phase durations-using two depth cameras and AI-based pose estimation. A total of 80 walking trials performed by a healthy adult (39 years, 1.74 m, 73.0 kg) were acquired across four sessions over 1 week, with complete system shutdown and restart between sessions to rigorously challenge operational stability. Session-level means, within-session standard deviations, and Relative Error Measurement (REM%) between sessions were used to quantify robustness and session-to-session variability. Across all four sessions, gait outputs remained within published normative ranges for healthy adults. REM% values for primary spatiotemporal parameters were consistently below or around 5%, and standard deviations were low and comparable to values reported for marker-based systems, indicating stable, repeatable measurements over repeated restarts and days. Real-time computation and visualisation at frame rates up to approximately 80 frames per second further distinguish this system from traditional post hoc workflows. These findings characterise the operational robustness and real-time capabilities of the dual-depth SANE system in a controlled, single-participant setting and support its use as a practical, marker-less gait assessment tool in clinical and research environments, while motivating future studies on larger and pathological cohorts.
Marine and coastal ecosystems are among the least observable yet most rapidly changing environments, where climate impacts, pollution, and biodiversity loss demand monitoring and intervention at scales that manual sampling and single-robot deployments cannot sustain. This paper argues for a conceptual shift in ecological monitoring and restoration toward networked robotic ecosystems, adopting cooperative swarms of autonomous aquatic robots coupled to in-situ digital twins and human-in-the-loop supervision. We use the REMORA project as a concrete instantiation of this paradigm, outlining an integrated architecture in which multiple low-cost, persistent robots perform distributed sensing and targeted interventions, while a continuously updated digital twin fuses multi-source data to support predictive assessment, "what-if" scenario exploration, and decision support. A dedicated human-swarm interaction layer enables non-roboticist stakeholders (e.g., environmental managers and aquaculture operators) to specify intent, manage exceptions, and maintain trust without micromanaging individual units. By linking advances in swarm intelligence, sensor integration, AI-enabled analysis, and interactive decision-making, the proposed approach targets durable, scalable, and socially relevant solutions for ecological monitoring and ecosystem management, spanning aquaculture, marinas, and broader coastal environments, and provides a roadmap for translating these technologies from controlled trials to sustained field operations.
Robot navigation in shared environments may influence how robots are socially interpreted by nearby users, beyond its purely functional role in task execution. Prior work has shown that individual navigation features, such as proximity regulation, approach direction, or speed adaptation, can affect user perception; however, fewer studies have examined how these behaviors operate when combined within a single socially aware navigation framework. Understanding how integrated navigation behaviors shape user perception is important for designing robots that can operate appropriately and comfortably in human-centered environments. We conducted a perceptual evaluation of a socially aware navigation framework that integrates five socially relevant navigation behaviors, including approach direction constraints, side-consistent passing behavior, proximity-based speed adaptation, and user-oriented motion cues. A controlled in-person user study was conducted with 40 participants comparing this socially aware navigation strategy against a baseline costmap-based obstacle avoidance strategy during repeated navigation encounters with a Pepper robot in an office-like environment. The socially aware navigation strategy was associated with more positive evaluations across multiple social and movement-related dimensions, including warmth, likeability, animacy, perceived safety, naturalness of motion, and overall movement preference, while maintaining comparable ratings of competence and comfort. Open-ended responses further indicated that participants frequently described navigation behavior in experiential and social terms related to comfort, awareness, friendliness, and attentiveness. These findings suggest that combined socially informed navigation behaviors can meaningfully influence how robot motion is perceived during human-robot interaction. More broadly, the results highlight the importance of designing navigation systems that are not only collision-free and efficient, but also socially interpretable and considerate of nearby users in shared environments.
Accurate estimation of upper-limb kinematics is essential for applications such as rehabilitation assessment and assistive robotics, yet remains challenging in real-world scenarios involving occlusion and physical human interaction. While vision-based pose estimation methods have advanced significantly, their ability to recover reliable joint kinematics under such conditions remains unclear. This paper presents a systematic comparison of vision-based and wearable sensing approaches for upper-limb pose estimation during assisted dressing tasks. A monocular RGB-based convolutional neural network (CNN) and a temporally smoothed variant (CNN_temporal) are evaluated alongside a wearable IMU-based reconstruction method. All approaches are compared against an inverse kinematics (IK) reference derived from VICON motion capture data using participant-specific kinematic models. Performance is assessed using both positional error, measured via global and shoulder-centred mean per-joint position error (MPJPE), and kinematic agreement, measured via elbow flexion/extension angle error. Experiments on a real-world dataset of assisted dressing trials, involving an occupational therapist and three participants, demonstrate that IMU-based estimation provides consistently accurate and stable joint-angle reconstruction (e.g., ∼ 12 ° mean absolute error). In contrast, vision-based methods achieve reasonable positional accuracy (MPJPE ∼ 0.20 m) but exhibit substantially larger errors in joint-angle estimation (often exceeding 80 ° ), particularly under occlusion. Temporal smoothing improves positional consistency but does not preserve kinematic fidelity. These results highlight a fundamental limitation of current vision-based approaches for tasks requiring accurate joint kinematics. The findings suggest that integrating inertial sensing or incorporating biomechanical constraints may be necessary to achieve reliable pose estimation in real-world assistive scenarios.
SpatialPrompting is a practical, training-free and model-agnostic framework for 3D spatial reasoning with off-the-shelf multimodal large language models, requiring no fine-tuning or 3D-specific inputs. It adopts a pose-aware, keyframe-driven prompting strategy: we select a compact, diverse set of frames using vision-language similarity, Mahalanobis distance, field of view, and image sharpness, and then verbalize each camera pose to enable multi-view, viewpoint-aware reasoning within a single structured prompt. Evaluations on ScanQA and SQA3D using GPT-4o and Gemini show that SpatialPrompting achieves competitive performance on ScanQA compared to existing approaches, while remaining slightly below the best fine-tuned methods on SQA3D under a training-free setting. On our Complex Spatial QA (CSQA) dataset, the proposed method improves accuracy from 62 to 78 (+16) over a GPT-4o baseline and consistently outperforms query-based keyframe selection. Furthermore, experiments across multiple models-including GPT-4o, Gemini, and Qwen-demonstrate that the framework generalizes across both proprietary and open-source models. By eliminating task-specific training and enabling reuse across different models without retraining, SpatialPrompting shifts the cost from model-specific training to flexible inference, offering a scalable and effective approach for real-world spatial reasoning as multimodal models continue to evolve.
Densely populated enclosed environments, characterized by complex thermal stratification and overlapping breathing zones, represent high-risk clusters for the airborne transmission of respiratory pathogens. To address the dual challenges of physical boundary fidelity and prohibitive computational latency in traditional solvers, this study proposes a physics-validated, data-driven framework for the rapid spatiotemporal prediction of sneeze-induced pollutant dispersion. A CFD-Robotics Twin system was developed, utilizing an anthropomorphic manipulator governed by a Transformer-based Imitation Learning policy to replicate non-linear human sneezing kinematics. To ensure rigorous physical fidelity, the simulated flow fields and particle trajectories were validated through a dual-stage experimental benchmark involving anemometer measurements and robotic-arm-nozzle discharge tests. On this basis, a spatial-preserving U-ConvLSTM architecture was developed. By integrating a convolutional U-Net encoder with ConvLSTM layers and symmetric skip-connections, the model maintains absolute physical coordinates (x,y) and bypasses the topological destruction inherent in conventional flattened architectures. Evaluation via mass error ( ε mass ) and Center of Mass distance confirms strict adherence to Eulerian conservation laws. Results demonstrate that the surrogate model achieves a computational acceleration of three orders of magnitude while maintaining high structural similarity (Structural Similarity Index Measure = 0.992). Furthermore, the framework translates abstract concentration fields into practical engineering metrics, including the Dynamic Safety Radius and vertical exposure windows. This research provides a scientifically rigorous tool for real-time risk assessment and the optimization of indoor spatial layouts to mitigate pathogen exposure.
This article explores the self-efficacy of guide robot (GR) users and assesses the value of integrating GRs into organizational workflows. The study consisted of two stages, both conducted in Estonia. First, we carried out a preliminary quantitative pilot study by applying the newly developed Guide Robot User Self-Efficacy Scale (GRUSES) in controlled and uncontrolled organizational settings. This pilot stage examined users' confidence in using GRs across demographic variables, prior robot experience, and interaction contexts, and generated initial insights for the qualitative stage. In the main stage, we conducted semi-structured interviews with three stakeholder groups: GR end users, organizational administrators, and GR distributors. The preliminary survey indicated that prior robot experience was associated with higher self-efficacy, whereas age and gender differences were not statistically significant in this sample. Users' self-efficacy was lower in uncontrolled real-life use than in controlled guided scenarios, although this difference should be interpreted cautiously because the groups were not randomly assigned. The qualitative interviews, which form the core of the study, identified technical, user-related, and organizational barriers to integrating GRs into everyday workflows. Based on these findings, we propose an exploratory three-actor framework linking robot capability, user readiness, and organizational readiness. The article also provides recommendations for guide-robot deployment in service organizations. As the GRUSES instrument remains under development, the survey results are interpreted as exploratory and hypothesis-generating.
Tactile sensing provides critical contact feedback for precision robotic micro-assembly, particularly in visually occluded environments common in 3C (Computer, Communication, and Consumer Electronics) manufacturing. Magnetic tactile sensors are especially promising due to their high sensitivity, fast response, and compact structure. However, the lack of effective physics-based simulation tools remains a key bottleneck for applying magnetic tactile sensing in reinforcement learning-based assembly policy training. To address this limitation, we propose a unified framework that integrates physics-based elastomer simulation, tactile-oriented point-cloud representation learning, and Real-to-Sim cross-modal mapping. This framework establishes a physically grounded intermediate representation for magnetic tactile sensing, enabling unified modeling and alignment across domains. The elastomer deformation is modeled using a particle-based formulation and simulated via the Material Point Method (MPM), achieving 0.02 mm spatial resolution with approximately 1 GB GPU memory. To enable efficient learning on high-dimensional tactile data, we propose Point-PAMAE, a masked point-cloud autoencoder with grid-based partitioning, a multi-scale dynamic graph convolutional encoder, and a position-aware decoder. The proposed method reduces partitioning overhead by 43.54% and achieves over 88% compression with a Chamfer Distance of 0.015. Furthermore, a latent-space Real-to-Sim mapping model is developed to project real magnetic signals into the simulated deformation feature space. Experimental results demonstrate that the mapped representations preserve contact-relevant geometric structures and enable reliable cross-domain tactile alignment. These results indicate that the proposed framework provides an efficient representation for magnetic tactile sensing, supporting future Sim-to-Real deployment in precision robotic assembly.
Long-term visual localization has the potential to reduce cost and improve mapping quality in optical benthic monitoring with autonomous underwater vehicles Despite this potential, long-term visual localization in benthic environments remains understudied, primarily due to the lack of curated datasets for benchmarking. Moreover, limited georeferencing accuracy and image footprints necessitate precise geometric information for accurate ground-truthing. In this work, we address these gaps by presenting SEALOC, a curated dataset for long-term visual localization in benthic environments, and a novel method to ground-truth visual localization results for near-nadir underwater imagery. Our dataset comprises georeferenced AUV imagery from five benthic reference sites, revisited over periods up to 6 years, and includes raw and color-corrected stereo imagery, camera calibrations, and sub-decimeter registered camera poses. To our knowledge, this is the first curated underwater dataset for long-term visual localization spanning multiple sites and photic-zone habitats. Our ground-truthing method estimates 3D seafloor image footprints and links camera views with overlapping footprints, ensuring that ground-truth links reflect shared visual content. Building on this dataset and ground truth, we benchmark eight state-of-the-art visual place recognition (VPR) methods and find that Recall@K is significantly lower on our dataset than on established terrestrial and underwater benchmarks. Finally, we compare our footprint-based ground truth to a traditional location-based ground truth and show that distance-threshold ground-truthing can overestimate visual place recognition Recall@K at sites with rugged terrain and altitude variations. Together, the curated dataset, ground-truthing method, and VPR benchmark provide a stepping stone for advancing long-term visual localization in dynamic benthic environments.
Federated learning enables multiple autonomous vehicles (AVs) to collaboratively train machine learning models while preserving data privacy. However, performance degrades significantly under non-independent and identically distributed (non-IID) data conditions commonly encountered in real-world driving scenarios. Existing aggregation methods, particularly Federated Averaging (FedAvg), struggle to effectively handle client update divergence, leading to inefficient communication, unstable convergence, and increased privacy risks. To address these challenges, we propose a Dynamic Variance-Aware Federated Tuning (DV-FedTune) framework for object detection in autonomous driving systems using YOLOv12. The proposed framework dynamically adjusts client contributions through a variance-aware aggregation strategy that jointly models update consistency, variance-based diversity, and loss-guided reliability using a round-adaptive weighting mechanism. Comprehensive experiments were conducted on the KITTI object detection dataset under various non-IID federated learning configurations involving different numbers of clients, local training durations, and communication rounds. The results demonstrate that DV-FedTune consistently outperforms FedAvg, Exponential Moving Average in Federated Learning (EWHFed), and VINOEffiFedAV in terms of communication efficiency, computational cost, and model performance while maintaining stronger privacy preservation. The proposed framework achieves stable aggregation behavior and effective parameter utilization as the federated network scales. These findings indicate that DV-FedTune provides an efficient and privacy-preserving federated learning solution for distributed object detection in autonomous vehicle environments operating under heterogeneous data distributions.
Socially aware robot navigation requires robots to move among people in ways that respect human social norms, comfort, and perceived safety. Proxemics, the regulation of interpersonal space, plays a central role in this process. Applied HRI work often relies on simplified, static representations of personal space, overlooking the dynamic, asymmetric, and context-dependent nature of proxemic behavior observed in real-world interactions. The literature reflects a clear progression from simplified, concentric representations of proxemics toward increasingly context-sensitive and interaction-dependent models. This evolution indicates a growing consensus that interpersonal comfort cannot be adequately captured by a single, universal geometric shape. Instead, proxemic representations vary as a function of interaction context, task demands, cultural norms, and environmental constraints. To build on this evolution, we propose a comprehensive taxonomy of proxemics for socially aware robot navigation addressing gaps in the literature. Grounded in an extensive review of proxemics-related HRI studies published between 2020 and 2025, the taxonomy was developed through a hybrid methodology that integrates a top-down analysis of established HRI taxonomies and an AI exploratory approach with a bottom-up extraction of variables from 39 empirical studies. The resulting taxonomy systematically organizes proxemic dimensions into four interrelated clusters: Human, Robot, Environment, and Context. Together, these clusters capture the key variables shaping proxemic form (shape geometry and the scale of the personal zone boundary) and dynamics, including human activity and posture, robot design and behavior, environmental structure, task context, and the dynamic spatial properties of proxemics as captured by their metrics (the proxemics output variables). The proposed structured taxonomy of proxemics will inform the design of socially adaptive robot navigation systems and provide a foundation for future empirical research. Our analyses reveal significant gaps in current research practices, including limited consideration of interactions among multiple variables, overreliance on static laboratory settings, and insufficient integration of contextual and human-centered variables. To address these limitations, we propose future directions.
Multi-robot coordination under communication constraints is a fundamental challenge in autonomous systems, particularly in underwater environments, where low-bandwidth acoustic links restrict centralized planning and limit decentralized information propagation. In this study, we explore a "Mothership-passenger" paradigm, in which a capable underwater vehicle deploys and coordinates many lower-cost autonomous robots, enabling broad spatial coverage while reducing the mission risk and cost. We present two algorithms that tightly couple hierarchical guidance with decentralized planning to improve coordination efficiency under these constraints. The first algorithm integrates centralized and decentralized solutions to an underwater multi-robot orienteering problem, enabling globally coherent yet locally adaptive coordination among robots operating under stochastic travel costs and mission disruptions. The second algorithm introduces a learned policy on a Mothership vehicle that adaptively configures individual robot priorities, allowing globally informed guidance to shape local planning under communication and resource constraints. Through simulation and field experiments at an inland lake, we demonstrate that these contributions significantly improve coordination efficiency and robustness compared to fixed-behavior and fully decentralized baselines. These advances are directly motivated by the long-term goal of deploying coordinated robotic teams in isolated environments such as under-ice ocean regions, where the coordination and communication constraints studied here are most acute.
Imitation learning in complex, unstructured environments remains challenging due to the difficulty of grounding perception in physically meaningful representations and the need to model multimodal action distributions. Existing approaches often rely on unstructured pixel-level feature encodings or stochastic latent-variable decoders, which can lead to brittle attention in cluttered scenes. In this work, we present a novel integration of detector-based visual representations with conditional diffusion modeling (DINO + CDP) for real-world robotic imitation learning. Our framework utilizes a DINO object detection transformer to extract spatially-grounded object-query embeddings that serve as the conditioning signal for a diffusion-based policy. A primary contribution of this work is the systematic quantification of how scene complexity-measured via image entropy-affects robotic policy performance. By comparing rigid-object baselines with complex biological plant scenes, we demonstrate that organic morphology induces a measurable increase in pixel-level uncertainty that degrades standard pixel-centric models. Our results show that DINO + CDP mitigates this degradation by grounding action generation in stable object-level features. We evaluate our approach using a fully real-world manipulation dataset collected without simulation or synthetic pre-training. To isolate the impact of our architectural choices, we conduct a comparative study within a unified framework against convolutional (CNN-MLP), transformer-patch (ViT), and latent-variable (DETR + CVAE) variants. Experimental results in a robotic-arm biocell setup demonstrate that object-query-conditioned diffusion significantly improves task success rates, produces smoother trajectories, and exhibits superior robustness to high-entropy visual inputs, establishing a scalable pathway for imitation learning in challenging agricultural domains.