College students experience many stressors, resulting in high levels of anxiety and depression. Wearable technology provides unobtrusive sensor data that can be used for the early detection of mental illness. However, current research is limited concerning the variety of psychological instruments administered, physiological modalities, and time series parameters. In this research, we collect the Student Mental and Environmental Health (StudentMEH) Fitbit dataset from students at our institution during the pandemic. We provide a comprehensive assessment of the ability of predictive machine learning models to screen for depression, anxiety, and stress using different Fitbit modalities. Our findings indicate potential in physiological modalities such as heart rate and sleep to screen for mental illness with the F1 scores as high as 0.79 for anxiety, the former modality reaching 0.77 for stress screening, and the latter modality achieving 0.78 for depression. This research highlights the potential of wearable devices to support continuous mental health monitoring, the importance of identifying best data aggregation levels and appropriate modalities for screening for different mental ai
Background Commercially available activity monitors, such as the Fitbit , may encourage physical activity. However, the accuracy of the Fitbit in older adults remains unknown. This study aimed to determine (1) the criterion validity of Fitbit step counts compared to visual count and ActiGraph accelerometer step counts and (2) the accuracy of ActiGraph step counts compared to visual count in community-dwelling older people. Methods Thirty-two community-dwelling adults aged over 60 wore Fitbit and ActiGraph devices simultaneously during a 2 min walk test (2MWT) and then during waking hours over a 7-day period. A physiotherapist counted the steps taken during the 2MWT. Results There was excellent agreement between Fitbit and visually counted steps (intraclass correlation coefficient (ICC 2,1 )=0.88, 95% CI 0.76 to 0.94) from the 2MWT, and good agreement between Fitbit and ActiGraph (ICC 2,1 =0.66, 95% CI 0.41 to 0.82), and between ActiGraph and visually counted steps (ICC 2,1 =0.60, 95% CI 0.33 to 0.79). There was excellent agreement between the Fitbit and ActiGraph in average steps/day over 7 days (ICC 2,1 =0.94, 95% CI 0.88 to 0.97). Percentage agreement was closest for Fitbit steps compared to visual count (mean 0%, SD 4%) and least for Fitbit average steps/day compared to the ActiGraph (mean 13%, SD 25%). Conclusions The Fitbit accurately tracked steps during the 2MWT, but the ActiGraph appeared to underestimate steps. There was strong agreement between Fitbit and ActiGraph counted steps. The Fitbit tracker is sufficiently accurate to be used among community-dwelling older adults to monitor and give feedback on step counts.
OBJECTIVES: To examine the validity and reliability of the Fitbit Flex against direct observation for measuring steps in the laboratory and against the Actigraph for step counts in free-living conditions and for moderate-to-vigorous physical activity (MVPA) and activity energy expenditure (AEE) overall. METHODS: Twenty-five adults (12 females, 13 males) wore a Fitbit Flex and an Actigraph GT3X+ during a laboratory based protocol (including walking, incline walking, running and stepping) and free-living conditions during a single day period to examine measurement of steps, AEE and MVPA. Twenty-four of the participants attended a second session using the same protocol. RESULTS: Intraclass correlations (ICC) for test-retest reliability of the Fitbit Flex were strong for walking (ICC = 0.57), moderate for stair stepping (ICC = 0.34), and weak for incline walking (ICC = 0.22) and jogging (ICC = 0.26). The Fitbit significantly undercounted walking steps in the laboratory (absolute proportional difference: 21.2%, 95%CI 13.0-29.4%), but it was more accurate, despite slightly over counting, for both jogging (6.4%, 95%CI 3.7-9.0%) and stair stepping (15.5%, 95%CI 10.1-20.9%). The Fitbit had higher coefficients of variation (Cv) for step counts compared to direct observation and the Actigraph. In free-living conditions, the average MVPA minutes were lower in the Fitbit (35.4 minutes) compared to the Actigraph (54.6 minutes), but AEE was greater from the Fitbit (808.1 calories) versus the Actigraph (538.9 calories). The coefficients of variation were similar for AEE for the Actigraph (Cv = 36.0) and Fitbit (Cv = 35.0), but lower in the Actigraph (Cv = 25.5) for MVPA against the Fitbit (Cv = 32.7). CONCLUSION: The Fitbit Flex has moderate validity for measuring physical activity relative to direct observation and the Actigraph. Test-rest reliability of the Fitbit was dependant on activity type and had greater variation between sessions compared to the Actigraph. Physical activity surveillance studies using the Fitbit Flex should consider the potential effect of measurement reactivity and undercounting of steps.
BACKGROUND: Recent advances in sensor technologies have promoted the use of consumer-based accelerometers such as Fitbit Flex in epidemiological and clinical research; however, the validity of the Fitbit Flex in measuring sedentary behavior (SED) and physical activity (PA) has not been fully determined against previously validated research-grade accelerometers such as ActiGraph GT3X+. Therefore, the purpose of this study was to examine the concurrent validity of the Fitbit Flex against ActiGraph GT3X+ in a free-living condition. METHODS: A total of 65 participants (age: M = 42, SD = 14 years, female: 72%) each wore a Fitbit Flex and GT3X+ for seven consecutive days. After excluding sleep and non-wear time, time spent (min/day) in SED and moderate-to-vigorous PA (MVPA) were estimated using various cut-points for GT3X+ and brand-specific algorithms for Fitbit, respectively. Repeated measures one-way ANOVA and mean absolute percent errors (MAPE) served to examine differences and measurement errors in SED and MVPA estimates between Fitbit Flex and GT3X+, respectively. Pearson and Spearman correlations and Bland-Altman (BA) plots were used to evaluate the association and potential systematic bias between Fitbit Flex and GT3X+. PROC MIXED procedure in SAS was used to examine the equivalence (i.e., the 90% confidence interval with ±10% equivalence zone) between the devices. RESULTS: Fitbit Flex produced similar SED and low MAPE (mean difference [MD] = 37 min/day, P = .21, MAPE = 6.8%), but significantly higher MVPA and relatively large MAPE (MD = 59-77 min/day, P < .0001, MAPE = 56.6-74.3%) compared with the estimates from GT3X+ using three different cut-points. The correlations between Fitbit Flex and GT3X+ were consistently higher for SED (r = 0.90, ρ = 0.86, P < .01), but weaker for MVPA (r = 0.65-0.76, ρ = 0.69-0.79, P < .01). BA plots revealed that there is no apparent bias in estimating SED. CONCLUSION: In comparison with the GT3X+ accelerometer, the Fitbit Flex provided comparatively accurate estimates of SED, but the Fitbit Flex overestimated MVPA under free-living conditions. Future investigations using the Fitbit Flex should be aware of present findings.
Epidemic levels of inactivity are associated with chronic diseases and rising healthcare costs. To address this, accelerometers have been used to track levels of activity. The Fitbit and Fitbit Ultra are some of the newest commercially available accelerometers. The purpose of this study was to determine the reliability and validity of the Fitbit and Fitbit Ultra. Twenty-three subjects were fitted with two Fitbit and Fitbit Ultra accelerometers, two industry-standard accelerometers and an indirect calorimetry device. Subjects participated in 6-min bouts of treadmill walking, jogging and stair stepping. Results indicate the Fitbit and Fitbit Ultra are reliable and valid for activity monitoring (step counts) and determining energy expenditure while walking and jogging without an incline. The Fitbit and standard accelerometers under-estimated energy expenditure compared to indirect calorimetry for inclined activities. These data suggest the Fitbit and Fitbit Ultra are reliable and valid for monitoring over-ground energy expenditure.
Wearable devices are widely used for heart rate (HR) monitoring, yet their accuracy across diverse body compositions and skin tones remains uncertain. This study evaluated four wrist worn devices (Apple, Fitbit, Samsung, Garmin) in 58 Hispanic adults with Fitzpatrick skin types III to V during a cycling protocol alternating moderate (0.64 to 0.76 HRmax) and vigorous (0.77 to 0.95 HRmax) intensities. Criterion HR was obtained using a Polar H10 ECG, and accuracy was assessed using mean absolute error, mean absolute percentage error (MAPE), bias, and intraclass correlation coefficients. All devices showed significant deviation from criterion measures. Apple and Garmin demonstrated the lowest error, whereas Fitbit and Samsung exhibited greater inaccuracies. Higher BMI and darker skin tones were associated with increased MAPE. These biases disproportionately affect higher risk populations, underscoring the need for improved algorithms to ensure equitable health monitoring.
Smartwatches are widely used to estimate caloric expenditure for weight management, clinical decision making, and public health monitoring. These devices combine photoplethysmography, accelerometry, and proprietary algorithms. However, prior studies report substantial error, and the influence of moderators such as skin tone and body fat percentage (BF) remains underexamined. This study tested whether smartwatch brand, BF, and Fitzpatrick skin type (III to V) predict caloric expenditure error relative to indirect calorimetry. Fifty eight Hispanic adults completed a single laboratory visit including a ten minute recumbent cycling protocol with alternating two minute moderate and vigorous intensity intervals, bracketed by rest and recovery. Participants wore four consumer devices: Apple Watch Series 8, Fitbit Sense 2, Samsung Galaxy Watch 5, and Garmin Forerunner 955. Energy expenditure was measured using a COSMED K5 metabolic system. After device specific data quality filtering, valid participant device pairings ranged from 44 to 52 per brand. One sample tests showed significant mean bias for three devices: Apple, Garmin, and Samsung. Fitbit showed no significant overall bias, althou
Sleep traits are shaped by genetic and environmental factors and may influence many health conditions. The All of Us Research Program, which includes EHR, physical measurements, genomic data, and wearable data across ancestry groups, provides an opportunity to study genetic and non-genetic contributors to sleep-related health outcomes. We examined associations between genetic predispositions to chronotype, sleep duration, and short sleep and health outcomes across ancestries, as well as the role of measured sleep duration. We used All of Us genome-wide association study results, including ancestry-specific and meta-analyses for 3,414 phenotypes, to identify phenotypes associated with 455 sleep-related SNPs. Cross-sectional and longitudinal analyses (n = 212,529) evaluated associations between polygenic risk scores (PRS) and anthropometric and metabolic measures from EHR. A subgroup analysis (n = 7,655) assessed sleep duration using Fitbit data. Across six ancestry groups, SNP analysis identified 61 phenotypes linked to 29 sleep-trait-associated SNPs. The chronotype SNP rs1421085 in FTO showed the strongest associations with obesity, diabetes, and cardiovascular conditions, mainly i
Low physical activity is a known risk factor for major depressive disorder (MDD), but changes in activity before a first clinical diagnosis remain unclear, especially using long-term objective measurements. This study characterized trajectories of wearable-measured physical activity during the year preceding incident MDD diagnosis. We conducted a retrospective nested case-control study using linked electronic health record and Fitbit data from the All of Us Research Program. Adults with at least 6 months of valid wearable data in the year before diagnosis were eligible. Incident MDD cases were matched to controls on age, sex, body mass index, and index time (up to four controls per case). Daily step counts and moderate-to-vigorous physical activity (MVPA) were aggregated into monthly averages. Linear mixed-effects models compared trajectories from 12 months before diagnosis to diagnosis. Within cases, contrasts identified when activity first significantly deviated from levels 12 months prior. The cohort included 4,104 participants (829 cases and 3,275 controls; 81.7% women; median age 48.4 years). Compared with controls, cases showed consistently lower activity and significant down
Total knee arthroplasty (TKA) and total hip arthroplasty (THA) improve symptoms in end-stage osteoarthritis, yet long-term objective characterization of perioperative physical activity trajectories remains limited. We conducted a longitudinal observational study within the All of Us Research Program dataset, linking electronic health records with continuous Fitbit-derived step count data over a four-year perioperative window (two years before and two years after arthroplasty). Piecewise linear mixed-effects models characterized preoperative declines and postoperative recovery trajectories, and time-to-recovery was evaluated using Kaplan-Meier curves and Cox proportional hazards models under remote and immediate preoperative physical activity baseline definitions. Among 238 participants (147 TKA; 91 THA), both procedures exhibited progressive preoperative decline with distinct procedure-specific patterns and staged postoperative recovery: rapid improvement during weeks 1-6, decelerating gains through weeks 7-19/20, and subsequent stabilization through week 104. Recovery to remote and immediate baselines differed in timing (median 22 vs 13 weeks) and associated predictors. Higher imm
Language models excel at diagnostic assessments on curated medical case-studies and vignettes, performing on par with, or better than, clinical professionals. However, existing studies focus on complex scenarios with rich context making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. We deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants (N=13,917) to interact with five AI agents. This corpus captures diverse communication and a realistic distribution of illnesses from a real world population. A subset of 1,228 participants reported a clinician-provided diagnosis, and 517 of these were further evaluated by a panel of clinicians during over 250 hours of annotation. SymptomAI DDx were significantly more accurate (OR = 2.56, p < 0.001) than those from independent clinicians given the same dialogue in a blinded randomized comparison. Moreover, agentic strategies which conduct a dedicated symptom interview that elicit additional symptom information before providing a diagnosis, perform sub
The incorporation of generative artificial intelligence into personal health applications presents a transformative opportunity for personalized, data-driven health and fitness guidance, yet also poses challenges related to user safety, model accuracy, and personal privacy. To address these challenges, a novel, principle-based framework was developed and validated for the systematic evaluation of LLMs applied to personal health and wellness. First, the development of the Fitbit Insights explorer, a large language model (LLM)-powered system designed to help users interpret their personal health data, is described. Subsequently, the safety, helpfulness, accuracy, relevance, and personalization (SHARP) principle-based framework is introduced as an end-to-end operational methodology that integrates comprehensive evaluation techniques including human evaluation by generalists and clinical specialists, autorater assessments, and adversarial testing, into an iterative development lifecycle. Through the application of this framework to the Fitbit Insights explorer in a staged deployment involving over 13,000 consented users, challenges not apparent during initial testing were systematicall
Consistent physical inactivity poses a major global health challenge. Mobile health (mHealth) interventions, particularly Just-in-Time Adaptive Interventions (JITAIs), offer a promising avenue for scalable, personalized physical activity (PA) promotion. However, developing and evaluating such interventions at scale, while integrating robust behavioral science, presents methodological hurdles. The PEARL study was the first large-scale, four-arm randomized controlled trial to assess a reinforcement learning (RL) algorithm, informed by health behavior change theory, to personalize the content and timing of PA nudges via a Fitbit app. We enrolled and randomized 13,463 Fitbit users into four study arms: control, random, fixed, and RL. The control arm received no nudges. The other three arms received nudges from a bank of 155 nudges based on behavioral science principles. The random arm received nudges selected at random. The fixed arm received nudges based on a pre-set logic from survey responses about PA barriers. The RL group received nudges selected by an adaptive RL algorithm. We included 7,711 participants in primary analyses (mean age 42.1, 86.3% female, baseline steps 5,618.2). W
Cardio Load, introduced by Google in 2024, is a measure of cardiovascular work (also known as training load) resulting from all the user's activities across the day. It is based on heart rate reserve and captures both activity intensity and duration. Thanks to feedback from users and internal research, we introduce adaptive and personalized targets which will be set weekly. This feature will be available in the Public Preview of the Fitbit app after September 2025. This white paper provides a comprehensive overview of Cardio Load (CL) and how weekly CL targets are established, with examples shown to illustrate the effect of varying CL on the weekly target. We compare Cardio Load and Active Zone Minutes (AZMs), highlighting their distinct purposes, i.e. AZMs for health guidelines and CL for performance measurement. We highlight that CL is accumulated both during active workouts and incidental daily activities, so users are able top-up their CL score with small bouts of activity across the day.
This paper presents striking new data about the scale of Google's involvement in the global digital and corporate landscape, head and shoulders above the other big tech firms. While public attention and some antitrust scrutiny has focused on these firms' mergers and acquisitions (M&A) activities, Google has also been amassing an empire of more than 6,000 companies which it has acquired, supported or invested in, across the digital economy and beyond. The power of Google over the digital markets infrastructure and dynamics is likely greater than previously documented. We also trace the antitrust failures that have led to this state of affairs. In particular, we explore the role of neoclassical economics practiced both inside the regulatory authorities and by consultants on the outside. Their unduly narrow approach has obscured harms from vertical and conglomerate concentrations of market power and erected ever higher hurdles for enforcement action, as we demonstrate using examples of the failure to intervene in the Google/DoubleClick and Google/Fitbit mergers. Our lessons from the past failures can inform the current approach towards one of the biggest ever big tech M&A deal
Missing data is among the most prominent challenges in the analysis of physical activity (PA) data collected from wearable devices, with the threat of nonignorabile missingness arising when patterns of device wear relate to underlying activity patterns. We offer a rigorous consideration of assumptions about missing data mechanisms in the context of the common modeling paradigm of state space models with a finite, meaningful, set of underlying PA states. Focusing in particular on hidden Markov models, we identify inherent limitations in the presence of missing data when covariates are required to satisfy common missing data assumptions. In response to this limitation, we propose a Bayesian non-homogeneous state space model that can accommodate covariate dependence in the transitions between latent activity states, which in this case relates to whether patients' routine behavior can inform how they transition between PA states and thus support imputation of missing PA data. We show the benefits of the proposed model for missing data imputation and inference for relevant PA summaries. Our development advances analytic capacity to confront the ubiquitous challenge of missing data when
We present PhysioLLM, an interactive system that leverages large language models (LLMs) to provide personalized health understanding and exploration by integrating physiological data from wearables with contextual information. Unlike commercial health apps for wearables, our system offers a comprehensive statistical analysis component that discovers correlations and trends in user data, allowing users to ask questions in natural language and receive generated personalized insights, and guides them to develop actionable goals. As a case study, we focus on improving sleep quality, given its measurability through physiological data and its importance to general well-being. Through a user study with 24 Fitbit watch users, we demonstrate that PhysioLLM outperforms both the Fitbit App alone and a generic LLM chatbot in facilitating a deeper, personalized understanding of health data and supporting actionable steps toward personal health goals.
Background: The rise of mobile technology and health apps has increased the use of person-generated health data (PGHD). PGHD holds significant potential for clinical decision-making but remains challenging to manage. Objective: This study aimed to enhance the clinical utilization of wearable health data by developing the Validation and Inspection Tool for Armband-Based Lifelog Data (VITAL), a pipeline for data integration, visualization, and quality management, and evaluating its usability. Methods: The study followed a structured process of requirement gathering, tool implementation, and usability evaluation. Requirements were identified through input from four clinicians. Wearable health data from Samsung, Apple, Fitbit, and Xiaomi devices were integrated into a standardized dataframe at 10-minute intervals, focusing on biometrics, activity, and sleep. Features of VITAL support data integration, visualization, and quality management. Usability evaluation involved seven clinicians performing tasks, completing the Unified Theory of Acceptance and Use of Technology (UTAUT) survey, and participating in interviews to identify usability issues. Results: VITAL successfully integrated we
Wearable sensors, such as smartwatches, have become increasingly prevalent across domains like healthcare, sports, and education, enabling continuous monitoring of physiological and behavioral data. In the context of education, these technologies offer new opportunities to study cognitive and affective processes such as engagement, attention, and performance. However, the lack of scalable, synchronized, and high-resolution tools for multimodal data acquisition continues to be a significant barrier to the widespread adoption of Multimodal Learning Analytics in real-world educational settings. This paper presents two complementary tools developed to address these challenges: Watch-DMLT, a data acquisition application for Fitbit Sense 2 smartwatches that enables real-time, multi-user monitoring of physiological and motion signals; and ViSeDOPS, a dashboard-based visualization system for analyzing synchronized multimodal data collected during oral presentations. We report on a classroom deployment involving 65 students and up to 16 smartwatches, where data streams including heart rate, motion, gaze, video, and contextual annotations were captured and analyzed. Results demonstrate the f
The autonomic nervous system (ANS) is activated during stress, which can have negative effects on cardiovascular health, sleep, the immune system, and mental health. While there are ways to quantify ANS activity in laboratories, there is a paucity of methods that have been validated in real-world contexts. We present the Fitbit Body Response Algorithm, an approach to continuous remote measurement of ANS activation through widely available remote wrist-based sensors. The design was validated via two experiments, a Trier Social Stress Test (n = 45) and ecological momentary assessments (EMA) of perceived stress (n=87), providing both controlled and ecologically valid test data. Model performance predicting perceived stress when using all available sensor modalities was consistent with expectations (accuracy=0.85) and outperformed models with access to only a subset of the signals. We discuss and address challenges to sensing that arise in real world settings that do not present in conventional lab environments.