Natural audiovisual perception may not be fully captured by decomposing movies into auditory and visual streams. I introduce a computational-counterfactual framework that keeps movie viewing intact while varying only AI-derived descriptions of the same clips. Using 7 Tesla movie fMRI imaging data from 176 participants, I tested whether cortical responses were better predicted by native audiovisual semantics than by a dimension-matched additive reconstruction from audio-only and video-only descriptions. The native model outperformed the matched additive baseline under content-aware purged cross-validation, with strongest gains in auditory, visual, and dorsal attention systems. Representational-similarity, feature-replacement, and content-gating analyses showed that the advantage reflected feature- and network-specific routing linked to coherent audiovisual semantic emergence rather than raw auditory-visual discrepancy. The effect survived stronger temporal purging and repeat-content exclusion, suggesting that intact movie viewing evokes cortical structure aligned with native audiovisual meaning beyond additive unimodal semantics.
Automated bird monitoring plays a crucial role in ecoinformatics and biodiversity conservation. Despite significant advancements in fine-grained visual classification, targets in complex wild environments frequently encounter interferences such as long capture distances, severe visual occlusion, and strong background noise, rendering single-modal perception highly susceptible to performance bottlenecks. To advance research in audiovisual multimodal bird classification, this paper introduces AVB81, a multimodal dataset tailored for fine-grained bird recognition. Covering 81 bird species in North America, the dataset comprises 3247 fixed-length 10 s field videos, meticulously supplemented with 5418 independent audios and 7083 high-quality static images. Based on this dataset, we systematically conduct single-modal performance evaluations and cross-paradigm audiovisual fusion experiments, establishing a comprehensive and in-depth multimodal evaluation benchmark. The experimental results demonstrate that in video classification tasks, audiovisual multimodal fusion methods significantly outperform single-modal baselines. Notably, the mid-fusion strategy based on deep semantic spaces achieves the optimal performance. Furthermore, the experiments objectively reveal the formidable challenges associated with audio recognition within videos captured in complex wild habitats. Overall, AVB81 exhibits exceptional adaptability for multimodal research, providing a highly challenging evaluation benchmark for multimodal semantic modeling in complex scenarios and laying a solid data foundation for the development and validation of next-generation intelligent ecological monitoring systems.
Semantic memory decline is increasingly recognized as an early feature of Alzheimer's disease (AD) and amnestic mild cognitive impairment (MCI), yet the neural dynamics supporting automatic and controlled semantic retrieval in healthy aging remain poorly defined. This study examined task-evoked oscillatory activity during audiovisual object recognition in young adults (YA; N = 27), healthy older adults (OA; N = 33), and individuals with amnestic MCI (N = 21). Participants judged object orientation while viewing living and nonliving images paired with congruent or incongruent characteristic sounds, allowing semantic relationships to be manipulated under implicit retrieval demands. Accuracy was high across groups, although participants with MCI showed reduced performance under the more perceptually challenging inverted conditions. Reaction times were slower in OA than YA and further slowed in MCI, with group differences varying by object animacy and semantic congruency. Event-related spectral perturbation analyses revealed distinct group-related patterns. Healthy older adults showed reduced early and increased late theta activity in frontocentral and parieto-occipital regions, consistent with delayed recruitment of control-related and perceptual-attentional processes. In contrast, MCI participants showed elevated and less condition-sensitive posterior alpha power, together with task- and condition-specific differences in frontocentral theta. The principal pattern of delayed theta recruitment in healthy aging and elevated posterior alpha in MCI was also observed in a supplementary task requiring explicit audiovisual semantic judgments. These findings provide preliminary evidence that healthy aging and amnestic MCI are associated with partly distinct patterns of task-evoked oscillatory activity during audiovisual semantic processing.
The development of audiovisual speech perception in tone-language-speaking children remains debated, and this study addressed this issue by examining Cantonese-speaking children, extending prior work on Mandarin speakers to a tone language with greater phonological complexity. Using the McGurk paradigm, 82 typically developing Cantonese-speaking children (42 girls, 40 boys; aged 4-11 years) and 21 adults (11 females, 10 males) from Hong Kong completed audiovisual speech perception tasks in quiet, moderate background noise (10 dB speech-to-noise ratio, SNR), and severe background noise (-10 dB SNR). Results showed developmental shifts in sensory dominance among Cantonese-speaking children around age 10: from relying mainly on auditory cues to integrating auditory and visual cues in quiet and 10 dB conditions, and from audiovisual-integrated to visual-dominant processing in the -10 dB condition. This developmental pattern aligns with the predictions of the Statistically Optimal Hypothesis. These findings also indicate that Cantonese-speaking children undergo this shift later than Mandarin-speaking peers, highlighting the adaptive nature of this developmental process. Understanding speech involves hearing sounds and seeing a speaker's mouth movements. This study examined how Cantonese-speaking children aged 4 to 11 develop this ability and how it changes in noisy environments. Younger children relied mainly on what they heard. Around age 10, they began to combine sound and visual information more effectively in quiet and moderately noisy settings. In very noisy conditions, both children and adults depended more on visual cues because sound became less reliable. This developmental shift occurs later than previously reported for children speaking Mandarin. The findings suggest that children gradually learn to use the most reliable source of information depending on listening conditions and highlight how language background and noise shape speech perception development.
A growing body of literature suggests that there may be deficits in audiovisual multisensory integration in schizophrenia at both the nonsocial and social levels of stimuli. The present study provides a meta-analytic synthesis of the literature on audiovisual multisensory integration in schizophrenia to evaluate if there are deficits in multisensory integration in schizophrenia relative to healthy controls (HCs) and whether these deficits are more prominent with certain types of stimuli. Initial searches identified 1552 articles for review, with 17 studies deemed appropriate for inclusion in the meta-analysis. Findings confirm that individuals with schizophrenia exhibit impaired multisensory integration relative to HCs, d = 0.60, 95% CI [0.46, 0.75]. This impairment was seen regardless of stimulus type, across both social, d = 0.71, 95% CI [0.50, 0.91], and nonsocial, d = 0.49, 95% CI [0.27, 0.71] stimuli and did not significantly differ ( p = 0.150). These findings strengthen the past literature suggesting individuals with schizophrenia do exhibit deficits in multisensory integration across multiple levels of stimuli. Clinical implications are discussed.
Hearing-impaired (HI) individuals experience substantial difficulties with speech perception in reverberant environments with multiple simultaneous talkers. Hearing aids equipped with directional microphones are designed to support listening in such situations, but their effectiveness can be compromised by changes in target talker location and listener head movements. The present study used a complex audiovisual search task to investigate performance and orienting behavior for listeners with and without hearing loss under unaided and aided conditions. The authors hypothesized that hearing loss would lead to poorer performance and more complex orienting behavior, and that directional microphone processing would affect performance and orienting behavior further due to altered spatial cues. Twenty normal-hearing (NH) and 22 HI participants completed a task requiring them to locate and identify a target narrative in a moderately reverberant audiovisual environment containing multiple competing narratives. The environment was rendered using a 64-channel loudspeaker array and virtual-reality glasses, with 15 possible target azimuths. Head and eye movements were tracked using the built-in sensors of the virtual-reality glasses. Testing was conducted under unaided conditions and aided conditions with either omnidirectional or directional microphone settings. Task performance was assessed using response accuracy, localization error, and response time. Orienting behavior was quantified using misorientation rate, head-gaze ratio, scan-path efficiency, and head and gaze orientation offsets. Compared with the NH participants, the HI participants showed significantly lower response accuracy, larger localization errors, and longer response times across all conditions. They also exhibited less direct search behavior, characterized by more frequent initial head turns away from the target and a greater reliance on eye rather than head movements. When aided, the NH participants showed reduced accuracy and increased localization errors, particularly with the directional microphone setting, and they relied slightly more on head movements. For both groups, eye-movement complexity increased under aided conditions, with the HI participants showing more complex gaze patterns. Unlike the NH participants, the HI individuals often aligned their gaze-but not their head-with the target location when responding. The results demonstrate fundamental differences in task performance and orienting behavior between HI and NH listeners in complex multi-talker environments. Although directional microphones are designed to improve speech understanding in noise, they introduce challenges in tasks that require accurate target localization. Overall, these findings highlight both the impact of hearing loss and the limitations of hearing aid directionality in complex listening environments.
Urban air pollution severely impacts pedestrians on sidewalks, yet traditional fixed-site monitoring lacks the spatial coverage needed for fine-scale exposure assessment due to high costs. This study proposes a novel machine-learning framework to predict short-term sidewalk PM2.5 and PM1 concentrations using multimodal audiovisual features extracted from self-collected street-view videos, alongside meteorological and background pollution data. Based on a mobile monitoring campaign in Shenzhen, China, we evaluated multiple models (linear regression, XGBoost, and LightGBM) across different temporal resolutions (10 s and 1 min) and validation strategies. LightGBM achieved the best performance, yielding R2 values of 0.64-0.65 for 10 s predictions and 0.80 for 1 min predictions under random cross-validation. Under rigorous spatial cross-validation, the model maintained moderate generalizability, with R2 reaching 0.41-0.48 at the 1 min resolution. Furthermore, developing a hybrid model that incorporated static geospatial context further improved the overall predictive accuracy. Variable interpretation revealed that while background PM and meteorology were dominant predictors, dynamic audio-derived features and visual indicators provided substantial additional predictive power. These findings demonstrate that integrating multimodal audiovisual sensing with ancillary data enables scalable, high-resolution estimation of street-level PM, effectively complementing conventional monitoring for urban air-quality management.
to identify the risk perceptions of professionals, patients, and caregivers/family members regarding the proper disposal of waste from home peritoneal dialysis, with a view to verifying content/topics to be considered in the development of audiovisual material intended for patients and healthcare professionals. a qualitative study was conducted, including 17 semi-structured interviews with patients and caregivers, and an online focus group with eight experts in the fields of health, communication, and waste management, and one patient who is a radiologist undergoing peritoneal dialysis. Data were analyzed using thematic analysis. doubts were reported regarding proper disposal, lack of standardized guidelines, and suggestions for selective collection and reverse logistics. Two categories emerged: "From care to waste"; and "Communication strategies of healthcare professionals". the importance of accessible educational strategies, with emphasis on risk communication and health literacy, to promote safe and sustainable practices in home care is highlighted. identificar las percepciones de riesgo de profesionales, pacientes y cuidadores/familiares con respecto a la correcta eliminación de los residuos de la diálisis peritoneal domiciliaria, para verificar el contenido/temas que deben considerarse en el desarrollo de material audiovisual destinado a pacientes y profesionales sanitarios. se realizó un estudio cualitativo que incluyó 17 entrevistas semiestructuradas con pacientes y cuidadores, un grupo focal en línea con ocho expertos en salud, comunicación y gestión de residuos, y un paciente radiólogo en diálisis peritoneal. Los datos se analizaron mediante análisis temático. se reportaron dudas sobre la eliminación adecuada de residuos, la falta de guías estandarizadas y sugerencias para la recolección selectiva y la logística inversa. Surgieron dos categorías: “De la atención al desecho”; y “Estrategias de comunicación de los profesionales de la salud”. se destaca la importancia de estrategias educativas accesibles, con énfasis en la comunicación de riesgos y la alfabetización en salud, para promover prácticas seguras y sostenibles en la atención domiciliaria.
Developmental dyslexia has been linked to atypical neural processing of the temporal dynamics of speech, but there has been disagreement concerning whether faster or slower dynamics are impaired. According to the temporal sampling (TS) theory, dyslexia arises from impaired entrainment of low-frequency neural oscillations-particularly in the delta (1-4 Hz) and theta (4-8 Hz) bands-to the rhythmic modulations of speech. This hypothesis was tested for adults with and without dyslexia using electroencephalography (EEG) and employing a rhythmic audiovisual speech paradigm previously used with dyslexic children. Participants viewed a "talking head" repeating the syllable "ba" at a 2-Hz rate. Measures were neural phase entrainment and band power in the delta, theta, beta (15-25 Hz), and low gamma (25-40 Hz) bands. Phase-amplitude coupling (PAC) and phase-phase coupling (PPC) were also assessed for delta-theta, delta-beta, theta-beta, delta-gamma, and theta-gamma interactions. Both groups exhibited significant delta- and theta-band phase entrainment; however, the two groups differed significantly in the preferred phase for the theta band. While the control group showed consistent beta- and low gamma-band phase entrainment, this was not observed for the dyslexic group. There was significantly greater delta-band power for the dyslexic group across the whole brain and in the right temporal region. These findings indicate that temporal sampling deficits appear to persist into adulthood and show a multi-timescale pattern, with altered low-frequency processing and atypical higher-frequency entrainment. The data also generate developmental hypotheses about possible compensation or reorganisation mechanisms that could be examined in future longitudinal research.
Patient comprehension of bowel preparation instructions is essential for adherence to purgative and dietary requirements. Inadequate bowel preparation reduces visualization quality, increases procedure length, and diminishes both diagnostic accuracy and patient experience. Evidence supports the use of audiovisual (AV) education to improve comprehension and adherence compared to standard written instruction. This quality improvement project implemented an AV educational intervention in an outpatient endoscopy unit and evaluated its impact on bowel preparation adequacy and patient experience. Conducted at a 238-bed community hospital in Virginia, the project used two iterative Plan-Do-Study-Act (PDSA) cycles with a total of 121 patients. Data included demographics, bowel preparation adequacy using the Ottawa Bowel Preparation Scale, purgative type, viewing status, and procedure times. Patient experience was measured using standardized Press Ganey survey data. Bowel preparation adequacy improved from 80.4% at baseline to 90.5% in the second PDSA cycle, meeting the American Gastroenterological Association (AGA) benchmark. Patient experience scores improved from 82.39% to 90.96%, exceeding the institutional benchmark of 88%. Implementation of AV education improved bowel preparation adequacy and enhanced patient experience, supporting its integration into standard pre-colonoscopy education and quality improvement practices.
Dental anxiety affects roughly 15% of the general population, yet the relative contribution of procedure type, patient age, and distraction interventions to objective autonomic arousal during dental treatment remains poorly characterised. This study investigated the effects of these three factors on electrodermal activity (EDA) in a clinical dental setting. Skin conductance was continuously recorded in 65 adult patients (age 18-29: n=23; 30-55: n=19; ≥56: n=23) undergoing dental procedures under local anaesthesia at a walk-in dental clinic. Patients were randomly assigned to audiovisual intervention (n=30) or control (n=35), and EDA was measured with a wrist-worn sensor and analyzed in STATISTICA 13.0 using nonparametric tests. Extractions produced significantly elevated mean EDA (mean (M)=10.85 μS, standard deviation (SD)=5.00) compared with non-extraction procedures (M=5.20 μS, SD=5.60; Z=4.25, P<0.001, η2=0.28). Older patients (≥56 years: median (Mdn)=2.49 µS) showed lower EDA than middle-aged patients (30-55 years: Mdn=7.81 µS; H=7.96, P=0.019, η2=0.09), consistent with physiological aging and stress resilience. Procedure duration did not correlate with EDA (r(s)=0.15, P=0.23). Anti-stress ball users showed elevated EDA during extractions (Mdn=13.12 vs. 8.60 μS, Z=-2.22, P=0.026), most plausibly attributable to bilateral hand immobilisation imposed by the study protocol. Among non-extraction procedures, augmented reality glasses (AR) were associated with the lowest median EDA (Mdn=2.35 μS, n=12), followed by music glasses (Mdn=2.73 µS, n=4) and headphones (Mdn=4.01 µS, n=3), though group differences were non-significant (H=1.42, P=0.49). Autonomic arousal was more strongly determined by procedure invasiveness (η2=0.28) than by patient age (η2=0.09), with the invasiveness effect driven primarily by the extraction subgroup. AR glasses were well tolerated and associated with the lowest EDA in non-extraction procedures, warranting evaluation in an adequately powered trial. Adequately powered randomised controlled trials restricted to single procedure types and incorporating validated anxiety scales are required before clinical recommendations can be made.
In noisy environments, visible speech articulations improve listening comprehension. The benefit derives from several sources, including articulatory timing and shape. Recent research has shown that visual cortex encodes a categorical representation of articulatory features and that visual speech can benefit both acoustic and phonetic feature processing separately. The present study advances the hypothesis that the shape of the articulators specifically influences the categorization of auditory speech in terms of its phonetic features. We tested this by linearly modeling electroencephalographic responses to natural, continuous speech (in noise) in terms of the acoustic and articulatory features of the speech. We compared the performance of these models in conditions where the speech was accompanied by a natural video of the speaker with their mouth visible, and a video where their mouth was covered by a dynamic ellipse obscuring articulatory shape but preserving dynamics. The dynamic mask reduced comprehension, neural processing of phonetic features, the associated multisensory benefits, and indices of visual-only linguistic processing over occipital scalp. Our findings support substantial visual involvement in speech comprehension, derived largely from the shape of the articulators. They also corroborate several proposals involving audiovisual speech processing hierarchy and the nature of the information contained in visible speech.
[This corrects the article DOI: 10.3389/fnins.2025.1536688.].
By 9 months, infants recognize own-race faces from static images, but struggle when familiarized with naturalistic audiovisual (AV) talking faces, despite daily exposure to dynamic, speaking faces. This study examined whether AV speech or AV mouth chewing motion disrupts face recognition and whether selective attention to the mouth or cognitive load mediates this effect. We recorded the eye gaze of 10-month-olds (N = 133) and 20-month-olds (N = 125) during a Visual Paired Comparison task in three conditions: Talking (AV, speech, dynamic), Chewing (AV, no speech, dynamic), and Static (no AV, no speech, static). Infants focused more on the mouth in Chewing, the eyes in Static, and equally on both in Talking. However, recognition occurred only in the Static and Chewing conditions, independently of age. These findings suggest that AV speech, not just motion, increases cognitive load and disrupts face memorization regardless of selective attention, impacting early face processing.
Human creativity is often studied in relation to idea-generation and problem-solving; yet a growing body of evidence suggests that creativity-related traits may also shape how we perceive and understand the world around us. In this study, we examined whether individual differences in creative disposition are associated with differences in crossmodal associations triggered by musical stimuli. A total of 97 healthy adult participants completed a crossmodal association task wherein they rated the extent to which short musical melodies could be associated with illustrations of everyday objects, followed by questionnaires assessing creative personal identity, creative self-efficacy, openness to experience, and creative activities. Using a Gaussian Mixture model-based clustering approach, we identified two participant profiles characterised by consistently high versus low creativity-related scores. By means of a Bayesian Zero-Inflated Beta-distributed Mixed Model, credible evidence was obtained that individuals from the former cluster, characterised by a broadly developed creative profile, gave higher association scores. Despite some uncertainty regarding its practical relevance, the effect showed a 95.82% probability of existence, suggesting that creativity-related traits may influence the crossmodal associations between musical and visual percepts.
Children and adolescents with neurodevelopmental disorders (NDDs) frequently experience high dental anxiety and behavioral challenges that complicate routine dental care. Nonpharmacological behavior management strategies have been increasingly implemented; however, their effectiveness has not been comprehensively quantified. This systematic review and meta-analysis aimed to evaluate the effects of nonpharmacological interventions on anxiety and cooperation during dental procedures in pediatric patients with NDDs. A systematic search was conducted in 5 electronic databases: PubMed, EMBASE, CENTRAL, Scopus, and Web of Science from database inception until February 2, 2026. Randomized clinical trials, nonrandomized interventional studies, and observational studies were included, focusing on children and adolescents (age ≤ 21 years) diagnosed with NDDs, who received nonpharmacological behavior management techniques-including sensory-adapted dental environment (SADE), distraction techniques, desensitization and other types of audiovisual interventions-during or before dental treatment. The Risk of Bias 2 (RoB2) tool for randomized controlled trials (RCTs), the ROBINS-I tool for nonrandomized studies of interventions, and the MINORS-tool for single-armed studies were used to assess the risk of bias. A total of 42 studies involving 1,809 NDD patients were included, with 22 being synthesized for meta-analysis. Audiovisual distractions and SADE were associated with reduced anxiety levels when measured with objective tools (standardized mean difference: -0.70, 95% CI: -1.11 to -0.29), Venham Scale (mean difference: -0.55, 95% CI: -0.91 to -0.18), and improved cooperation according to the Frankl Scale (mean difference: 0.45, 95% CI: 0.22-0.68). Most outcomes were classified as having moderate to very low certainty of evidence. Audiovisual distractions and SADEs may serve as effective approaches to both improve cooperation and reduce anxiety during dental care for children and adolescents with NDDs. Clinicians should consider individualized strategies, and continued research is needed to inform dental care protocols for this vulnerable population in everyday practice.
Childhood diarrhea remains a leading cause of morbidity and mortality in low- and middle-income countries. Educational interventions targeting caregivers may promote preventive behaviors, but the overall evidence base has not yet been comprehensively synthesized. This systematic review evaluates the effectiveness and types of community-based educational interventions in caregivers of children under five years of age to prevent diarrhea. We systematically searched MEDLINE/PubMed, EMBASE, Cochrane Library and LILACS with relevant keywords and MeSH terms from inception to December 2025. Of 1,917 records screened, 15 controlled trials (1985-2018) from Asia, Africa, and Latin America met inclusion criteria. Interventions included home visits, audiovisual materials, group education sessions, and behavior change communication (BCC) strategies. The primary outcomes were diarrhea incidence and prevention and hygiene-related behavioral changes. Risk of bias was assessed using RoB-2 for all studies. Most interventions reported positive outcomes, such as reduced diarrhea incidence and improved caregiver knowledge and enhanced hygiene behaviors. Multicomponent approaches, especially those combining home visits with audiovisual tools and BCC, were particularly effective. However, heterogeneity in follow-up periods, population characteristics, and settings limited comparability. Several studies had high or unclear risk of bias, often due to inadequate reporting of randomization or lack of blinding. Community-based educational strategies show potential for reducing childhood diarrhea, particularly when tailored to local contexts and combined multicomponent approaches. Future high-quality, standardized RCTs are needed to build a more robust evidence base and guide policy and program development.
Associative equivalence learning, the ability to form connections between different stimuli based on shared outcomes, plays a fundamental role in human cognition. In this study, we investigated how stimulus modality (visual vs. audiovisual) and the semantic content of the visual stimuli affect performance in associative equivalence learning. We applied four tests (FaceFish, Polygon, SoundFace, and SoundPolygon) based on the Rutgers Acquired Equivalence Test, in a sample of 117 healthy adult participants. The tests differed in modality and semantic content of visual stimuli. All tests were divided into an acquisition phase, where participants learned associations based on shared outcomes, and a test phase, where they retrieved the learned associations and applied the acquired equivalence to new stimulus pairs. Our results revealed that visual semantic content had a stronger and more consistent effect on performance than stimulus modality, particularly in the acquisition phase. Audiovisual presentation offered some facilitative effects (lower error ratios), especially when combined with visual stimuli with higher semantic content, but did not compensate for the limitations of simplified stimuli. These findings support the critical role of visual stimulus features in shaping learning efficiency and suggest that the benefits of multisensory input may depend on the semantic content of visual stimuli. The results also validate previous findings obtained from smaller samples and highlight directions for future research, including developmental, multisensory, and neurophysiological investigations of associative learning.
In the era of digital interconnection, cultural soft power extends beyond macro-policy communication, increasingly manifesting in the subjective perceptions and group identity evaluations formed by netizens during daily online interactions. This study explores the psychological mechanisms linking domestic netizens' digital engagement with their perceived global influence of Chinese traditional festivals, collective self-esteem, and sense of digital cultural empowerment. Using a cross-sectional survey (N = 1,540), we measured digital interaction levels, audiovisual modality dependency, and core cultural-psychological indicators. Data were statistically analyzed via confirmatory factor analysis, multi-group comparisons, and the PROCESS macro. Results indicated that digital engagement was significantly and positively associated with perceived global influence. Moreover, perceived global influence significantly mediated the relationship between digital interaction and collective self-esteem. The data also supported a serial indirect path from digital interaction to digital cultural empowerment, sequentially mediated by perceived global influence and collective self-esteem. Notably, neither audiovisual modality dependency nor participants' overseas experience significantly moderated the core association between digital engagement and perceived global influence. The findings support a statistical association consistent with the "cognition-emotion-empowerment" serial framework, suggesting that cultural identity formation and psychological empowerment in the digital age are more closely associated with the active frequency of everyday online interaction than with specific media formats (e.g., short videos vs. text) or direct intercultural experiences. Highlighting the characteristics of "mediated migration," this study provides empirical, user-centric insights for optimizing the global digital communication strategies of traditional cultural heritage. Finally, because this study relies on cross-sectional survey data, all highlighted paths represent statistical associations rather than strict causal mechanisms.
The use of #pathology and #pathologist is increasing on TikTok, a short-form video platform with over 1 billion users. As of July 2024, #pathology had over 21 000 posts and 468 million views, while #pathologist had over 3000 posts. It could be a useful tool for education and recruitment if the content is accurate and engaging. To assess the accuracy, engagement, and educational value of this growing content, we conducted a cross-sectional study analyzing 105 English-language TikTok videos identified using these keywords over a 72-hour period. Videos were evaluated using the Patient Education Assessment Tool for audiovisual material (PEMAT-AV) for audiovisual quality, the Global Quality Scale (GQS) for overall content quality, and a modified JAMA benchmark score for information accountability. Additionally, a harm-benefit score categorized educational impact. Statistical analysis revealed that educational content demonstrated significantly higher quality scores on GQS (P = .0001) and PEMAT-AV (P = .0007) than other content types. In contrast, medical profile type was associated with higher PEMAT-AV (P = .0147), JAMA (P = .0332), and harm-benefit (P = .0004) scores. While the volume of pathology-specific content is currently limited, the high average engagement metrics-298 100 followers, 1.5 million views, and 67 884 likes per video-indicate substantial user interest. The higher-quality scores for educational and medical content suggest that pathology content creators are generally knowledgeable and accurately represent the field on TikTok, highlighting TikTok's potential as a valuable platform for disseminating reliable pathology-related information.