共找到 20 条结果
Political polarisation on structured discussion platforms such as Reddit differs fundamentally from that on broadcast platforms such as Twitter/X, yet most prior work targets the latter. We present an end-to-end framework for measuring and analysing polarisation dynamics, applied to the r/Brexit subreddit (871K submissions, November 2015 -- February 2021). We construct r/Brexit, a crowd-annotated stance dataset of 5,024 labelled submissions (inter-annotator agreement = 0.804), and train a domain-adapted BERT classifier. We introduce a continuous polarity metric that replaces discrete stance categories, revealing fine-grained opinion spectra across 27 politically-defined periods. Our analysis yields three findings: (a) future stance prediction is confounded by survivorship bias: who remains active is self-selected on engagement, not stance, biasing any longitudinal model toward a non-representative minority; (b) echo chambers are quantifiably dominant, with nearly 40% of interactions between like-minded users; (c) user current polarity is the dominant predictor of future polarity, with echo-chamber immersion as the secondary predictive signal. These findings reveal that Reddit's par
Family-firm scholarship offers competing predictions about whether family control protects or threatens market integrity. We argue that the answer depends on how family involvement is exercised. Drawing on socioemotional wealth and agency-entrenchment perspectives, we examine 8,634 U.S. firm-years (2007-2018) and link family-firm constructs to exchange-generated surveillance flags from NASDAQ SMARTS. Founder-CEO control is associated with approximately 9.5% fewer flags, family governance involvement with 21.3% more, and deep multi-generational family control with 47.1% more. The findings reveal heterogeneous identity and entrenchment mechanisms within family firms and connect family-firm governance to a market-integrity outcome previously absent from the literature.
We investigate the the itinerant ferromagnetism in a dipolar Fermi atomic system with the anisotropic spin-orbit coupling (SOC),which is traditionally explored with isotropic contact interaction.We first study the ferromagnetism transition boundaries and the properties of the ground states through the density and spin-flip distribution in momentum space, and we find that both the anisotropy and the magnitude of the SOC play an important role in this process. We propose a helpful scheme and a quantum control method which can be applied to conquering the difficulties of previous experimental observation of itinerant ferromagnetism. Our further study reveals that exotic Fermi surfaces and an abnormal phase region can exist in this system by controlling the anisotropy of SOC, which can provide constructive suggestions for the research and the application of a dipolar Fermi gas. Furthermore, we also calculate the ferromagnetism transition temperature and novel distributions in momentum space at finite temperature beyond the ground states from the perspective of experiment.
Conventional quantum routing operates under the entrenched assumption that pathfinding is a prerequisite for routing. This classical-inspired routing model imposes a restricting design option, which prevents scaling the quantumness to the network functioning. In this paper, we proposed a novel entanglement-driven routing framework that exploits multipartite entanglement complementation for enabling simultaneous 1-hop connectivity among all non-adjacent source-destination pairs. This changes the notion of ``remoteness'' in the entanglement graph, activated by entanglement. We extend this framework to inter-domain quantum networks and design a polynomial-time algorithm. Such an algorithm allows to select and parallelize multiple requests, bypassing NP-complete path discovery. Performance analysis shows the proposed routing strategy achieves up to $60\%$ hop reduction, with the algorithm enabling efficient parallelism and strong scalability in inter-domain quantum networks.
I critique a set of entrenched methodological conventions that collectively create systemic dysfunction in statistical research for clinical decisions. These include: (1) the prevalent use of hypothesis tests to compare treatments, (2) remoteness from patient care of the methods used to evaluate the accuracy of predictions of patient outcomes, (3) poor practice of meta-analysis to combine findings across studies, and (4) widespread research with incredible certitude. It appears that the dysfunction is held in place by three factors: (i) rudimentary instruction in statistical methodology received by medical students and residents, (ii) reliance of clinical researchers on consulting biostatisticians, wo act as statistical gatekeepers in evaluation of grant proposals and paper submissions, and (iii) institutional practices of research funding agencies, medical journals, and governmental bodies that regulate medical treatment. I conjecture that systemic changes are necessary to break the existing impasse, moving statistical research to a better equilibrium.
Well-documented research on physics graduate education has demonstrated long-standing issues that hinder equitable student access and participation. Addressing these challenges can be particularly difficult because they are often rooted in entrenched disciplinary and departmental cultures that tend to be rigid and resistant to change. In this work, we aim to cultivate a data-driven culture of cyclic self-reflection and action to proactively identify and address issues that affect student well-being and success in both a new physics and a long-standing astrophysics graduate program within a single institution. Drawing primarily on qualitative data (open-ended survey responses and focus-group interview data) from 15 students in both programs, we collaborated with the program leadership to identify actionable steps for improvement. In this paper, we present findings on student experiences across the two programs and discuss implications for research and practice. More broadly, this work provides a framework for graduate programs seeking to build a data-driven culture that improves student experiences.
Embodied AI is widely discussed as a job-displacement problem. The deeper risk, however, is governance lag: the inability of public institutions to keep pace with how fast the technology spreads through the physical economy. As reusable robotic platforms are combined with increasingly general AI models, embodied AI may scale across manufacturing, logistics, care, and infrastructure faster than governance systems can observe, interpret, and respond. We argue that this lag appears in three connected forms: observational, institutional, and distributive. The central policy challenge, therefore, is not automation alone, but whether governance and compliance systems can adapt before disruption becomes entrenched.
Current bias evaluations in Instruction Text-to-Speech (ITTS) often rely on univariate testing, overlooking the compositional structure of social cues. In this work, we investigate gender bias by modeling prompts as combinations of Social Status, Career stereotypes, and Persona descriptors. Analyzing open-source ITTS models, we uncover systematic interaction effects where social dimensions modulate one another, creating complex bias patterns missed by univariate baselines. Crucially, our findings indicate that these biases extend beyond surface-level artifacts, demonstrating strong associations with the semantic priors of pre-trained text encoders and the skewed distributions inherent in training data. We further demonstrate that generic diversity prompting is insufficient to override these entrenched patterns, underscoring the need for compositional analysis to diagnose latent risks in generative speech.
We critically discuss the apparent lack of logical rigor pervading the debate on quantum nonlocality. Strong convictions often prevail over rational assessment, leading to the acceptance of loose ideas that become entrenched dogmas. The lack of sound rationales and adherence to the rules of logical inference lead to widely adopted antinomies that receive little conceptual scrutiny.
The demand for Explainable AI (XAI) has triggered an explosion of methods, producing a landscape so fragmented that we now rely on surveys of surveys. Yet, fundamental challenges persist: conflicting metrics, failed sanity checks, and unresolved debates over robustness and fairness. The only consensus on how to achieve explainability is a lack of one. This has led many to point to the absence of a ground truth for defining ``the'' correct explanation as the main culprit. This position paper posits that the persistent discord in XAI arises not from an absent ground truth but from a ground truth that exists, albeit as an elusive and challenging target: the causal model that governs the relevant system. By reframing XAI queries about data, models, or decisions as causal inquiries, we prove the necessity and sufficiency of causal models for XAI. We contend that without this causal grounding, XAI remains unmoored. Ultimately, we encourage the community to converge around advanced concept and causal discovery to escape this entrenched uncertainty.
Linguistic insights may help make Large Language Model (LLM) training more efficient. We trained Meta's OPT model on the 100M word BabyLM dataset, and evaluated it on the BLiMP benchmark, which consists of 67 classes, each defined by sentence pairs that differ in a targeted syntactic or semantic rule violation. We tested the model's preference for grammatical over ungrammatical sentences across training iterations and grammatical types. In nearly one-third of the BLiMP classes, OPT fails to consistently assign a higher likelihood to grammatical sentences, even after extensive training. When it fails, it often establishes a clear (erroneous) separation of the likelihoods at an early stage of processing and sustains this to the end of our training phase. We hypothesize that this mis-categorization is costly because it creates entrenched biases that must, eventually, be reversed in order for the model to perform well. We probe this phenomenon using a mixture of qualitative (based on linguistic theory and the theory of Deep Learning) and quantitative (based on numerical testing) assessments. Our qualitative assessments indicate that only some BLiMP tests are meaningful guides. We concl
Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood. Could the self-bootstrapping nature of these methods amplify undesired behavior in models? We test seven consistency training methods on 108 model organisms: open-source models (7B--70B) fine-tuned to exhibit various forms of controlled misaligned behavior. We find that outcomes vary significantly: consistency training generally suppresses reward hacking and emergent misalignment but amplifies sycophancy. We present evidence that distribution shifts induced by the consistency labeling process, rather than variation in the selection operators, may be the primary driver of systematic alignment effects. Finally, we present a unifying theoretical framework to derive conditions under which consistency training will amplify or suppress misalignment. In total, our study establishes that consistency training is not alignment-neutral, and that its use in critical systems should be carefully audited.
This paper critically examines the machine learning (ML) modeling of humans in three case studies of well-being technologies. Through a critical technical approach, it examines how these apps were experienced in daily life (technology in use) to surface breakdowns and to identify the assumptions about the "human" body entrenched in the ML models (technology design). To address these issues, this paper applies agential realism to decenter foundational assumptions, such as body regularity and health/illness binaries, and speculates more inclusive design and ML modeling paths that acknowledge irregularity, human-system entanglements, and uncertain transitions. This work is among the first to explore the implications of decentering theories in computational modeling of human bodies and well-being, offering insights for more inclusive technologies and speculations toward posthuman-centered ML modeling.
In this paper, we document a paradigm gap in the combinatorial possibilities of verbs and aspect in Urdu: the perfective form of the -ya: kar construction (e.g. ro-ya: ki: cry-Pfv do.Pfv) is sharply ungrammatical in modern Urdu and Hindi, despite being freely attested in 19th century literature. We investigate this diachronic shift through historical text analysis, a large-scale corpus study which confirms the stark absence of perfective forms and subjective evaluation tasks with native speakers, who judge perfective examples as highly unnatural. We argue that this gap arose from a fundamental morphosyntactic conflict: the construction's requirement for a nominative subject and an invariant participle clashes with the core grammatical rule that transitive perfective assign ergative case. This conflict rendered the perfective form unstable, and its functional replacement by other constructions allowed the gap to become entrenched in the modern grammar.
Large language model (LLM)-driven AI systems may exhibit an inference failure mode we term `neural howlround,' a self-reinforcing cognitive loop where certain highly weighted inputs become dominant, leading to entrenched response patterns resistant to correction. This paper explores the mechanisms underlying this phenomenon, which is distinct from model collapse and biased salience weighting. We propose an attenuation-based correction mechanism that dynamically introduces counterbalancing adjustments and can restore adaptive reasoning, even in `locked-in' AI systems. Additionally, we discuss some other related effects arising from improperly managed reinforcement. Finally, we outline potential applications of this mitigation strategy for improving AI robustness in real-world decision-making tasks.
As large language models (LLMs) become more widely deployed, it is crucial to examine their ethical tendencies. Building on research on fairness and discrimination in AI, we investigate whether LLMs exhibit speciesist bias -- discrimination based on species membership -- and how they value non-human animals. We systematically examine this issue across three paradigms: (1) SpeciesismBench, a 1,003-item benchmark assessing recognition and moral evaluation of speciesist statements; (2) established psychological measures comparing model responses with those of human participants; (3) text-generation tasks probing elaboration on, or resistance to, speciesist rationalizations. In our benchmark, LLMs reliably detected speciesist statements but rarely condemned them, often treating speciesist attitudes as morally acceptable. On psychological measures, results were mixed: LLMs expressed slightly lower explicit speciesism than people, yet in direct trade-offs they more often chose to save one human over multiple animals. A tentative interpretation is that LLMs may weight cognitive capacity rather than species per se: when capacities were equal, they showed no species preference, and when an
Despite widespread debunking, many psychological myths remain deeply entrenched. This paper investigates whether Large Language Models (LLMs) mimic human behaviour of myth belief and explores methods to mitigate such tendencies. Using 50 popular psychological myths, we evaluate myth belief across multiple LLMs under different prompting strategies, including retrieval-augmented generation and swaying prompts. Results show that LLMs exhibit significantly lower myth belief rates than humans, though user prompting can influence responses. RAG proves effective in reducing myth belief and reveals latent debiasing potential within LLMs. Our findings contribute to the emerging field of Machine Psychology and highlight how cognitive science methods can inform the evaluation and development of LLM-based systems.
This paper examines the role of cousin marriage in shaping women's autonomy, household status, and labor supply in Pakistan. Existing research offers contradictory claims: some suggest that cousin marriage improves women's position within the household, while others argue it limits their freedoms and economic opportunities. Using data from 15,068 married women in the Pakistan Demographic and Health Survey 2017-18, this study provides new quantitative evidence. Results indicate a modest negative association between cousin marriage and women's participation in paid work, alongside a stronger link to home-based and unpaid labor. Women in cousin marriages do not appear to gain household status relative to those in non-cousin marriages and are more likely to justify spousal violence, reflecting entrenched patriarchal norms. These findings suggest that cousin marriage may reinforce traditional gender roles and constrain women's economic and social autonomy.
Recent advances in reasoning techniques have substantially improved the performance of large language models (LLMs), raising expectations for their ability to provide accurate, truthful, and reliable information. However, emerging evidence suggests that iterative reasoning may foster belief entrenchment and confirmation bias, rather than enhancing truth-seeking behavior. In this study, we propose a systematic evaluation framework for belief entrenchment in LLM reasoning by leveraging the Martingale property from Bayesian statistics. This property implies that, under rational belief updating, the expected value of future beliefs should remain equal to the current belief, i.e., belief updates are unpredictable from the current belief. We propose the unsupervised, regression-based Martingale Score to measure violations of this property, which signal deviation from the Bayesian ability of updating on new evidence. In open-ended problem domains including event forecasting, value-laden questions, and academic paper review, we find such violations to be widespread across models and setups, where the current belief positively predicts future belief updates, a phenomenon which we term belie
Multi-Agent Debate (MAD) has emerged as a promising inference scaling method for Large Language Model (LLM) reasoning. However, it frequently suffers from belief entrenchment, where agents reinforce shared errors rather than correcting them. Going beyond merely identifying this failure, we decompose it into two distinct root causes: (1) the model's biased $\textit{static initial belief}$ and (2) $\textit{homogenized debate dynamics}$ that amplify the majority view regardless of correctness. To address these sequentially, we propose $\textbf{DReaMAD}$ $($$\textbf{D}$iverse $\textbf{Rea}$soning via $\textbf{M}$ulti-$\textbf{A}$gent $\textbf{D}$ebate with Refined Prompt$)$. Our framework first rectifies the static belief via strategic prior knowledge elicitation, then reshapes the debate dynamics by enforcing perspective diversity. Validated on our new $\textit{MetaNIM Arena}$ benchmark, $\textbf{DReaMAD}$ significantly mitigates entrenchment, achieving a +9.5\% accuracy gain over ReAct prompting and a +19.0\% higher win rate than standard MAD.