Pretrained deep learning model sharing holds tremendous value for researchers and enterprises alike. It allows them to apply deep learning by fine-tuning models at a fraction of the cost of training a brand-new model. However, model sharing exposes end-users to cyber threats that leverage the models for malicious purposes. Attackers can use model sharing by hiding self-executing malware inside neural network parameters and then distributing them for unsuspecting users to unknowingly directly execute them, or indirectly as a dependency in another software. In this work, we propose NeuPerm, a simple yet effec- tive way of disrupting such malware by leveraging the theoretical property of neural network permutation symmetry. Our method has little to no effect on model performance at all, and we empirically show it successfully disrupts state-of-the-art attacks that were only previously addressed using quantization, a highly complex process. NeuPerm is shown to work on LLMs, a feat that no other previous similar works have achieved. The source code is available at https://github.com/danigil/NeuPerm.git.
Foreign information operations conducted by Russian and Chinese actors exploit the United States' permissive information environment. These campaigns threaten democratic institutions and the broader Westphalian model. Yet, existing detection and mitigation strategies often fail to identify active information campaigns in real time. This paper introduces ChestyBot, a pragmatics-based language model that detects unlabeled foreign malign influence tweets with up to 98.34% accuracy. The model supports a novel framework to disrupt foreign influence operations in their formative stages.
We study how targeted content injection can strategically disrupt social networks. Using the Friedkin-Johnsen (FJ) model, we utilize a measure of social dissensus and show that (i) simple FJ variants cannot significantly perturb the network, (ii) extending the model enables valid graph structures where disruption at equilibrium exceeds the initial state, and (iii) altering an individual's inherent opinion can maximize disruption. Building on these insights, we design a reinforcement learning framework to fine-tune a Large Language Model (LLM) for generating disruption-oriented text. Experiments on synthetic and real-world data confirm that tuned LLMs can approach theoretical disruption limits. Our findings raise important considerations for content moderation, adversarial information campaigns, and generative model regulation.
We introduce the Adversarial Confusion Attack, a new class of threats against multimodal large language models (MLLMs). Unlike jailbreaks or targeted misclassification, the goal is to induce systematic disruption that makes the model generate incoherent or confidently incorrect outputs. Practical applications include embedding such adversarial images into websites to prevent MLLM-powered AI Agents from operating reliably. The proposed attack maximizes next-token entropy using a small ensemble of open-source MLLMs. In the white-box setting, we show that a single adversarial image can disrupt all models in the ensemble, both in the full-image and Adversarial CAPTCHA settings. Despite relying on a basic adversarial technique (PGD), the attack generates perturbations that transfer to both unseen open-source (e.g., Qwen3-VL) and proprietary (e.g., GPT-5.1) models.
Galaxies like the Milky Way are surrounded by complex populations of satellites at all stages of tidal disruption. In this paper, we present a dynamical study of the disrupting satellite galaxies in the Auriga simulations that are orbiting 28 distinct Milky Way-mass hosts across three resolutions. We find that the satellite galaxy populations are highly disrupted. The majority of satellites that remain fully intact at present day were accreted recently without experiencing more than one pericentre ($n_{\rm peri} \lesssim 1$) and have large apocentres ($r_{\rm apo} \gtrsim 200$ kpc) and pericentres ($r_{\rm peri} \gtrsim 50$ kpc). The remaining satellites have experienced significant tidal disruption and, given full knowledge of the system, would be classified as stellar streams. We find stellar streams in Auriga across the range of pericentres and apocentres of the known Milky Way dwarf galaxy streams and, interestingly, overlapping significantly with the Milky Way intact satellite population. We find no significant change in satellite orbital distributions across resolution. However, we do see substantial halo-to-halo variance of $(r_\text{peri}, r_\text{apo})$ distributions acros
The fabrication of visual misinformation on the web and social media has increased exponentially with the advent of foundational text-to-image diffusion models. Namely, Stable Diffusion inpainters allow the synthesis of maliciously inpainted images of personal and private figures, and copyrighted contents, also known as deepfakes. To combat such generations, a disruption framework, namely Photoguard, has been proposed, where it adds adversarial noise to the context image to disrupt their inpainting synthesis. While their framework suggested a diffusion-friendly approach, the disruption is not sufficiently strong and it requires a significant amount of GPU and time to immunize the context image. In our work, we re-examine both the minimal and favorable conditions for a successful inpainting disruption, proposing DDD, a "Digression guided Diffusion Disruption" framework. First, we identify the most adversarially vulnerable diffusion timestep range with respect to the hidden space. Within this scope of noised manifold, we pose the problem as a semantic digression optimization. We maximize the distance between the inpainting instance's hidden states and a semantic-aware hidden state ce
Face-swapping DeepFakes have become an escalating societal concern, attracting increasing attention in recent years. To counter this, we investigate a new proactive defense framework to prevent individuals from being victimized in DeepFake videos. The core idea of this framework is to contaminate the inputs of DeepFake models by disrupting face detectors, based on the observation that face detectors are commonly used to automatically extract victim faces in most DeepFake techniques. Once the face detectors malfunction, the faces will not be correctly extracted, thereby impairing the training or synthesis stages of DeepFake models. To achieve this, we describe a strategy named {\em FacePoison}, which fools face detectors by adding dedicated adversarial perturbations to video frames. Building upon this, we introduce {\em VideoFacePoison}, an extended strategy that can efficiently propagate FacePoison across video frames instead of applying it individually to each frame, thus significantly reducing the computational overhead while retaining favorable attack performance. This framework is validated on five face detectors, and extensive experiments against eleven different DeepFake mode
In a hierarchically formed Universe, galaxies accrete smaller systems that tidally disrupt as they evolve in the host's potential. We present a complete catalogue of disrupting galaxies accreted onto Milky Way-mass haloes from the Auriga suite of cosmological magnetohydrodynamic zoom-in simulations. We classify accretion events as intact satellites, stellar streams, or phase-mixed systems based on automated criteria calibrated to a visually classified sample, and match accretions to their counterparts in haloes re-simulated at higher resolution. Most satellites with a bound progenitor at the present day have lost substantial amounts of stellar mass -- 67 per cent have $f_\text{bound} < 0.97$ (our threshold of lost stellar mass to no longer be considered intact), while 53 per cent satisfy a more stringent $f_\text{bound} < 0.8$. Streams typically outnumber intact systems, contribute a smaller fraction of overall accreted stars, and are substantial contributors at intermediate distances from the host centre ($\sim$0.1 to $\sim$0.7$R_\text{200m}$, or $\sim$35 to $\sim$250 kpc for the Milky Way). We also identify accretion events that disrupt to form streams around massive intact
We do a morphological, kinematic and chemical analysis of the disrupting cluster UBC 274 (2.5 Gyr, $d=1778$ pc) to study its global properties. We use HDBSCAN to obtain a new membership list up to 50 pc from its centre and up to magnitude $G=19$ using Gaia EDR3 data. We use high resolution and high signal-to-noise spectra to obtain atmospheric parameters of 6 giants and subgiants, and individual abundances of 18 chemical species. The cluster has a highly eccentric (0.93) component, tilted $\sim$10 deg with respect to the plane of the Galaxy, which is morphologically compatible with the result of a test-particle simulation of a disrupting cluster. Our abundance analysis shows that the cluster has a subsolar metallicity of [Fe/H]$=-0.08\pm0.02$. Its chemical pattern is compatible with that of Ruprecht 147, of similar age but located closer to the Sun, with the remarkable exception of neutron-capture elements, which present an overabundance of $[n\mathrm{/Fe]}\sim0.1$. The cluster's elongated morphology is associated with the internal part of its tidal tail, following the expected dynamical process of disruption. We find a significant sign of mass segregation where the most massive st
Recent years have seen fast development in synthesizing realistic human faces using AI technologies. Such fake faces can be weaponized to cause negative personal and social impact. In this work, we develop technologies to defend individuals from becoming victims of recent AI synthesized fake videos by sabotaging would-be training data. This is achieved by disrupting deep neural network (DNN) based face detection method with specially designed imperceptible adversarial perturbations to reduce the quality of the detected faces. We describe attacking schemes under white-box, gray-box and black-box settings, each with decreasing information about the DNN based face detectors. We empirically show the effectiveness of our methods in disrupting state-of-the-art DNN based face detectors on several datasets.
We consider a new class of multi-period network interdiction problems, where interdiction and restructuring decisions are decided upon before the network is operated and implemented throughout the time horizon. We discuss how we apply this new problem to disrupting domestic sex trafficking networks, and introduce a variant where a second cooperating attacker has the ability to interdict victims and prevent the recruitment of prospective victims. This problem is modeled as a bilevel mixed integer linear program (BMILP), and is solved using column-and-constraint generation with partial information. We also simplify the BMILP when all interdictions are implemented before the network is operated. Modeling-based augmentations are proposed to significantly improve the solution time in a majority of instances tested. We apply our method to synthetic domestic sex trafficking networks, and discuss policy implications from our model. In particular, we show how preventing the recruitment of prospective victims may be as essential to disrupting sex trafficking as interdicting existing participants.
Face modification systems using deep learning have become increasingly powerful and accessible. Given images of a person's face, such systems can generate new images of that same person under different expressions and poses. Some systems can also modify targeted attributes such as hair color or age. This type of manipulated images and video have been coined Deepfakes. In order to prevent a malicious user from generating modified images of a person without their consent we tackle the new problem of generating adversarial attacks against such image translation systems, which disrupt the resulting output image. We call this problem disrupting deepfakes. Most image translation architectures are generative models conditioned on an attribute (e.g. put a smile on this person's face). We are first to propose and successfully apply (1) class transferable adversarial attacks that generalize to different classes, which means that the attacker does not need to have knowledge about the conditioning class, and (2) adversarial training for generative adversarial networks (GANs) as a first step towards robust image translation networks. Finally, in gray-box scenarios, blurring can mount a successf
This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify Smoothing (DR-Smoothing). Specifically, we integrate a two-stage prompt processing scheme-first disrupting the input prompt, then rectifying it-into the conventional smoothing defense framework. This disrupt-and-rectify approach improves upon previous disrupt-only approaches by restoring out-of-distribution disrupted prompts to an in-distribution form, thereby reducing the risk of unpredictable LLM behavior. In addition, this two-stage scheme offers a distinct advantage in striking a balance between harmlessness and helpfulness in jailbreaking defense. Notably, we present a theoretical analysis for generic smoothing framework, offering a tight bound for the defense success probability and the requirements on the disruption strength. Our approach can defend against both token-level and prompt-level jailbreaking attacks, under both established and adaptive attacking scenarios. Extensive experime
Criminal networks, such as the Sicilian Mafia, pose substantial threats to public safety, national security, and economic stability. Outdated disruption methods with a focus on removing influential individuals or key players have proven ineffective due to the covertness of the network. Thus, researchers have been trying to apply Social Network Analysis (SNA) techniques, such as centrality-based measures, to identify key players. However, removing individuals with high centrality often proves to be inefficient, as it does not mimic the real-world scenarios that Law Enforcement Agencies (LEAs) face. For instance, the operational costs limit the LEAs from exploiting the results of the centrality-based methods. This study proposes a multi-objective optimisation framework like the Weighted Sum Genetic Algorithm (WS-GA) and the Non-dominated Sorting Genetic Algorithm II (NSGA-II) to identify disruption strategies that balance two conflicting goals, maximising fragmentation and minimising operational cost which is captured by the spatial distance between nodes and the nearest LEA headquarters. The study utilises the "Montagna Operation" dataset for the experiments. The results demonstrate
Large Language Models (LLM) are disrupting science and research in different subjects and industries. Here we report a minimum-viable-product (MVP) web application called $\textbf{ScienceSage}$. It leverages generative artificial intelligence (GenAI) to help researchers disrupt the speed, magnitude and scope of product innovation. $\textbf{ScienceSage}$ enables researchers to build, store, update and query a knowledge base (KB). A KB codifies user's knowledge/information of a given domain in both vector index and knowledge graph (KG) index for efficient information retrieval and query. The knowledge/information can be extracted from user's textual documents, images, videos, audios and/or the research reports generated based on a research question and the latest relevant information on internet. The same set of KBs interconnect three functions on $\textbf{ScienceSage}$: 'Generate Research Report', 'Chat With Your Documents' and 'Chat With Anything'. We share our learning to encourage discussion and improvement of GenAI's role in scientific research.
Stars that orbit too close to a black hole can be ripped apart by strong tides, producing a type of luminous transient event called a ``tidal disruption event" (TDE). Tidal disruption events of stars by supermassive black holes (SMBHs) provide windows into the nuclei of galaxies at size scales that are difficult to observe directly outside our own galactic neighborhood. They provide a unique opportunity to study these supermassive black holes under feeding conditions that change dramatically over ~week-month timescales, and that regularly reach super-Eddington mass inflow rates. Their light curves are dependent on the properties of the disrupting black hole, and can be used to help constrain the lower mass end of the SMBH mass function -- a region of parameter space that is difficult to access with classic dynamical mass measurements.
Despite extensive research on scientific disruption, two questions remain: why disruption has declined amid growing knowledge, and why disruptive work receives fewer and delayed citations. One way to address these questions is to identify an intrinsic, paper-level property that reliably predicts disruption and explains both patterns. Here, we propose a novel measure, knowledge independence, capturing the extent to which a paper draws on references that do not cite one another. Analyzing 114 million publications, we find that knowledge independence strongly predicts disruption and mediates the disruptive advantage of small, onsite, and fresh teams. Its long-term decline, nonreproducible by null models, provides a mechanistic explanation for the parallel decline in disruption. Causal and simulation evidence further indicates that knowledge independence drives the persistent trade-off between disruption and impact. Taken together, these findings fill a critical gap in understanding scientific innovation, revealing a universal law: Knowledge independence breeds disruption but limits recognition.
Market design research in economics naturally focusses on how to improve market efficiency. Our objective here is exactly the opposite - how to design interventions that make a market less efficient. Our research is inspired by the growth of illicit markets online where reducing their efficiency may reduce societal harm. Using a web-based experiment, we find that a partial disruption to delivery is an effective method to decrease market efficiency. The decrease is borne by sellers who sell fewer goods and have lower earnings. A consequence of a disruption to delivery, however, is an increase in market concentration because it facilitates the emergence of a dominant seller. In contrast, we find that attacks on seller ratings are ineffective at reducing market efficiency. This study paves the way for evidence-based, causally driven investigations to aid policies to disrupt cybercrime and other illicit markets.
In many planning applications, we might be interested in finding plans that minimally modify the initial state to achieve the goals. We refer to this concept as plan disruption. In this paper, we formally introduce it, and define various planning-based compilations that aim to jointly optimize both the sum of action costs and plan disruption. Experimental results in different benchmarks show that the reformulated task can be effectively solved in practice to generate plans that balance both objectives.
Many disruptions are caused by resistive wall tearing modes (RWTM). A database of DIII-D locked mode disruptions provides two main disruption criteria, which are shown to be signatures of RWTMs. The first is that the q = 2 rational surface must be sufficiently close the resistive wall surrounding the plasma to interact with it. If active feedback is used, this implies that RWTMs can be prevented from causing major disruptions. This is demonstrated in simulations. The second criterion is that the current profile is sufficiently peaked. This is caused by edge cooling, such as by impurity radiation and turbulence, which suppress edge current and temperature. This implies the disruptions are not caused by neoclassical tearing modes (NTM), because the bootstrap current is also suppressed. The dependence of the critical internal inductance on elongation is given, which suggests that elongation might be used as an actuator to prevent disruptions. At high $β,$ resistive wall modes (RWM) can be stabilized with feedback. Feedback also stabilizes high $β$ RWTMs, as shown in NSTX data and in simulations. These results suggest that RWTM disruptions in ITER might be prevented using the resonant