This paper introduces embodied communication, a new wireless communication modality in which information is imprinted onto environmental states and recovered by the receiver through sensing. No dedicated communication transmitter is activated, and no additional communication spectrum is occupied; instead, the sensed environment itself becomes the carrier of information. The key insight is that sensing must be reinterpreted for communication. Rather than asking how accurately an unknown physical state can be estimated, embodied communication asks how reliably two states can be distinguished. We formalize this idea through a multi-snapshot radio frequency (RF) sensing model and derive a sensing-induced reliability field that quantifies the distinguishability between physical states. This field turns embodied symbol design into a geometric packing problem shaped by the sensing resolution of the infrastructure. For this embodied channel, we characterize the finite-snapshot $ε$-capacity through achievable designs and converses. We develop lattice-based codebooks, obtain a closed-form hexagonal design under a main-lobe approximation, and establish information-theoretic and geometric uppe
Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advancements, the remote sensing community has begun to adopt VLMs for remote sensing vision-language tasks, including scene understanding, image captioning, and visual question answering. However, existing remote sensing VLMs typically rely on closed-set scene understanding and focus on generic scene descriptions, yet lack the ability to incorporate external knowledge. This limitation hinders their capacity for semantic reasoning over complex or context-dependent queries that involve domain-specific or world knowledge. To address these challenges, we first introduced a multimodal Remote Sensing World Knowledge (RSWK) dataset, which comprises high-resolution satellite imagery and detailed textual descriptions for 14,141 well-known landmarks from 175 countries, integrating both remote sensing domain knowledge and broader world knowledge. Building upon this dataset, we proposed a novel Remote Sensing Retrieval-Augmented Generation (RS-RAG) framework, which consists of two key components. The Multi-Modal Knowledge Vector Database Construction modul
Rydberg Atomic REceivers (RAREs) have demonstrated remarkable capabilities for radio-frequency signal measurement, enabling advanced quantum wireless sensing. Existing RARE-based sensing systems popularly adopt the heterodyne detection methodology, which requires an additional reference source to serve as an atomic mixer. However, this approach entails a bulky transceiver architecture and is limited in the supportable sensing bandwidth. To address these limitations, we propose a self-heterodyne sensing paradigm where the transmitter's self-interference naturally provides the reference signal. We demonstrate that a self-heterodyne RARE functions as an atomic autocorrelator, eliminating the need for external reference sources while supporting substantially wider bandwidth than conventional heterodyne methods. Next, a two-stage algorithm is devised to perform target ranging in self-heterodyne RARE systems. This algorithm is shown to closely approach the Cramer-Rao lower bound. Furthermore, we introduce the power-trajectory ($P$-trajectory) design for RAREs, which maximizes the sensing sensitivity through time-varying transmission power control. An internal noise (ITN)-limited $P$-traj
Open-vocabulary semantic segmentation (OVSS) in remote sensing images aims to segment categories beyond a fixed label space. Recent SAM 3-based methods provide a promising training-free foundation, yet three key issues remain: (1) a single class-name prompt lacks sufficient semantic coverage for complex remote sensing categories; (2) expanding each category into multiple prompts introduces redundant online text encoding; and (3) directly aggregating multiple prompt responses propagates noisy activations into the final prediction. To address these issues, we propose ProC-SAM3, which calibrates SAM 3's prompt interface for remote sensing OVSS from three complementary aspects. First, we construct an offline prompt pool where a Category Matcher groups MLLM-generated candidates into per-category sets, and Expansion Constraints further refine each set using category-specific prior knowledge. Second, the resulting text embeddings are cached and reused across all test images, eliminating repeated text encoding. Third, we introduce Presence-Guided Residual Fusion to gate unreliable decoder outputs by prompt presence and confidence, followed by peak-preserving class aggregation that retains
Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote sensing domain has made significant progress. The resulting models benefit from the absorption of extensive general knowledge and demonstrate strong performance across a variety of remote sensing data analysis tasks. Moreover, they are capable of interacting with users in a conversational manner. In this paper, we aim to provide the remote sensing community with a timely and comprehensive review of the developments in VLM using the two-stage paradigm. Specifically, we first cover a taxonomy of VLM in remote sensing: contrastive learning, visual instruction tuning, and text-conditioned image generation. For each category, we detail the commonly used network architecture and pre-training objectives. Second, we conduct a thorough review of existing works, examining foundation models and task-specific adaptation methods in contrastive-based VLM, architectural upgrades, training strategies and model capabilities in instruction-based VLM, as well as gene
Recently, the flourishing large language models(LLM), especially ChatGPT, have shown exceptional performance in language understanding, reasoning, and interaction, attracting users and researchers from multiple fields and domains. Although LLMs have shown great capacity to perform human-like task accomplishment in natural language and natural image, their potential in handling remote sensing interpretation tasks has not yet been fully explored. Moreover, the lack of automation in remote sensing task planning hinders the accessibility of remote sensing interpretation techniques, especially to non-remote sensing experts from multiple research fields. To this end, we present Remote Sensing ChatGPT, an LLM-powered agent that utilizes ChatGPT to connect various AI-based remote sensing models to solve complicated interpretation tasks. More specifically, given a user request and a remote sensing image, we utilized ChatGPT to understand user requests, perform task planning according to the tasks' functions, execute each subtask iteratively, and generate the final response according to the output of each subtask. Considering that LLM is trained with natural language and is not capable of di
Integrated sensing and communication (ISAC) has emerged as a key technology for future wireless networks by enabling communication and environmental sensing through a common waveform and hardware platform. Among the candidate waveforms for ISAC, Affine Frequency Division Multiplexing (AFDM) had attracted significant attention due to its robustness in high-mobility environments, but it suffers from a high peak-to-average power ratio (PAPR). In this paper, we propose a sensing-aware chirp-subcarrier reservation (CSR) framework that reduces PAPR while improving ranging performance. The proposed method combines low-complexity gradient-based PAPR minimization with a randomized local search that exploits the phase sensitivity of the AFDM autocorrelation function to suppress delay low-ambiguity-zone (LAZ) sidelobes. Numerical results show that the proposed scheme achieves significant PAPR reduction together with significant sidelobe suppression, resulting in improved weak-target detection performance.
Remote sensing change detection is essential for environmental monitoring, urban planning, and related applications. However, current methods often struggle to capture long-range dependencies while maintaining computational efficiency. Although Transformers can effectively model global context, their quadratic complexity poses scalability challenges, and existing linear attention approaches frequently fail to capture intricate spatiotemporal relationships. Drawing inspiration from the recent success of Titans in language tasks, we present ChangeTitans, the Titans-based framework for remote sensing change detection. Specifically, we propose VTitans, the first Titans-based vision backbone that integrates neural memory with segmented local attention, thereby capturing long-range dependencies while mitigating computational overhead. Next, we present a hierarchical VTitans-Adapter to refine multi-scale features across different network layers. Finally, we introduce TS-CBAM, a two-stream fusion module leveraging cross-temporal attention to suppress pseudo-changes and enhance detection accuracy. Experimental evaluations on four benchmark datasets (LEVIR-CD, WHU-CD, LEVIR-CD+, and SYSU-CD)
Vision-language models (VLMs) have shown significant promise in remote sensing applications, particularly for land-use and land-cover (LULC) mapping via zero-shot classification and retrieval. However, current approaches face several key challenges, such as the dependence on caption-based supervision, which is often not available or very limited in terms of the covered semantics, and the fact of being adapted from generic VLM architectures that are suitable for very high resolution images. Consequently, these models tend to prioritize spatial context over spectral and temporal information, limiting their effectiveness for medium-resolution remote sensing imagery. In this work, we present TimeSenCLIP, a lightweight VLM for remote sensing time series, using a cross-view temporal contrastive framework to align multispectral Sentinel-2 time series with geo-tagged ground-level imagery, without requiring textual annotations. Unlike prior VLMs, TimeSenCLIP emphasizes temporal and spectral signals over spatial context, investigating whether single-pixel time series contain sufficient information for solving a variety of tasks.
Quantum sensor networks promise precision advantages over classical and single-sensor strategies, in particular when the estimator is non-local. We address the problem of finding such estimators through a framework we connote spatial quantum sensing: given an underlying field interrogated by a network of quantum sensors at fixed positions, construct an estimator for a property of the field, for example, distinguishing a source of signal, or evaluating the field or its derivatives at an arbitrary point. We first treat polynomial fields, casting the task as an interpolation problem, and then generalize to fields modeled by analytic functions, which yields general least-squares estimators. A central and largely unaddressed question is under what conditions on sensor placement these estimators are well-defined and error-free. For $m$-dimensional arrays we give explicit constructions and proofs in the interpolation setting using algebraic geometry, and establish necessary and sufficient conditions in the general case. Comparing a non-local entangled protocol with the best local strategy, we show that entanglement yields maximal precision in distributed sensing under global resource cons
Ultra High Resolution (UHR) remote sensing imagery (RSI) (e.g. 100,000 $\times$ 100,000 pixels or more) poses a significant challenge for current Remote Sensing Multimodal Large Language Models (RSMLLMs). If choose to resize the UHR image to standard input image size, the extensive spatial and contextual information that UHR images contain will be neglected. Otherwise, the original size of these images often exceeds the token limits of standard RSMLLMs, making it difficult to process the entire image and capture long-range dependencies to answer the query based on the abundant visual context. In this paper, we introduce ImageRAG for RS, a training-free framework to address the complexities of analyzing UHR remote sensing imagery. By transforming UHR remote sensing image analysis task to image's long context selection task, we design an innovative image contextual retrieval mechanism based on the Retrieval-Augmented Generation (RAG) technique, denoted as ImageRAG. ImageRAG's core innovation lies in its ability to selectively retrieve and focus on the most relevant portions of the UHR image as visual contexts that pertain to a given query. Fast path and slow path are proposed in this
As a common method in the field of computer vision, spatial attention mechanism has been widely used in semantic segmentation of remote sensing images due to its outstanding long-range dependency modeling capability. However, remote sensing images are usually characterized by complex backgrounds and large intra-class variance that would degrade their analysis performance. While vanilla spatial attention mechanisms are based on dense affine operations, they tend to introduce a large amount of background contextual information and lack of consideration for intrinsic spatial correlation. To deal with such limitations, this paper proposes a novel scene-Coupling semantic mask network, which reconstructs the vanilla attention with scene coupling and local global semantic masks strategies. Specifically, scene coupling module decomposes scene information into global representations and object distributions, which are then embedded in the attention affinity processes. This Strategy effectively utilizes the intrinsic spatial correlation between features so that improve the process of attention modeling. Meanwhile, local global semantic masks module indirectly correlate pixels with the global
The dominating waveform in 5G is orthogonal frequency division multiplexing (OFDM). OFDM will remain a promising waveform candidate for joint communication and sensing (JCAS) in 6G since OFDM can provide excellent data transmission capability and accurate sensing information. This paper proposes a novel OFDM-based diagonal waveform structure and corresponding signal processing algorithm. This approach allocates the sensing signals along the diagonal of the time-frequency resource block. Therefore, the sensing signals in a linear structure span both the frequency and time domains. The range and velocity of the object can be estimated simultaneously by applying 1D-discrete Fourier transform (DFT) to the diagonal sensing signals. Compared to the conventional 2D-DFT OFDM radar algorithm, the computational complexity of the proposed algorithm is low. In addition, the sensing overhead can be substantially reduced. The performance of the proposed waveform is evaluated using simulation and analysis of results.
On-device computing, or edge computing, is becoming increasingly important for remote sensing, particularly in applications like deep network-based perception on on-orbit satellites and unmanned aerial vehicles (UAVs). In these scenarios, two brain-like capabilities are crucial for remote sensing models: (1) high energy efficiency, allowing the model to operate on edge devices with limited computing resources, and (2) online adaptation, enabling the model to quickly adapt to environmental variations, weather changes, and sensor drift. This work addresses these needs by proposing an online adaptation framework based on spiking neural networks (SNNs) for remote sensing. Starting with a pretrained SNN model, we design an efficient, unsupervised online adaptation algorithm, which adopts an approximation of the BPTT algorithm and only involves forward-in-time computation that significantly reduces the computational complexity of SNN adaptation learning. Besides, we propose an adaptive activation scaling scheme to boost online SNN adaptation performance, particularly in low time-steps. Furthermore, for the more challenging remote sensing detection task, we propose a confidence-based inst
Recent real-time detection transformers have gained popularity due to their simplicity and efficiency. However, these detectors do not explicitly model object rotation, especially in remote sensing imagery where objects appear at arbitrary angles, leading to challenges in angle representation, matching cost, and training stability. In this paper, we propose a real-time oriented object detection transformer, the first real-time end-to-end oriented object detector to the best of our knowledge, that addresses the above issues. Specifically, angle distribution refinement is proposed to reformulate angle regression as an iterative refinement of probability distributions, thereby capturing the uncertainty of object rotation and providing a more fine-grained angle representation. Then, we incorporate a Chamfer distance cost into bipartite matching, measuring box distance via vertex sets, enabling more accurate geometric alignment and eliminating ambiguous matches. Moreover, we propose oriented contrastive denoising to stabilize training and analyze four noise modes. We observe that a ground truth can be assigned to different index queries across different decoder layers, and analyze this
Stereo matching in remote sensing has recently garnered increased attention, primarily focusing on supervised learning. However, datasets with ground truth generated by expensive airbone Lidar exhibit limited quantity and diversity, constraining the effectiveness of supervised networks. In contrast, unsupervised learning methods can leverage the increasing availability of very-high-resolution (VHR) remote sensing images, offering considerable potential in the realm of stereo matching. Motivated by this intuition, we propose a novel unsupervised stereo matching network for VHR remote sensing images. A light-weight module to bridge confidence with predicted error is introduced to refine the core model. Robust unsupervised losses are formulated to enhance network convergence. The experimental results on US3D and WHU-Stereo datasets demonstrate that the proposed network achieves superior accuracy compared to other unsupervised networks and exhibits better generalization capabilities than supervised models. Our code will be available at https://github.com/Elenairene/CBEM.
We study an auto-calibration problem in which a transform-sparse signal is acquired via compressive sensing by multiple sensors in parallel, but with unknown calibration parameters of the sensors. This inverse problem has an important application in pMRI reconstruction, where the calibration parameters of the receiver coils are often difficult and costly to obtain explicitly, but nonetheless are a fundamental requirement for high-precision reconstructions. Most auto-calibration strategies for this problem involve solving a challenging biconvex optimization problem, which lacks reconstruction guarantees. In this work, we transform the auto-calibrated parallel compressive sensing problem to a convex optimization problem using the idea of `lifting'. By exploiting sparsity structures in the signal and the redundancy introduced by multiple sensors, we solve a mixed-norm minimization problem to recover the underlying signal and the sensing parameters simultaneously. Our method provides robust and stable recovery guarantees that take into account the presence of noise and sparsity deficiencies in the signals. As such, it offers a theoretically guaranteed approach to auto-calibrated parall
Wave front sensing of the surface of equal phase for a propagating electromagnetic wave is a vital technology in fields ranging from real time adaptive optics, to high accuracy metrology, to medical optometry. We have developed a new method of wavefront sensing that makes a direct measurement of the electromagnetic phase distribution, or path-length delay, across an optical wavefront. The method is based on techniques developed in radio astronomical interferometric imaging. The method employs optical interferometry using a 2-D aperture mask, a Fourier transform of the interferogram to derive interferometric visibilities, and self-calibration of the complex visibilities to derive the voltage amplitude and phase gains at each hole in the mask, corresponding to corrections for non-uniform illumination and wavefront distortions across the aperture, respectively. The derived self-calibration gain phases are linearly proportional to the electromagnetic path-length distribution to each hole in the aperture mask, relative to the path-length to the reference hole, and hence represent a wavefront sensor with a precision of a small fraction of a wavelength. The method was tested at $λ=400\,$n
Distributed quantum sensing enables the estimation of multiple parameters encoded in spatially separated probes. While traditional quantum sensing is often focused on estimating a single parameter with maximum precision, distributed quantum sensing seeks to estimate some function of multiple parameters that are only locally accessible for each party involved. In such settings it is natural to not want to give away more information than is necessary. To address this, we use the concept of privacy with respect to a function, ensuring that only information about the target function is available to all the parties, and no other information. We define a measure of privacy (essentially how close we are to this condition being satisfied), and show it satisfies a set of naturally desirable properties of such a measure. Using this privacy measure, we identify and construct entangled resources states that ensure privacy for a given function under different resource distributions and encoding dynamics, characterized by Hamiltonian evolution. For separable and parallel Hamiltonians, we prove that the GHZ state is the only private state for certain linear functions, with the minimum amount of r
Data augmentation has shown significant advancements in computer vision to improve model performance over the years, particularly in scenarios with limited and insufficient data. Currently, most studies focus on adjusting the image or its features to expand the size, quality, and variety of samples during training in various tasks including object detection. However, we argue that it is necessary to investigate bounding box transformations as a data augmentation technique rather than image-level transformations, especially in aerial imagery due to potentially inconsistent bounding box annotations. Hence, this letter presents a thorough investigation of bounding box transformation in terms of scaling, rotation, and translation for remote sensing object detection. We call this augmentation strategy NBBOX (Noise Injection into Bounding Box). We conduct extensive experiments on DOTA and DIOR-R, both well-known datasets that include a variety of rotated generic objects in aerial images. Experimental results show that our approach significantly improves remote sensing object detection without whistles and bells and it is more time-efficient than other state-of-the-art augmentation strate