Domain-Specific architectures with accelerators for machine learning and signal processing require efficient bulk data movement and high-bandwidth access to large datasets. Such capabilities are often absent from minimal open-source microcontrollers (MCUs). We present HyperCroc, an extension to the end-to-end open-source RISC-V Croc system-on-chip (SoC) integrating a silicon-proven HyperBus controller for off-chip DRAM and Flash memory access and a DMA engine, providing a practical MCU-class platform with streamlined plug-in support for domain-specific acceleration. HyperBus offers a low-pin-count PSDRAM interface at up to 400 MB/s, enabling bandwidth-scaled dataset access, while the DMA engine enables autonomous, high-throughput transfers without CPU intervention. HyperCroc preserves Croc's open-source synthesis and physical implementation flow targeting IHP's open 130 nm process design kit (PDK); the full chip can be implemented in under one hour on a consumer-grade workstation. We further report first silicon measurements from MLEM, the first Croc tapeout, confirming that the silicon is fully functional at 72 MHz @ 1.2 V and validating the end-to-end flow.
The demand for domain-specific systems-on-chip (SoCs) in artificial intelligence, robotics, and automotive systems is increasing the need for engineers with hands-on expertise on very-large-scale integration (VLSI) design from architecture specification to fabricated silicon. Yet, most VLSI courses rely on restrictively licensed electronic design automation tools and process design kits (PDKs), as well as closed-source hardware designs. We present an end-to-end open-source domain-specific SoC design and fabrication flow built around Croc, a highly customizable RISC-V platform. Built from open-source SystemVerilog intellectual property blocks and integrated with an end-to-end open-source design flow in a 130nm open PDK, Croc enables tapeout projects supporting multiple domain customization options: instruction-set extensions, accelerator co-processors, and peripherals. In our first open-source course experience using Croc, 65 students completed 33 projects, 30 of which produced manufacturable layouts. 18 designs were selected as tapeout candidates, and five were fabricated. A first baseline chip has already been successfully characterized in silicon, demonstrating microcontroller-cl
Ensuring a continuous and growing influx of skilled chip designers and a smooth path from education to innovation are key goals for several national and international "Chips Acts". Silicon democratization can greatly benefit from end-to-end (from silicon technology to software) free and open-source (OS) platforms. We present Croc, an extensible RISC-V microcontroller platform explicitly targeted at hands-on teaching and innovation. Croc features a streamlined OS synthesis and an end-to-end OS implementation flow, ensuring full, unconstrained access to the design, the design automation tools, and the implementation technology. Croc uses the industry-proven, open-source CVE2 core, implementing the RV32I(EMC) instruction set architecture (ISA), enabling students to define and implement their own ISA extensions. MLEM, a tapeout of Croc in IHP's open 130 nm node completed in eight weeks by a team of just two students, demonstrates the platform's viability for hands-on teaching in schools, universities, or even on a self-education path. In spring 2025, ETH Zurich will utilize Croc for its curricular VLSI class, involving up to 80 students, producing up to 40 OS application-specific integ
The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) generation tasks. Human-based meta-evaluation is costly and time-intensive, and automated alternatives are scarce. We address this gap and propose CROC: a scalable framework for automated Contrastive Robustness Checks that systematically probes and quantifies metric robustness by synthesizing contrastive test cases across a comprehensive taxonomy of image properties. With CROC, we generate a pseudo-labeled dataset (CROC$^{syn}$) of over 1 million contrastive prompt-image pairs to enable a fine-grained comparison of evaluation metrics. We also use this dataset to train CROCScore, a new metric that achieves state-of-the-art performance among open-source methods, demonstrating an additional key application of our framework. To complement this dataset, we introduce a human-supervised benchmark (CROC$^{hum}$) targeting especially challenging categories. Our results highlight robustness issues in existing metrics: for example, many fail on prompts involving negation, and all tested open-source metrics fail on at least 24% of cases involving correct
Graph Neural Networks (GNNs) are widely used as the engine for various graph-related tasks, with their effectiveness in analyzing graph-structured data. However, training robust GNNs often demands abundant labeled data, which is a critical bottleneck in real-world applications. This limitation severely impedes progress in Graph Anomaly Detection (GAD), where anomalies are inherently rare, costly to label, and may actively camouflage their patterns to evade detection. To address these problems, we propose Context Refactoring Contrast (CRoC), a simple yet effective framework that trains GNNs for GAD by jointly leveraging limited labeled and abundant unlabeled data. Different from previous works, CRoC exploits the class imbalance inherent in GAD to refactor the context of each node, which builds augmented graphs by recomposing the attributes of nodes while preserving their interaction patterns. Furthermore, CRoC encodes heterogeneous relations separately and integrates them into the message-passing process, enhancing the model's capacity to capture complex interaction semantics. These operations preserve node semantics while encouraging robustness to adversarial camouflage, enabling G
Recent cosmological surveys and datasets have highlighted a variety of tensions to the concordance model of our universe, $Λ$CDM. Of particular interest is the Hubble tension, the $5.5σ$ discrepancy between measurements of the Hubble constant $H_0$ using high redshift CMB data from Planck ($67.27\pm0.60$km$\text{s}^{-1}\text{Mpc}^{-1}$) and low redshift supernovae from SH0ES ($73.2\pm1.3$km$\text{s}^{-1}\text{Mpc}^{-1}$). To avoid stepping on any toes, we have initiated the CROCS collaboration to resolve this tension, gathering experts from across many fields of cosmology, astrophysics, astronomy, machine learning, data science, philosophy, and astrology. In this paper, we present findings from CROCS Data Release 1, corresponding to the first $\sim3$ days and 27 minutes (rest frame) of observation. We perform a robust statistical analysis, showing that Planck and SH0ES both suffer from imperial biasing systematics (IBS) at $5σ$ significance. Accounting for these errors by converting to metric units reconciles the high and low redshift data, with $H_0 = 69.00\pm0.420$km$\text{s}^{-1}\text{Mpc}^{-1}$. We thus report that our results are sufficient to end the Hubble tension for good.
With grid operators confronting rising uncertainty from renewable integration and a broader push toward electrification, Demand-Side Management (DSM) -- particularly Demand Response (DR) -- has attracted significant attention as a cost-effective mechanism for balancing modern electricity systems. Unprecedented volumes of consumption data from a continuing global deployment of smart meters enable consumer segmentation based on real usage behaviours, promising to inform the design of more effective DSM and DR programs. However, existing clustering-based segmentation methods insufficiently reflect the behavioural diversity of consumers, often relying on rigid temporal alignment, and faltering in the presence of anomalies, missing data, or large-scale deployments. To address these challenges, we propose a novel two-stage clustering framework -- Clustered Representations Optimising Consumer Segmentation (CROCS). In the first stage, each consumer's daily load profiles are clustered independently to form a Representative Load Set (RLS), providing a compact summary of their typical diurnal consumption behaviours. In the second stage, consumers are clustered using the Weighted Sum of Minimu
The low-redshift mass-metallicity relation (MZR) is well studied, but the high-redshift MZR remains difficult to observe. To study the early MZR further, we analyze the Cosmic Reionization on Computers (CROC) simulations with a focus on the MZR from redshifts 5 to 10. We find that, across all redshifts, CROC galaxies exhibit similar stellar-phase and gas-phase MZRs that flatten with higher stellar mass. We attribute this flattening to the inaccurate star formation and feedback modeling in CROC (star formation is overly suppressed in massive CROC galaxies). In addition, we show that the ratio between stellar metallicity and gas metallicity ($Z_*/Z_{gas}$) decreases as stellar age increases, meaning that in CROC galaxies, gas accretion rate is lower than metal production rate. With JWST we will be able to compare our predictions to observations of the Epoch of Reionization and understand better early galaxy formation.
Clock synchronization is a key function in embedded wireless systems and networks. This issue is equally important and more challenging in IoT systems nowadays, which often include heterogeneous wireless devices that follow different wireless standards. Conventional solutions to this problem employ gateway-based indirect synchronization, which suffers low accuracy. This paper for the first time studies the problem of cross-technology clock synchronization. Our proposal called Crocs synchronizes WiFi and ZigBee devices by direct cross-technology communication. Crocs decouples the synchronization signal from the transmission of a timestamp. By incorporating a barker-code based beacon for time alignment and cross-technology transmission of timestamps, Crocs achieves robust and accurate synchronization among WiFi and ZigBee devices, with the synchronization error lower than 1 millisecond. We further make attempts to implement different cross-technology communication methods in Crocs and provide insight findings with regard to the achievable accuracy and expected overhead.
Learning dense visual representations without labels is an arduous task and more so from scene-centric data. We propose to tackle this challenging problem by proposing a Cross-view consistency objective with an Online Clustering mechanism (CrOC) to discover and segment the semantics of the views. In the absence of hand-crafted priors, the resulting method is more generalizable and does not require a cumbersome pre-processing step. More importantly, the clustering algorithm conjointly operates on the features of both views, thereby elegantly bypassing the issue of content not represented in both views and the ambiguous matching of objects from one crop to the other. We demonstrate excellent performance on linear and unsupervised segmentation transfer tasks on various datasets and similarly for video object segmentation. Our code and pre-trained models are publicly available at https://github.com/stegmuel/CrOC.
We introduce a model for the explicit evolution of interstellar dust in a cosmological galaxy formation simulation. We post-process a simulation from the Cosmic Reionization on Computers project (CROC, Gnedin 2014), integrating an ordinary differential equation for the evolution of the dust-to-gas ratio along pathlines in the simulation sampled with a tracer particle technique. This model incorporates the effects of dust grain production in asymptotic giant branch (AGB) star winds and supernovae (SN), grain growth due to the accretion of heavy elements from the gas phase of the interstellar medium (ISM), and grain destruction due to thermal sputtering in the high temperature gas of supernova remnants (SNRs). A main conclusion of our analysis is the importance of a carefully chosen dust destruction model, for which different reasonable parameterizations can predict very different values at the $\sim 100$ pc resolution of the ISM in our simulations. We run this dust model on the single most massive galaxy in a 10$h^{-1}$ co-moving Mpc box, which attains a stellar mass of $\sim 2\times10^9 M_{\odot}$ by $z=5$. We find that the model is capable of reproducing dust masses and dust-sensi
Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure, and multi-agent systems. In safety- and mission-critical deployments, such decisions must be accompanied by statistical reliability guarantees rather than by point estimates alone. Conformal changepoint localization (CONCH) and conformal root cause analysis (CROC) meet this need by returning confidence sets that contain the true changepoint, or the true root-cause stream, with a user-specified probability, without parametric assumptions on the data-generating process. In practice, however, observations are frequently corrupted, e.g., by outliers, sensor faults, or adversarial perturbations. While the finite-sample coverage of these procedures is preserved under contamination, the resulting confidence sets can become uninformatively large. Adopting a Huber-type contamination model, this paper proposes weighted CONCH (W-CONCH) and weighted CROC (W-CROC), which downweight observations that are likely to be corrupted with the goal of reducing confidence set
We study distribution-free root cause analysis in multi-stream data, where an evolving underlying system is observed through multiple data streams that may each undergo distributional changes at unknown timepoints. In such settings, the stream exhibiting the earliest change provides a natural starting point for investigating the underlying cause, which we refer to as the root-cause index. Leveraging conformal $p$-values, we propose a novel framework, Conformal Root Cause Analysis (CROC), which constructs finite-sample valid confidence sets for the root-cause index under minimal assumptions: the data streams are independent, and within each stream the pre- and post-change observations are sampled exchangeably from arbitrary and unknown distributions. We further establish a universality property, showing that any distribution-free method for root cause localization can be represented within the CROC framework. In addition, under mild regularity conditions and principled score design, our method yields asymptotically sharp confidence sets that efficiently isolate the root cause. We further extend CROC to efficiently handle cross-stream dependence when present. Extensive simulations de
Quasar absorption lines provide a unique window to the relationship between galaxies and the intergalactic medium during the Epoch of Reionization. In particular, high redshift quasars enable measurements of the neutral hydrogen content of the universe. However, the limited sample size of observed quasar spectra, particularly at the highest redshifts, hampers our ability to fully characterize the intergalactic medium during this epoch from observations alone. In this work, we characterize the distributions of mean opacities of the intergalactic medium in simulations from the Cosmic Reionization on Computers (CROC) project. We find that the distribution of mean opacities along sightlines follows a non-trivial distribution that cannot be easily approximated by a known distribution. When comparing the cumulative distribution function of mean opacities measurements in subsamples of sample sizes similar to observational measurements from the literature, we find consistency between CROC and observations at redshifts $z\lesssim 5.7$. However, at higher redshifts ($z\gtrsim5.7$), the cumulative distribution function of mean opacities from CROC is notably narrower than those from observed q
Image-to-3D models increasingly rely on hierarchical generation to disentangle geometry and texture. However, the design choices underlying these two-stage models--particularly the optimal choice of intermediate geometric representations--remain largely understudied. To investigate this, we introduce unPIC (undo-a-Picture), a modular framework for empirical analysis of image-to-3D pipelines. By factorizing the generation process into a multiview-geometry prior followed by an appearance decoder, unPIC enables a rigorous comparison of intermediate geometry representations. Through this framework, we identify that a specific representation, Camera-Relative Object Coordinates (CROCS), significantly outperforms alternatives such as depth maps, pretrained visual features, and other pointmap-based representations. We demonstrate that CROCS is not only easier for the first-stage geometry prior to predict, but also serves as an effective conditioning signal for ensuring 360-degree consistency during appearance decoding. Another advantage is that CROCS enables fully feedforward, direct 3D point cloud generation without requiring a separate post-hoc reconstruction step. Our unPIC formulation
Risk prediction that capitalizes on emerging genetic findings holds great promise for improving public health and clinical care. However, recent risk prediction research has shown that predictive tests formed on existing common genetic loci, including those from genome-wide association studies, have lacked sufficient accuracy for clinical use. Because most rare variants on the genome have not yet been studied for their role in risk prediction, future disease prediction discoveries should shift toward a more comprehensive risk prediction strategy that takes into account both common and rare variants. We are proposing a collapsing receiver operating characteristic CROC approach for risk prediction research on both common and rare variants. The new approach is an extension of a previously developed forward ROC FROC approach, with additional procedures for handling rare variants. The approach was evaluated through the use of 533 single-nucleotide polymorphisms SNPs in 37 candidate genes from the Genetic Analysis Workshop 17 mini-exome data set. We found that a prediction model built on all SNPs gained more accuracy AUC = 0.605 than one built on common variants alone AUC = 0.585. We fur
I compare the power spectra of the radiation fields from two recent sets of fully coupled simulations that model cosmic reionization: "Cosmic Reionization On Computers" (CROC) and "Thesan". While both simulations have similar power spectra of the radiation sources, the power spectra of the photoionization rate are significantly different at the same values of cosmic time or the same values of the mean neutral hydrogen fraction. However, the power spectra of the photoionization rate can be matched at large scales for the two simulations when the matching snapshots are allowed to vary independently. I.e., on large scales, the clustering of the radiation field in two simulations evolves similarly, but the exact timing of this evolution is different in different simulations and is not parameterized by an easily interpretable physical quantity like the mean neutral fraction or the mean free path. On small scales, large differences are present and remain partially unexplained. Both CROC and Thesan use the Variable Eddington Tensor approximation for modeling radiative transfer, but adopt different closure relations (optically thin OTVET versus M1). The role of this key difference is teste
We investigate the properties of cosmological ionization fronts during the Epoch of Reionization using the CROC simulations. By analyzing reionization timing maps, we characterize ionization front velocities and curvatures and their dependence on the density structure of the intergalactic medium (IGM). The velocity distribution of ionization fronts in the simulations indicates that while the barrier-crossing analytical model captures the overall shape in high-velocity regions, it fails to reproduce the low-velocity tail, highlighting the non-Gaussian nature of the IGM's density field. Ionization front velocities are inversely correlated with local density, propagating faster in underdense regions and more slowly in overdense environments. Faster ionization fronts also lead to higher post-ionization temperatures, reaching a plateau at $\sim 2 \times 10^4$ K for velocities exceeding 3000 km/s. Examining curvature statistics further establishes a connection between ionization front structure and the normalized density contrast $ν$, with trends in overdense regions aligning well with barrier-crossing model predictions, while deviations appear in underdense environments due to model lim
Recently, several observational detections of damping-wing-like features at the edges of ``dark gaps" in the spectra of distant quasars (the ``Malloy-Lidz effect") have been reported, rendering strong support for the existence of ``neutral islands" in the universe at redshifts as low as $z<5.5$. We apply the procedure from one of these works, Zhu et al (2024), to the outputs of fully coupled cosmological simulations from two recent large projects, ``Cosmic Reionization On Computers" (CROC) and ``Thesan". Synthetic spectra in both simulations have statistics of dark gaps similar to observations, but do not exhibit the damping wing features. Moreover, a toy model with neutral islands added ``by hand" only reproduces the observational results when the fraction of neutral islands among all dark gaps approaches 90%. I.e., simulations and observations appear to produce two distinct ``populations" of dark gaps. In addition, in the simulations, the neutral islands at $z=5.9$ should be short-lived and should not extend to $z<5.5$. A possible explanation for this discrepancy is that both simulations underestimate the fluctuations in the photoionization rate and, hence, miss a populatio
To fully exploit the increased luminosity of the HL-LHC, the CMS Inner Tracker is undergoing a major upgrade to withstand extreme radiation levels and data rates, while improving granularity and reducing material budget. The upgraded modules employ thin planar and 3D silicon pixel sensors, bump bonded to a new radiation-hard readout chip, the CROC, designed in 65 nm CMOS technology and powered via a serial scheme. Ensuring the quality of bump bonding between sensors and readout chips is critical for efficient detector operation. This work presents the qualification procedures developed to identify missing or defective bumps using multiple test methods, including crosstalk analysis, reverse/forward bias testing, and X-ray or beta source imaging. Results from prototype modules are presented, and advantages and limitations of each method are discussed.