OB associations are primordial tracers of star formation and Galactic structure. Originally defined about 80 years ago, their historical membership lists have been superseded thanks to the precise astrometry from ESA's \textit{Gaia}'s satellite. Recent studies have however been mostly focused on individual OB associations or limited by the coverage of spectroscopic surveys. In this paper, we exploit a complete census of $\sim$25,000 O- and B-type stars within 1 kpc of the Sun to produce a highly-reliable catalogue of 56 OB associations using the HDBSCAN clustering algorithm, increasing the number of known OB associations by a factor of two within this volume. We assess the validity of this catalogue by crossmatching our OB association members with other catalogues of OB associations, star clusters and young stellar groups, confirming the high-confidence of our census of OB associations. We characterize these OB associations physically (total initial stellar mass, number of OB stars, ...) and kinematically (velocity dispersion, linear expansion ages, ...). The majority of the OB associations (38 out of 56) exhibit a significant expansion pattern in at least one direction, including
Human word associations are a well-known method of gaining insight into the internal mental lexicon, but the responses spontaneously offered by human participants to word cues are not always predictable as they may be influenced by personal experience, emotions or individual cognitive styles. The ability to form associative links between seemingly unrelated concepts can be the driving mechanisms of creativity. We perform a comparison of the associative behaviour of humans compared to large language models. More specifically, we explore associations to emotionally loaded words and try to determine whether large language models generate associations in a similar way to humans. We find that the overlap between humans and LLMs is moderate, but also that the associations of LLMs tend to amplify the underlying emotional load of the stimulus, and that they tend to be more predictable and less creative than human ones.
Dense retrieval systems rank passages by embedding similarity to a query, but multi-hop questions require passages that are associatively related through shared reasoning chains. We introduce Association-Augmented Retrieval (AAR), a lightweight transductive reranking method that trains a small MLP (4.2M parameters) to learn associative relationships between passages in embedding space using contrastive learning on co-occurrence annotations. At inference time, AAR reranks an initial dense retrieval candidate set using bi-directional association scoring. On HotpotQA, AAR improves passage Recall@5 from 0.831 to 0.916 (+8.6 points) without evaluation-set tuning, with gains concentrated on hard questions where the dense baseline fails (+28.5 points). On MuSiQue, AAR achieves +10.1 points in the transductive setting. An inductive model trained on training-split associations and evaluated on unseen validation associations shows no significant improvement, suggesting that the method captures corpus-specific co-occurrences rather than transferable patterns. Ablation studies support this interpretation: training on semantically similar but non-associated passage pairs degrades retrieval belo
OB associations are important probes of recent star formation and Galactic structure. In this study, we focus on the Auriga constellation, an important region of star formation due to its numerous young stars, star-forming regions and open clusters. We show using \textit{Gaia} data that its two previously documented OB associations, Aur OB1 and OB2, are too extended in proper motion and distance to be genuine associations, encouraging us to revisit the census of OB associations in Auriga with modern techniques. We identify 5617 candidate OB stars across the region using photometry, astrometry and our SED fitting code, grouping these into 5 high-confidence OB associations using HDBSCAN. Three of these are replacements to the historical pair of associations - Aur OB2 is divided between a foreground and a background association - while the other two associations are completely new. We connect these OB associations to the surrounding open clusters and star-forming regions, analyse them physically and kinematically, constraining their ages through a combination of 3D kinematic traceback, the position of their members in the HR diagram and their connection to clusters of known age. Four
Existing works examining Vision-Language Models (VLMs) for social biases predominantly focus on a limited set of documented bias associations, such as gender:profession or race:crime. This narrow scope often overlooks a vast range of unexamined implicit associations, restricting the identification and, hence, mitigation of such biases. We address this gap by probing VLMs to (1) uncover hidden, implicit associations across 9 bias dimensions. We systematically explore diverse input and output modalities and (2) demonstrate how biased associations vary in their negativity, toxicity, and extremity. Our work (3) identifies subtle and extreme biases that are typically not recognized by existing methodologies. We make the Dataset of retrieved associations, (Dora), publicly available here https://github.com/chahatraj/BiasDora.
Despite the remarkable performance of foundation vision-language models, the shared representation space for text and vision can also encode harmful label associations detrimental to fairness. While prior work has uncovered bias in vision-language models' (VLMs) classification performance across geography, work has been limited along the important axis of harmful label associations due to a lack of rich, labeled data. In this work, we investigate harmful label associations in the recently released Casual Conversations datasets containing more than 70,000 videos. We study bias in the frequency of harmful label associations across self-provided labels for age, gender, apparent skin tone, and physical adornments across several leading VLMs. We find that VLMs are $4-7$x more likely to harmfully classify individuals with darker skin tones. We also find scaling transformer encoder model size leads to higher confidence in harmful predictions. Finally, we find improvements on standard vision tasks across VLMs does not address disparities in harmful label associations.
In recent decades, traditional drug research and development have been facing challenges such as high cost, long timelines, and high risks. To address these issues, many computational approaches have been suggested for predicting the relationship between drugs and diseases through drug repositioning, aiming to reduce the cost, development cycle, and risks associated with developing new drugs. Researchers have explored different computational methods to predict drug-disease associations, including drug side effects-disease associations, drug-target associations, and miRNAdisease associations. In this comprehensive review, we focus on recent advances in predicting drug-disease association methods for drug repositioning. We first categorize these methods into several groups, including neural network-based algorithms, matrixbased algorithms, recommendation algorithms, link-based reasoning algorithms, and text mining and semantic reasoning. Then, we compare the prediction performance of existing drug-disease association prediction algorithms. Lastly, we delve into the present challenges and future prospects concerning drug-disease associations.
We study the properties of associations of dwarf galaxies and their dependence on the environment. Associations of dwarf galaxies are extended systems composed exclusively of dwarf galaxies, considering as dwarf galaxies those galaxies less massive than $M_{\star, \rm max} = 10^{9.0}$ ${\rm M}_{\odot}\,h^{-1}$. We identify these particular systems using a semi-analytical model of galaxy formation coupled to a dark matter only simulation in the $Λ$ Cold Dark Matter cosmological model. To classify the environment, we estimate eigenvalues from the tidal field of the dark matter particle distribution of the simulation. We find that the majority, two thirds, of associations are located in filaments ($ \sim 67$ per cent), followed by walls ($ \sim 26 $ per cent), while only a small fraction of them are in knots ($ \sim 6 $ per cent) and voids ($ \sim 1 $ per cent). Associations located in more dense environments present significantly higher velocity dispersion than those located in less dense environments, evidencing that the environment plays a fundamental role in their dynamical properties. However, this connection between velocity dispersion and the environment depends exclusively on
OB associations play an important role in Galactic evolution, though their origins and dynamics remain poorly studied, with only a small number of systems analysed in detail. In this paper we revisit the existence and membership of the Cygnus OB associations. We find that of the historical OB associations only Cyg OB2 and OB3 stand out as real groups. We search for new OB stars using a combination of photometry, astrometry, evolutionary models and an SED fitting process, identifying 4680 probable OB stars with a reliability of $>$90\%. From this sample we search for OB associations using a new and flexible clustering technique, identifying 6 new OB associations. Two of these are similar to the associations Cyg OB2 and OB3, though the others bear no relationship to any existing systems. We characterize the properties of the new associations, including their velocity dispersions and total stellar masses, all of which are consistent with typical values for OB associations. We search for evidence of expansion and find that all are expanding, albeit anistropically, with stronger and more significant expansion in the direction of Galactic longitude. We also identify two large-scale (1
We develop a method to identify and determine the physical properties of stellar associations using Hubble Space Telescope (HST) NUV-U-B-V-I imaging of nearby galaxies from the PHANGS-HST survey. We apply a watershed algorithm to density maps constructed from point source catalogues Gaussian smoothed to multiple physical scales from 8 to 64 pc. We develop our method on two galaxies that span the distance range in the PHANGS-HST sample: NGC 3351 (10 Mpc), NGC 1566 (18 Mpc). We test our algorithm with different parameters such as the choice of detection band for the point source catalogue (NUV or V), source density image filtering methods, and absolute magnitude limits. We characterise the properties of the resulting multi-scale associations, including sizes, number of tracer stars, number of associations, photometry, as well as ages, masses, and reddening from Spectral Energy Distribution fitting. Our method successfully identifies structures that occupy loci in the UBVI colour-colour diagram consistent with previously published catalogues of clusters and associations. The median ages of the associations increases from log(age/yr) = 6.6 to log(age/yr) = 6.9 as the spatial scale incr
Transcriptome-wide association studies (TWAS) are powerful tools for identifying gene-level associations by integrating genome-wide association studies and gene expression data. However, most TWAS methods focus on linear associations between genes and traits, ignoring the complex nonlinear relationships that may be present in biological systems. To address this limitation, we propose a novel framework, QTWAS, which integrates a quantile-based gene expression model into the TWAS model, allowing for the discovery of nonlinear and heterogeneous gene-trait associations. Via comprehensive simulations and applications to both continuous and binary traits, we demonstrate that the proposed model is more powerful than conventional TWAS in identifying gene-trait associations.
Quasar-galaxy associations, if they result from the effect of gravitational lensing by foreground galaxies, depend sensitively on the shape of the quasar number counts. Two kinds of quasar number-magnitude relations are predicted to produce quite different properties in quasar-galaxy associations: the counts of Boyle, Shanks and Peterson (1988; BSP) provide both positive and ``negative" associations between distant quasars and foreground galaxies, relating closely with the knee ($B\approx19.15$) in these counts. However, Hawkins and Véron (1993; HV) quasar data lead to only a positive magnitude-independent quasar-galaxy association. The current observational evidence on quasar-galaxy associations, either positive or null, is shown to be the natural result of gravitational lensing if quasars follow the BSP number-magnitude relation. On the other hand, the HV counts are unable to produce the reported associations by the mechanism of gravitational lensing. It is emphasized that special attention should be paid to the limiting magnitudes in the selected quasar samples when one works on quasar-galaxy associations.
Gene-based testing is a commonly employed strategy in many genetic association studies. Gene-trait associations can be complex due to underlying population heterogeneity, gene-environment interactions, and various other reasons. Existing gene-based tests, such as Burden and Sequence Kernel Association Tests (SKAT), are based on detecting differences in a single summary statistic, such as the mean or the variance, and may miss or underestimate higher-order associations that could be scientifically interesting. In this paper, we propose a new family of gene-level association tests which integrate quantile rank score processes to better accommodate complex associations. The resulting test statistics have multiple advantages: (1) they are almost as efficient as the best existing tests when the associations are homogeneous across quantile levels, and have improved efficiency for complex and heterogeneous associations, (2) they provide useful insights on risk stratification, (3) the test statistics are distribution-free, and could hence accommodate a wide range of underlying distributions, and (4) they are computationally efficient. We established the asymptotic properties of the propose
We investigate the associations between background galaxies and foreground clusters of galaxies due to the effect of gravitational lensing by clusters of galaxies. Similar to the well-known quasar-galaxy ones, these associations depend sensitively on the shape of galaxy number-magnitude or number-flux relation, and both positive and ``negative" associations are found to be possible, depending on the limiting magnitude and/or the flux threshold in the surveys. We calculate the enhancement factors assuming a singular isothermal sphere model for clusters of galaxies and a pointlike model for background sources selected in three different wavelengths, $B$, $K$ and radio. Our results show that $K-$ selected galaxies might constitute the best sample to test the ``negative" associations while it is unlikely that one can actually observe any association for blue galaxies. We also point out that bright radio sources ($S>1$ Jy) can provide strong positive associations, which may have been already detected in 3CR sample.
In the SACY (Search for Associations Containing Young-stars) project we try to identify associations of stars younger than the Local Association among HIPPARCOS and/or TYCHO-2 stars later than G0 which are counterparts of the ROSAT X-ray bright sources. High-resolution spectra for the possible optical counterparts were obtained in order to assess both the youth and the spatial motion of each target. More than 1000 ROSAT sources were observed, covering a large area in the Southern Hemisphere. Associations are characterized mainly by the similarity in UVW velocity space of their proposed member, but other parameters, as evolutionary age, Li abundance and distribution in space must also be taken into account. We proposed a method to identify associations when proper motions and radial velocities are available, but no parallaxes. Using the method we found eleven associations in the SACY data.
Low-mass stars 0.1 ~< M ~< 1 Msun) in OB associations are key to addressing some of the most fundamental problems in star formation. The low-mass stellar populations of OB associations provide a snapshot of the fossil star-formation record of giant molecular cloud complexes. Large scale surveys have identified hundreds of members of nearby OB associations, and revealed that low-mass stars exist wherever high-mass stars have recently formed. The spatial distribution of low-mass members of OB associations demonstrate the existence of significant substructure ("subgroups"). This "discretized" sequence of stellar groups is consistent with an origin in short-lived parent molecular clouds within a Giant Molecular Cloud Complex. The low-mass population in each subgroup within an OB association exhibits little evidence for significant age spreads on time scales of ~10 Myr or greater, in agreement with a scenario of rapid star formation and cloud dissipation. The Initial Mass Function (IMF) of the stellar populations in OB associations in the mass range 0.1 ~< M ~< 1 Msun is largely consistent with the field IMF, and most low-mass pre-main sequence stars in the solar vicinity ar
When humans describe images they tend to use combinations of nouns and adjectives, corresponding to objects and their associated attributes respectively. To generate such a description automatically, one needs to model objects, attributes and their associations. Conventional methods require strong annotation of object and attribute locations, making them less scalable. In this paper, we model object-attribute associations from weakly labelled images, such as those widely available on media sharing sites (e.g. Flickr), where only image-level labels (either object or attributes) are given, without their locations and associations. This is achieved by introducing a novel weakly supervised non-parametric Bayesian model. Once learned, given a new image, our model can describe the image, including objects, attributes and their associations, as well as their locations and segmentation. Extensive experiments on benchmark datasets demonstrate that our weakly supervised model performs at par with strongly supervised models on tasks such as image description and retrieval based on object-attribute associations.
We have carried out a study of the early type stars in nearby OB associations spanning an age range of $\sim$ 3 to 16 Myr, with the aim of determining the fraction of stars which belong to the Herbig Ae/Be class. We studied the B, A, and F stars in the nearby ($\le 500$ pc) OB associations Upper Scorpius, Perseus OB2, Lacerta OB1, and Orion OB1, with membership determined from Hipparcos data. We obtained spectra for 440 Hipparcos stars in these associations, from which we determined accurate spectral types, visual extinctions, effective temperatures, luminosities and masses, using Hipparcos photometry. Using colors corrected for reddening, we find that the Herbig Ae/Be stars and the Classical Be stars (CBe) occupy clearly different regions in the JHK diagram. Thus, we use the location on the JHK diagram, as well as the presence of emission lines and of strong 12 microns flux relative to the visual to identify the Herbig Ae/Be stars in the associations. We find that the Herbig Ae/Be stars constitute a small fraction of the early type stellar population even in the younger associations. Comparing the data from associations with different ages and assuming that the near-infrared exces
The first gravitational wave (GW) - gamma-ray burst (GRB) association, GW170817/GRB 170817A, had an offset in time, with the GRB trigger time delayed by $\sim$1.7 s with respect to the merger time of the GW signal. We generally discuss the astrophysical origin of the delay time, $Δt$, of GW-GRB associations within the context of compact binary coalescence (CBC) -- short GRB (sGRB) associations and GW burst -- long GRB (lGRB) associations. In general, the delay time should include three terms, the time to launch a clean (relativistic) jet, $Δt_{\rm jet}$; the time for the jet to break out from the surrounding medium, $Δt_{\rm bo}$; and the time for the jet to reach the energy dissipation and GRB emission site, $Δt_{\rm GRB}$. For CBC-sGRB associations, $Δt_{\rm jet}$ and $Δt_{\rm bo}$ are correlated, and the final delay can be from 10 ms to a few seconds. For GWB-lGRB associations, $Δt_{\rm jet}$ and $Δt_{\rm bo}$ are independent. The latter is at least $\sim$10 s, so that $Δt$ of these associations is at least this long. For certain jet launching mechanisms of lGRBs, $Δt$ can be minutes or even hours long due to the extended engine waiting time $Δt_{\rm wait}$ to launch a jet. We d
Microbiota contribute to many dimensions of host phenotype, including disease. To link specific microbes to specific phenotypes, microbiome-wide association studies compare microbial abundances between two groups of samples. Abundance differences, however, reflect not only direct associations with the phenotype, but also indirect effects due to microbial interactions. We found that microbial interactions could easily generate a large number of spurious associations that provide no mechanistic insight. Using techniques from statistical physics, we developed a method to remove indirect associations and applied it to the largest dataset on pediatric inflammatory bowel disease. Our method corrected the inflation of p-values in standard association tests and showed that only a small subset of associations is directly linked to the disease. Direct associations had a much higher accuracy in separating cases from controls and pointed to immunomodulation, butyrate production, and the brain-gut axis as important factors in the inflammatory bowel disease.