Hong Kong' senior geography curriculum has included GIS since the early 2000s. However, GIS in secondary schools does not play a significant role in Hong Kong secondary geography education. Analyzing GIS benefits by literature review, it is believed that GIS should be included in both the senior and junior geography curriculum. Moreover, the literature review indicates that without clear instruction from the Hong Kong Education Bureau (EDB), low preparedness of Hong Kong geography teachers, and unsupportive attitudes from academia and textbook publishers, GIS cannot be implemented in secondary schools of Hong Kong. Therefore, suggestions are made for the EDB, geography teachers, academia and textbook publishers to facilitate GIS involvement in senior and junior geography curriculums. The EDB can develop clear guidelines for teachers, academia and textbook publishers' references, and offer student-centered GIS educational courses for teachers. It is important for teachers to be prepared for advanced GIS technology and to even learn along with students. Academics and textbook publishers can provide free GIS maps targeted at Hong Kong' junior and senior geography curriculums. Although
Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional Chinese script with Cantonese as the spoken form and its cultural context, remain underdeveloped. To address this gap, we introduce HKMMLU, a multi-task language understanding benchmark that evaluates Hong Kong's linguistic competence and socio-cultural knowledge. The HKMMLU includes 26,698 multi-choice questions across 66 subjects, organized into four categories: Science, Technology, Engineering, and Mathematics (STEM), Social Sciences, Humanities, and Other. To evaluate the multilingual understanding ability of LLMs, 90,550 Mandarin-Cantonese translation tasks were additionally included. We conduct comprehensive experiments on GPT-4o, Claude 3.7 Sonnet, and 18 open-source LLMs of varying sizes on HKMMLU. The results show that the best-performing model, DeepSeek-V3, struggles to achieve an accuracy of 75\%, significantly lower than that of MMLU and CMMLU. This performance gap highlights the need to improve LLMs' capabilities in Hong Kong-specific language and knowl
We study how legislation that restricts speech can induce online self-censorship and alter online discourse, using the recent Hong Kong national security law as a case study. We collect a dataset of 7 million historical Tweets from Hong Kong users, supplemented with historical snapshots of Tweet streams collected by other researchers. We compare online activity before and after enactment of the national security law, and we find that Hong Kong users demonstrate two types of self-censorship. First, Hong Kong users are more likely than a control group, sampled randomly from historical snapshots of Tweet streams, to remove past online activity. Specifically, Hong Kong users are over a third more likely than the control group to delete or restrict their account and over twice as likely to delete past posts. Second, we find that Hong Kong users post less often about politically sensitive topics that have been censored on social media in mainland China. This trend continues to increase.
The voting system in the Legislative Council of Hong Kong (Legco) is sometimes unicameral and sometimes bicameral, depending on whether the bill is proposed by the Hong Kong government. Therefore, although without any representative within Legco, the Hong Kong government has certain degree of legislative power --- as if there is a virtual representative of the Hong Kong government within the Legco. By introducing such a virtual representative of the Hong Kong government, we show that Legco is a three-dimensional voting system. We also calculate two power indices of the Hong Kong government through this virtual representative and consider the $C$-dimension and the $W$-dimension of Legco. Finally, some implications of this Legco model to the current constitutional reform in Hong Kong will be given.
We present the first study of the Public Register of Licensed Persons and Registered Institutions maintained by the Hong Kong Securities and Futures Commission (SFC) through the lens of complex network analysis. This dataset, spanning 21 years with daily granularity, provides a unique view of the evolving social network between licensed professionals and their affiliated firms in Hong Kong's financial sector. Leveraging large language models, we classify firms (e.g., asset managers, banks) and infer the likely nationality and gender of employees based on their names. This application enhances the dataset by adding rich demographic and organizational context, enabling more precise network analysis. Our preliminary findings reveal key structural features, offering new insights into the dynamics of Hong Kong's financial landscape. We release the structured dataset to enable further research, establishing a foundation for future studies that may inform recruitment strategies, policy-making, and risk management in the financial industry.
In the Hong Kong Observatory, the Analogue Forecast System (AFS) for precipitation has been providing useful reference in predicting possible daily rainfall scenarios for the next 9 days, by identifying historical cases with similar weather patterns to the latest output from the deterministic model of the European Centre for Medium-Range Weather Forecasts (ECMWF). Recent advances in machine learning allow more sophisticated models to be trained using historical data and the patterns of high-impact weather events to be represented more effectively. As such, an enhanced AFS has been developed using the deep learning technique autoencoder. The datasets of the fifth generation of the ECMWF Reanalysis (ERA5) are utilised where more meteorological elements in higher horizontal, vertical and temporal resolutions are available as compared to the previous ECMWF reanalysis products used in the existing AFS. The enhanced AFS features four major steps in generating the daily rain class forecasts: (1) preprocessing of gridded ERA5 and ECMWF model forecast, (2) feature extraction by the pretrained autoencoder, (3) application of optimised feature weightings based on historical cases, and (4) cal
Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large-scale unified medical datasets and hindering the development of powerful medical foundation models. In this work, we present the largest survey to date of medical image datasets, covering over 1,000 open-access datasets with a systematic catalog of their modalities, tasks, anatomies, annotations, limitations, and potential for integration. Our analysis exposes a landscape that is modest in scale, fragmented across narrowly scoped tasks, and unevenly distributed across organs and modalities, which in turn limits the utility of existing medical image datasets for developing versatile and robust medical foundation models. To turn fragmentation into scale, we propose a metadata-driven fusion paradigm (MDFP) that integrates public datasets with shared modalities or tasks, thereby transfor
This study delves into the media analysis of China's ambitious Belt and Road Initiative (BRI), which, in a polarized world, and furthermore, owing to the very polarizing nature of the initiative itself, has received both strong criticisms and conversely positive coverage in media from across the world. In that context, Hong Kong's dynamic media environment, with a particular focus on its drastically changing press freedom before and after the implementation of the National Security Law is of further interest. Leveraging data science techniques, this study employs Global Database of Events, Language, and Tone (GDELT) to comprehensively collect and analyse (English) news articles on the BRI. Through sentiment analysis, we uncover patterns in media coverage over different periods from several countries across the globe, and delve further to investigate the the media situation in the Hong Kong region. This work thus provides valuable insights into how the Belt and Road Initiative has been portrayed in the media and its evolving reception on the global stage, with a specific emphasis on the unique media landscape of Hong Kong. In an era characterised by increasing globalisation and inte
Results of the first comprehensive light pollution survey in Hong Kong are presented. The night-sky brightness was measured and monitored around the city using a portable light sensing device called the Sky Quality Meter over a 15-month period beginning in March 2008. A total of 1,957 data sets were taken at 199 distinct locations, including urban and rural sites covering all 18 Administrative Districts of Hong Kong. The survey shows that the environmental light pollution problem in Hong Kong is severe - the urban night-skies (sky brightness at 15.0 mag per arcsec square) are on average ~100 times brighter than at the darkest rural sites (20.1 mag per arcsec square), indicating that the high lighting densities in the densely populated residential and commercial areas lead to light pollution. In the worst polluted urban location studied, the night-sky at 13.2 mag per arcsec square can be over 500 times brighter than the darkest sites in Hong Kong. The observed night-sky brightness is found to be affected by human factors such as land utilization and population density of the observation sites, together with meteorological and/or environmental factors. Moreover, earlier night-skies (
The majority of inhabitants in Hong Kong are able to read and write in standard Chinese but use Cantonese as the primary spoken language in daily life. Spoken Cantonese can be transcribed into Chinese characters, which constitute the so-called written Cantonese. Written Cantonese exhibits significant lexical and grammatical differences from standard written Chinese. The rise of written Cantonese is increasingly evident in the cyber world. The growing interaction between Mandarin speakers and Cantonese speakers is leading to a clear demand for automatic translation between Chinese and Cantonese. This paper describes a transformer-based neural machine translation (NMT) system for written-Chinese-to-written-Cantonese translation. Given that parallel text data of Chinese and Cantonese are extremely scarce, a major focus of this study is on the effort of preparing good amount of training data for NMT. In addition to collecting 28K parallel sentences from previous linguistic studies and scattered internet resources, we devise an effective approach to obtaining 72K parallel sentences by automatically extracting pairs of semantically similar sentences from parallel articles on Chinese Wiki
Contrastive Language-Image Pre-training (CLIP) has demonstrated outstanding performance in global image understanding and zero-shot transfer through large-scale text-image alignment. However, the core of medical image analysis often lies in the fine-grained understanding of specific anatomical structures or lesion regions. Therefore, precisely comprehending region-of-interest (RoI) information provided by medical professionals or perception models becomes crucial. To address this need, we propose MedP-CLIP, a region-aware medical vision-language model (VLM). MedP-CLIP innovatively integrates medical prior knowledge and designs a feature-level region prompt integration mechanism, enabling it to flexibly respond to various prompt forms (e.g., points, bounding boxes, masks) while maintaining global contextual awareness when focusing on local regions. We pre-train the model on a meticulously constructed large-scale dataset (containing over 6.4 million medical images and 97.3 million region-level annotations), equipping it with cross-disease and cross-modality fine-grained spatial semantic understanding capabilities. Experiments demonstrate that MedP-CLIP significantly outperforms basel
We utilize a fundamentally different model of trading costs to look at the effect of the opening of the Hong Kong Shanghai Connect that links the stock exchanges in the two cities, arguably the biggest event in international business and finance since Christopher Columbus set sail for India. We design a novel methodology that compensates for the lack of data on trading costs in China. We estimate trading costs across similar positions on the dual listed set of securities in Hong Kong and China, hoping to provide useful pieces of information to help scale 'The Great Wall of Chinese Securities Trading Costs'. We then compare actual and estimated trading costs on a sample of real orders across the Hong Kong securities in the dual listed pair to establish the accuracy of our measurements. The primary question we seek to address is 'Which market would be better to trade to gain exposure to the same (or similar) set of securities or sectors?' We find that trading costs on Shanghai, which might have been lower than Hong Kong, might have become higher leading up to the Connect. What remains to be seen is whether this increase in trading costs is a temporary equilibrium due to the frenzy to
The Segment Anything Model (SAM) has recently gained popularity in the field of image segmentation due to its impressive capabilities in various segmentation tasks and its prompt-based interface. However, recent studies and individual experiments have shown that SAM underperforms in medical image segmentation, since the lack of the medical specific knowledge. This raises the question of how to enhance SAM's segmentation capability for medical images. In this paper, instead of fine-tuning the SAM model, we propose the Medical SAM Adapter (Med-SA), which incorporates domain-specific medical knowledge into the segmentation model using a light yet effective adaptation technique. In Med-SA, we propose Space-Depth Transpose (SD-Trans) to adapt 2D SAM to 3D medical images and Hyper-Prompting Adapter (HyP-Adpt) to achieve prompt-conditioned adaptation. We conduct comprehensive evaluation experiments on 17 medical image segmentation tasks across various image modalities. Med-SA outperforms several state-of-the-art (SOTA) medical image segmentation methods, while updating only 2\% of the parameters. Our code is released at https://github.com/KidsWithTokens/Medical-SAM-Adapter.
People are likely to engage in collective behaviour online during extreme events, such as the COVID-19 crisis, to express their awareness, actions and concerns. Hong Kong has implemented stringent public health and social measures (PHSMs) to curb COVID-19 epidemic waves since the first COVID-19 case was confirmed on 22 January 2020. People are likely to engage in collective behaviour online during extreme events, such as the COVID-19 crisis, to express their awareness, actions and concerns. Here, we offer a framework to evaluate interactions among individuals emotions, perception, and online behaviours in Hong Kong during the first two waves (February to June 2020) and found a strong correlation between online behaviours of Google search and the real-time reproduction numbers. To validate the model output of risk perception, we conducted 10 rounds of cross-sectional telephone surveys from February 1 through June 20 in 2020 to quantify risk perception levels over time. Compared with the survey results, the estimates of the risk perception of individuals using our network-based mechanistic model capture 80% of the trend of people risk perception (individuals who worried about being i
Image registration is essential for medical image applications where alignment of voxels across multiple images is needed for qualitative or quantitative analysis. With recent advancements in deep neural networks and parallel computing, deep learning-based medical image registration methods become competitive with their flexible modelling and fast inference capabilities. However, compared to traditional optimization-based registration methods, the speed advantage may come at the cost of registration performance at inference time. Besides, deep neural networks ideally demand large training datasets while optimization-based methods are training-free. To improve registration accuracy and data efficiency, we propose a novel image registration method, termed Recurrent Inference Image Registration (RIIR) network. RIIR is formulated as a meta-learning solver to the registration problem in an iterative manner. RIIR addresses the accuracy and data efficiency issues, by learning the update rule of optimization, with implicit regularization combined with explicit gradient input. We evaluated RIIR extensively on brain MRI and quantitative cardiac MRI datasets, in terms of both registration acc
Between 2003 and 2015 the prices of apartments in Hong Kong (adjusted for inflation) increased by a factor of 3.8. This is much higher than in the United States prior to the so-called subprime crisis of 2007. The analysis of this speculative episode confirms the mechanism and regularities already highlighted by the present authors in similar episodes in other countries. Based on these regularities, it is possible to predict the price trajectory over the time interval 2016-2025. It suggests that, unless appropriate relief is provided by the mainland, Hong Kong will experience a decade-long slump. Possible implications for its relations with Beijing are discussed at the end of the paper.
Reference class forecasting is a method to remove optimism bias and strategic misrepresentation in infrastructure projects and programmes. In 2012 the Hong Kong government's Development Bureau commissioned a feasibility study on reference class forecasting in Hong Kong - a first for the Asia-Pacific region. This study involved 25 roadwork projects, for which forecast costs and durations were compared with actual outcomes. The analysis established and verified the statistical distribution of the forecast accuracy at various stages of project development, and benchmarked the projects against a sample of 863 similar projects. The study contributed to the understanding of how to improve forecasts by de-biasing early estimates, explicitly considering the risk appetite of decision makers, and safeguarding public funding allocation by balancing exceedance and under-use of project budgets.
Background: Guillain-Barré Syndrome (GBS) is a common type of severe acute paralytic neuropathy and associated with other virus infections such as dengue fever and Zika. This study investigate the relationship between GBS, dengue, local meteorological factors in Hong Kong and global climatic factors from January 2000 to June 2016. Methods: The correlations between GBS, dengue, Multivariate El Nino Southern Oscillation Index (MEI) and local meteorological data were explored by the Spearman Rank correlations and cross-correlations between these time series. Poisson regression models were fitted to identify nonlinear associations between MEI and dengue. Cross wavelet analysis was applied to infer potential non-stationary oscillating associations among MEI, dengue and GBS. Findings : An increasing trend was found for both GBS cases and imported dengue cases in Hong Kong. We found a weak but statistically significant negative correlation between GBS and local meteorological factors. MEI explained over 12\% of dengue's variations from Poisson regression models. Wavelet analyses showed that there is possible non-stationary oscillating association between dengue and GBS from 2005 to 2015 i
Medical image segmentation has been significantly advanced with the rapid development of deep learning (DL) techniques. Existing DL-based segmentation models are typically discriminative; i.e., they aim to learn a mapping from the input image to segmentation masks. However, these discriminative methods neglect the underlying data distribution and intrinsic class characteristics, suffering from unstable feature space. In this work, we propose to complement discriminative segmentation methods with the knowledge of underlying data distribution from generative models. To that end, we propose a novel hybrid diffusion framework for medical image segmentation, termed HiDiff, which can synergize the strengths of existing discriminative segmentation models and new generative diffusion models. HiDiff comprises two key components: discriminative segmentor and diffusion refiner. First, we utilize any conventional trained segmentation models as discriminative segmentor, which can provide a segmentation mask prior for diffusion refiner. Second, we propose a novel binary Bernoulli diffusion model (BBDM) as the diffusion refiner, which can effectively, efficiently, and interactively refine the seg
Consequences from the 2019 anti-extradition protests in Hong Kong have been studied in many facets, but one topic of interest that has not been explored is the impact on the immigration of Bangladeshi immigrants into the city. This paper explores the value add of Bangladeshis to the Hong Kong, how the protests affected their mentality and consequently their immigration, and potentially longer-term detrimental effects on the city.