共找到 20 条结果
Historical Simulation (HS) and its extensions form a popular class of methods for estimating Value-at-Risk for portfolios of financial assets based on historical data. In this note, we seek to unify several ideas and models from throughout the literature into a single modeling framework. By explicitly defining a parametric model form for the asset returns and extracting the realized increments of the driving innovation process from historical data, we are able to reproduce the Historical Simulation, filtered Historical Simulation, and displaced Historical Simulation methods. This shows beyond a doubt that these methods need more underlying assumptions than what is often alluded to.
The correct detection of dense article layout and the recognition of characters in historical newspaper pages remains a challenging requirement for Natural Language Processing (NLP) and machine learning applications on historical newspapers in the field of digital history. Digital newspaper portals for historic Germany typically provide Optical Character Recognition (OCR) text, albeit of varying quality. Unfortunately, layout information is often missing, limiting this rich source's scope. Our dataset is designed to enable the training of layout and OCR modells for historic German-language newspapers. The Chronicling Germany dataset contains 693 annotated historical newspaper pages from the time period between 1852 and 1924. The paper presents a processing pipeline and establishes baseline results on in- and out-of-domain test data using this pipeline. Both our dataset and the corresponding baseline code are freely available online. This work creates a starting point for future research in the field of digital history and historic German language newspaper processing. Furthermore, it provides the opportunity to study a low-resource task in computer vision
We construct a four-dimensional diffeomorphism exhibiting a homoclinic tangency of the largest codimension, which admits a historic wandering domain of positive Lebesgue measure. Every orbit in this wandering domain exhibits historic behavior, in the sense that time averages do not converge. This example shows that homoclinic tangencies of the largest codimension can still give rise to positive Lebesgue measure sets with non-convergent statistical behavior.
Digital transformation in the built environment offers new opportunities to improve building maintenance through data-driven approaches. Smart monitoring, predictive modeling, and artificial intelligence can enhance decision-making and enable proactive strategies. The preservation of historic buildings is an important scenario where preventive maintenance is essential to ensure long-term sustainability while protecting heritage values. This thesis presents a comprehensive solution for data-driven smart maintenance of historic buildings, integrating Internet of Things (IoT), cloud computing, edge computing, ontology-based data modeling, and machine learning to improve indoor climate management, energy efficiency, and conservation practices. This thesis advances data-driven conservation of historic buildings by combining smart monitoring, digital twins, and artificial intelligence. The proposed methods enable preventive maintenance and pave the way for the next generation of heritage conservation strategies.
Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic orthographic normalization of the historical source material is pursued. This report proposes a normalization system for German literary texts from c. 1700-1900, trained on a parallel corpus. The proposed system makes use of a machine learning approach using Transformer language models, combining an encoder-decoder model to normalize individual word types, and a pre-trained causal language model to adjust these normalizations within their context. An extensive evaluation shows that the proposed system provides state-of-the-art accuracy, comparable with a much larger fully end-to-end sentence-based normalization system, fine-tuning a pre-trained Transformer large language model. However, the normalization of historical text remains a challenge due to difficulties for models to generalize, and the lack of extensive high-quality parallel data.
The global ocean meridional overturning circulation (GMOC) is central for ocean transport and climate variations. However, a comprehensive picture of its historical mean state and variability remains vague due to limitations in modelling and observing systems. Incorporating observations into models offers a viable approach to reconstructing climate history, yet achieving coherent estimates of GMOC has proven challenging due to difficulties in harmonizing ocean stratification. Here, we demonstrate that applying multiscale data assimilation scheme that integrates atmospheric and oceanic observations into multiple coupled models in a dynamically consistent way, the global ocean currents and GMOC over the past 80 years are retrieved. While the major historic events are printed in variability of the rebuilt GMOC, the timeseries of multisphere 3-dimensional physical variables representing the realistic historical evolution enable us to advance understanding of mechanisms of climate signal propagation cross spheres and give birth to Artificial Intelligence coupled big models, thus advancing the Earth science.
Arabic handwritten text recognition (HTR) is challenging, especially for historical texts, due to diverse writing styles and the intrinsic features of Arabic script. Additionally, Arabic handwriting datasets are smaller compared to English ones, making it difficult to train generalizable Arabic HTR models. To address these challenges, we propose HATFormer, a transformer-based encoder-decoder architecture that builds on a state-of-the-art English HTR model. By leveraging the transformer's attention mechanism, HATFormer captures spatial contextual information to address the intrinsic challenges of Arabic script through differentiating cursive characters, decomposing visual representations, and identifying diacritics. Our customization to historical handwritten Arabic includes an image processor for effective ViT information preprocessing, a text tokenizer for compact Arabic text representation, and a training pipeline that accounts for a limited amount of historic Arabic handwriting data. HATFormer achieves a character error rate (CER) of 8.6% on the largest public historical handwritten Arabic dataset, with a 51% improvement over the best baseline in the literature. HATFormer also a
Historic structures are important for our society but could be prone to structural deterioration due to long service durations and natural impacts. Monitoring the deterioration of historic structures becomes essential for stakeholders to take appropriate interventions. Existing work in the literature primarily focuses on assessing the structural damage at a given moment instead of evaluating the development of deterioration over time. To address this gap, we proposed a novel five-component digital twin framework to monitor time-varying changes in historic structures. A testbed of a casemate in Fort Soledad on the island of Guam was selected to validate our framework. Using this testbed, key implementation steps in our digital twin framework were performed. The findings from this study confirm that our digital twin framework can effectively monitor deterioration over time, which is an urgent need in the cultural heritage preservation community.
The latest developments in digital have provided large data sets that can increasingly easily be accessed and used. These data sets often contain indirect localisation information, such as historical addresses. Historical geocoding is the process of transforming the indirect localisation information to direct localisation that can be placed on a map, which enables spatial analysis and cross-referencing. Many efficient geocoders exist for current addresses, but they do not deal with the temporal aspect and are based on a strict hierarchy (..., city, street, house number) that is hard or impossible to use with historical data. Indeed historical data are full of uncertainties (temporal aspect, semantic aspect, spatial precision, confidence in historical source, ...) that can not be resolved, as there is no way to go back in time to check. We propose an open source, open data, extensible solution for geocoding that is based on the building of gazetteers composed of geohistorical objects extracted from historical topographical maps. Once the gazetteers are available, geocoding an historical address is a matter of finding the geohistorical object in the gazetteers that is the best match
In this study, we present a generalizable workflow to identify documents in a historic language with a nonstandard language and script combination, Armeno-Turkish. We introduce the task of detecting distinct patterns of multilinguality based on the frequency of structured language alternations within a document.
In this article, we present and discuss a user-study prototype, developed for Bakkehuset historic house museum in Copenhagen. We examine how the prototype - a digital sound installation - can expand visitors' experiences of the house and offer encounters with immaterial cultural heritage. Historic house museums often hold back on utilizing digital communication tools inside the houses, since a central purpose of this type of museum is to preserve an original environment. Digital communication tools however hold great potential for facilitating rich encounters with cultural heritage and in particular with the immaterial aspects of museum collections and their histories. In this article we present our design steps and choices, aiming at subtly and seamlessly adding a digital dimension to a historic house. Based on qualitative interviews, we evaluate how the sound installation at Bakkehuset is sensed, interpreted, and used by visitors as part of their museum experience. In turn, we shed light on the historic house museum as a distinct design context for designing hybrid visitor experiences and point to the potentials of digital communication tools in this context.
Instance segmentation of compound objects in XXL-CT imagery poses a unique challenge in non-destructive testing. This complexity arises from the lack of known reference segmentation labels, limited applicable segmentation tools, as well as partially degraded image quality. To asses recent advancements in the field of machine learning-based image segmentation, the "Instance Segmentation XXL-CT Challenge of a Historic Airplane" was conducted. The challenge aimed to explore automatic or interactive instance segmentation methods for an efficient delineation of the different aircraft components, such as screws, rivets, metal sheets or pressure tubes. We report the organization and outcome of this challenge and describe the capabilities and limitations of the submitted segmentation methods.
Smart maintenance of historic buildings involves integration of digital technologies and data analysis methods to help maintain functionalities of these buildings and preserve their heritage values. However, the maintenance of historic buildings is a long-term process. During the process, the digital transformation requires overcoming various challenges, such as stable and scalable storage and computing resources, a consistent format for organizing and representing building data, and a flexible design to integrate data analytics to deliver applications. This licentiate thesis aims to address these challenges by proposing a digitalization framework that integrates Internet of Things (IoT), cloud computing, ontology, and machine learning. IoT devices enable data collection from historic buildings to reveal their latest status. Using a public cloud platform brings stable and scalable resources for storing data, performing analytics, and deploying applications. Ontologies provide a clear and concise way to organize and represent building data, which makes it easier to understand the relationships between different building components and systems. Combined with IoT devices and ontologie
We consider a parametrised perturbation of a $\mathscr C^r$ diffeomorphism on a closed smooth Riemannian manifold with $r\geq 1$, modeled by nonautonomous dynamical systems. A point without time averages for a (nonautonomous) dynamical system is said to have historic behaviour. It is known that for any $\mathscr C^r$ diffeomorphism, the observability of historic behaviour, in the sense of the existence of a positive Lebesgue measure set consisting of points with historic behaviour, disappears under absolutely continuous, independent and identically distributed (i.i.d.) noise. On contrast, we show that the observability of historic behaviour can appear by a non-i.i.d. noise: we consider a contraction mapping for which the set of points with historic behaviour is of zero Lebesgue measure and provide an absolutely continuous, non-i.i.d. noise under which the set of points with historic behaviour is of positive Lebesgue measure.
Researchers and practitioners often face the issue of having to attribute an IP address to an organization. For current data this is comparably easy, using services like whois or other databases. Similarly, for historic data, several entities like the RIPE NCC provide websites that provide access to historic records. For large-scale network measurement work, though, researchers often have to attribute millions of addresses. For current data, Team Cymru provides a bulk whois service which allows bulk address attribution. However, at the time of writing, there is no service available that allows historic bulk attribution of IP addresses. Hence, in this paper, we introduce and evaluate our 'Back-to-the-Future whois' service, allowing historic bulk attribution of IP addresses on a daily granularity based on CAIDA Routeviews aggregates. We provide this service to the community for free, and also share our implementation so researchers can run instances themselves.
We introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language. This is a fundamentally important routine to historians and digital humanities researchers but has never been automated. We compile a high-quality gold-standard text summarisation dataset, which consists of historical German and Chinese news from hundreds of years ago summarised in modern German or Chinese. Based on cross-lingual transfer learning techniques, we propose a summarisation model that can be trained even with no cross-lingual (historical to modern) parallel data, and further benchmark it against state-of-the-art algorithms. We report automatic and human evaluations that distinguish the historic to modern language summarisation task from standard cross-lingual summarisation (i.e., modern to modern language), highlight the distinctness and value of our dataset, and demonstrate that our transfer learning approach outperforms standard cross-lingual benchmarks on this task.
Historic dress artifacts are a valuable source for human studies. In particular, they can provide important insights into the social aspects of their corresponding era. These insights are commonly drawn from garment pictures as well as the accompanying descriptions and are usually stored in a standardized and controlled vocabulary that accurately describes garments and costume items, called the Costume Core Vocabulary. Building an accurate Costume Core from garment descriptions can be challenging because the historic garment items are often donated, and the accompanying descriptions can be based on untrained individuals and use a language common to the period of the items. In this paper, we present an approach to use Natural Language Processing (NLP) to map the free-form text descriptions of the historic items to that of the controlled vocabulary provided by the Costume Core. Despite the limited dataset, we were able to train an NLP model based on the Universal Sentence Encoder to perform this mapping with more than 90% test accuracy for a subset of the Costume Core vocabulary. We describe our methodology, design choices, and development of our approach, and show the feasibility of
Using Caratheodory measures, we associate to each positive orbit ${\mathcal O}_{f}^{+}(x)$ of a measurable map $f$, a Borel measure $η_{x}$. We show that $η_{x}$ is $f$-invariant whenever $f$ is continuous or $η_{x}$ is a probability. These measures are used to study the \emph{historic} points of the system, that is, \emph{points with no Birkhoff averages}, and we construct topologically generic subset of \emph{wild historic points} for wide classes of dynamical models. We use properties of the measure $η_x$ to deduce some features of the dynamical system involved, like the \emph{existence of heteroclinic connections from the existence of open sets of historic points}.
We describe some bulk statistics of historical initial line outages and the implications for forming contingency lists and understanding which initial outages are likely to lead to further cascading. We use historical outage data to estimate the effect of weather on cascading via cause codes and via NOAA storm data. Bad weather significantly increases outage rates and interacts with cascading effects, and should be accounted for in cascading models and simulations. We suggest how weather effects can be incorporated into the OPA cascading simulation and validated. There are very good prospects for improving data processing and models for the bulk statistics of historical outage data so that cascading can be better understood and quantified.
We report a historic Ks-band light curve spanning over three decades of the FUor PGIR20dci recently discovered by Hillenbrand et al. (2021) . We find some minor variability of the object prior to the FUor outburst, an initial rather slow rise in brightness, followed in 2019 by a much steeper rise to the maximum.