Abstract:Archives Portal Europe (APE, www.archivesportaleurope.net) is the portal of European archives, an aggregator that connects on a single research point the catalogues and digitised archival material of all archives in and about Europe. It currently hosts material from more than 30 countries, and from a variety of archival institutions (such as State archives, city archives, university and parish archives, private institutions, and more). It is maintained by the Archives Portal Europe Foundation, an international consortium of State archives and other archival institutions that aim to connect the archival material of single institutions into one digital repository, in order to allow universal access to the archival heritage of Europe, promoting new forms of archival research beyond national or local boundaries. One of the research tools made available by Archives Portal Europe is by topics; however, these are currently maintained manually by the archivists, and the vast amount of archival material ingested in the portal makes it impossible to have a comprehensive body of topics that describe the whole of the APE repository. Archives are traditionally not organised by their subject content, but around the entity (person, organization, body) that created and/or collected the documents in the course of their activities. While this is an undisputed pillar of archival management, the availability of online digital repositories for archival research requires new tools for digital archival research, particularly when different archival traditions from different countries and different types of institutions are merged into a unique research portal. Topic detection becomes a fundamental tool to guide archival research and to allow archives to be accessible to potentially world-wide users, in a situation where national and linguistics barriers blur, or are re-defined. This paper presents the preliminary results and plan for future iterations of an AI tool for automated topic detection in a multi- lingual environment, where human-created taxonomies act as bases for the algorithms to aggregate relevant material around a specific topic. The development is based on supervised machine learning, with a combination of human inputs in different languages, and of the usage of Wikipedia pages to model the relevant vocabulary and entities.

Identifying epochs in text archives

Understanding Archives: Towards New Research Interfaces Relying on the Semantic Annotation of Documents

History playground: A tool for discovering temporal trends in massive textual corpora

Identifying Historical Travelogues in Large Text Corpora Using Machine Learning

Extracting Event-Centric Document Collections from Large-Scale Web Archives

A fully data-driven method to identify (correlated) changes in diachronic corpora

Unsilencing Colonial Archives via Automated Entity Recognition

Designing Search Tasks for Archive Search

Survey of Computational Approaches to Lexical Semantic Change

Effects of evolutionary linguistics in text classification

Metaphor, Popular Science, and Semantic Tagging: Distant reading with the Historical Thesaurus of English

Text Line Segmentation of Historical Documents: a Survey

Part-of-Speech Tagging for Historical English

Diachronic Document Dataset for Semantic Layout Analysis

Historical insights at scale: A corpus-wide machine learning analysis of early modern astronomic tables

Terminologies, mod{è}les de donn{é}es arch{é}ologiques et th{é}saurus documentaires

Named Entity Recognition and Classification on Historical Documents: A Survey

Event-based Access to Historical Italian War Memoirs

Text Mining and Data Visualization: Exploring Cultural Formations and Structural Changes in Fifty Years of Eighteenth-Century Poetry Criticism (1967–2018)

What's in a ? Cross-Lingual Topic Detection & Information Retrieval in Archives Portal Europe

Newswire: A Large-Scale Structured Database of a Century of Historical News