New Tools for Old Documents: Deploying Fichero on National Infrastructure

May, 2025

Client

Dr. Daniel Tubb, Associate Professor and Chair, Department of Anthropology, University of New Brunswick

ACENET Research Consultant

Dr. Serguei Vassiliev

Objective

The objective of the project was to create a modular AI workflow that automatically transcribes digitized historical documents and extracts structured information, including summaries, named entities, and library cataloguing metadata. The goal was to improve access to archival collections and enable researchers to analyze them more efficiently.

Challenges

The project’s main challenge was scaling an AI-based document processing pipeline from small test datasets to a large historical archive. Achieving this required optimizing the software to make effective use of high-performance computing resources while maintaining reliable transcription and metadata generation across thousands of scanned documents.

Results

Based on the group’s objectives, Serguei adapted and optimized the document processing workflow to run on the HPC cluster, enabling efficient processing of large collections of scanned archival material.

The workflow uses the open-source AI models Qwen2.5-VL-72B to convert images of historical handwritten documents into text and Llama-3.3-70B to automatically generate summaries, extract key information, and produce structured outputs for cataloguing and further analysis. Serguei also developed prompts to improve the quality, consistency, and reliability of the summaries and structured data generated by the system.

With these improvements, the research team successfully demonstrated automated transcription and cataloguing on a representative set of sample documents. The project also identified challenges related to organizing and navigating the underlying archival data, which the team is continuing to address as work progresses toward processing the full collection.

Future

Future work will focus on scaling the workflow to process the full archival collection, and improving the organization and accessibility of the resulting data. Additional enhancements, such as image pre-processing to improve transcription accuracy, are also being considered. In the longer term, the project aims to provide these tools to the broader Canadian research community to support large-scale analysis and digitization of historical archives.