A multi-stage agentic AI system for extracting information from large digital archives: case study on the Czechoslovak year 1968 in CIA’s FOIA collection

  • Černý J
  • Avramov K
  • Pendse L
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Purpose - This study aims to design, implement and evaluate a conceptual multi-stage artificial intelligence (AI) system for the systematic analysis of large, unstructured digital archives. Using the 1968 Prague Spring and subsequent Soviet invasion of Czechoslovakia as a case study, the paper demonstrates how such a system can automate the extraction of historical intelligence from declassified documents, creating a time-resolved narrative from non-machine-readable sources. Design/methodology/approach - A multi-stage agentic system comprising eight specialized agents was developed to deconstruct the historical research workflow. The system was applied to the corpus of declassified President's Daily Briefs from 1968 to 1969, sourced from the CIA's FOIA Electronic Reading Room. The methodology integrates optical character recognition (OCR) and expert-guided prompt engineering and introduces a novel evaluation framework to quantitatively and qualitatively compare the performance of four distinct LLMs (GPT-5, Claude Sonnet 4.5, Grok 4 and Magistral Medium) across the 2,122-page corpus. Findings - The system produced three key outputs: a comprehensive monthly summary of intelligence reporting, a structured list of key named entities and a thematic quantification of the content. Critically, the comparative analysis reveals significant trade-offs in performance: GPT-5 achieved the highest output quality (F1 score: 0.731), while Claude Sonnet 4.5 offered superior cost-efficiency and processing speed. Moreover, Claude Sonnet 4.5 and Grok 4 demonstrated flawless operational stability, while Mistral Magistral Medium proved most effective at text reduction. These findings underscore that while AI enhances efficiency, expert human oversight remains essential for ensuring interpretive nuance.Research limitations/implicationsA primary limitation is the data acquisition process, due to the lack of a public API for the canonical data source, which affects long-term reproducibility. Furthermore, the reliance on OCR introduces a layer of potential error into the source text. The study implies that fully automated historical analysis is not yet fully feasible; rather, a human-in-the-loop, collaborative approach is essential for credible results. Practical implications - The proposed framework provides a replicable model for historians, archivists, librarians and intelligence analysts to unlock insights from vast, unstructured document collections. It streamlines labor-intensive tasks (e.g. data discovery, text extraction, summarization), allowing researchers to focus on higher-level analysis and interpretation. Originality/value - This paper's primary novelty lies in presenting one of the first comprehensive frameworks for evaluating and benchmarking competing LLMs on complex historical analysis tasks using declassified intelligence documents. It moves beyond single-LLM case studies by offering a complete, end-to-end workflow that not only processes intelligence documents but also provides a replicable methodology for assessing the practical trade-offs (quality, cost and speed) between different AI models in a digital humanities' context.

Cite

CITATION STYLE

APA

Černý, J., Avramov, K., & Pendse, L. R. (2026). A multi-stage agentic AI system for extracting information from large digital archives: case study on the Czechoslovak year 1968 in CIA’s FOIA collection. The Electronic Library, 1–26. https://doi.org/10.1108/el-06-2025-0272

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free