Abstract
In the vast landscape of CERN’s internal documentation, finding and accessing relevant detailed information remains a complex and time-consuming task. To address this challenge, the AccGPT project (Accelerating Science GPT) aims for the development of an intelligent chatbot leveraging Natural Language Processing (NLP) technologies. We utilize open-source Large Language Models (LLMs) to create a specialized chatbot for CERN internal text-based knowledge retrieval, with the potential future extensions to code assistance and other functionalities. A promising first prototype utilising a Retrieval Augmented Generation (RAG) pipeline has already been developed and deployed. Ongoing improvements focus on enhancing the retrieval accuracy, integrating more powerful and larger LLMs, or fine-tuning with domain-specific data to improve domain accuracy and relevance. The chatbot’s user interface design and overall experience are being iteratively improved, and efforts are underway to prepare AccGPT for community-wide testing. Additionally, automated data scraping and preprocessing pipelines are being implemented to ensure an up-to-date, self-sustaining knowledge base.
Cite
CITATION STYLE
Rehm, F., Guerrieri, G., Guijarro, M., Vallecorsa, S., & Kain, V. (2025). AccGPT: A CERN Knowledge Retrieval Chatbot. In EPJ Web of Conferences (Vol. 337). EDP Sciences. https://doi.org/10.1051/epjconf/202533701279
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.