Abstract
One key challenge for instructors is creating high-quality educational content, such as programming practice questions for introductory programming courses. While Large Language Models (LLMs) show promise for this task, their output quality can be inconsistent, and it is often unclear how to systematically improve their performance. In this experience report, we present the development process for ContentGen, an open-source tool that generates programming questions within the context of data science instructional materials. We describe our process of designing the tool and iteratively improving the tool through prompt engineering. To evaluate our changes, we designed and open-sourced a dataset of 91 test cases based on our course materials and developed three metrics to assess the generated questions: Correctness, Contextual Fit, and Coherence. We compare three prompting strategies and find that providing detailed instructions and an automatically generated summary of recently covered instructional materials to the LLM substantially improves the quality of the generated questions across our metrics. A usability study with six data science instructors further suggests that our final prototype is perceived as usable and effective. Our work contributes a case study of evidence-based prompt engineering for an educational tool and offers a practical approach for instructors and tool designers to evaluate and enhance LLM-based content generation.
Author supplied keywords
Cite
CITATION STYLE
Yu, J., Wu, Y., Cha, G., Shah, A., & Lau, S. (2026). Improving LLM-Generated Educational Content: A Case Study on Prototyping, Prompt Engineering, and Evaluating a Tool for Generating Programming Problems for Data Science. In SIGCSE TS 2026 - Proceedings of the 57th ACM Technical Symposium on Computer Science Education V.1 (pp. 1186–1192). Association for Computing Machinery, Inc. https://doi.org/10.1145/3770762.3772619
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.