Abstract
Large language models (LLMs) such as GPT-5, Claude, and Grok are increasingly incorporated into engineering education, yet their effectiveness on discipline-specific cognitive tasks remains insufficiently characterized. Engineering learning relies heavily on conceptual understanding, multi-step problem solving, and higher-order reasoning, making it essential to evaluate how LLMs perform across these contexts. This study investigates the performance of three LLMs across four undergraduate engineering courses, Thermodynamics, Computer Programming, Systems and Control, and Circuits, using multiple-choice questions aligned with Bloom’s Taxonomy. We examine differences in accuracy and reasoning quality across cognitive levels, compare zero-shot and context-assisted prompting, and benchmark LLM performance against human learners in Computer Programming. Forty-eight instructor-designed questions (12 per course) spanning all Bloom levels were administered to each model. Responses were evaluated using accuracy, a rubric-based reasoning score, and a Composite Performance Index capturing completeness, efficiency, and quality of support. Human performance data were collected through an IRB-approved assessment with undergraduate students. Results show that Claude demonstrated the strongest overall performance and consistency, GPT-5 excelled in analytical reasoning, and Grok exhibited variable accuracy with limited higher-order reasoning. Providing contextual information improved GPT-5’s performance but had minimal impact on Claude or Grok. In Programming, Claude and GPT-5 outperformed human participants across all Bloom levels. These findings suggest that LLMs can support engineering learning when paired with instructional strategies emphasizing verification and critical evaluation. However, persistent limitations at higher cognitive levels highlight the importance of guided, responsible integration, supported by a practical instructional guide for effective LLM use across engineering tasks.
Author supplied keywords
Cite
CITATION STYLE
Barari, G., Basu, D., Lin, Y., & Ortega-Moody, J. (2026). Evaluating Large Language Models in Engineering Education: A Multi-Course Assessment Across Bloom’s Taxonomy. IEEE Access, 14, 95774–95790. https://doi.org/10.1109/ACCESS.2026.3686339
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.