A Feasibility Study on the Use of Large Language Models in Supporting Adaptive Learning
Home Research Details
Cheng Keat TAN, Qing Hao NG, Seh Yi Joseph TAN, Yin Ni NG

A Feasibility Study on the Use of Large Language Models in Supporting Adaptive Learning

0.0 (0 ratings)

Introduction

A feasibility study on the use of large language models in supporting adaptive learning. Study assesses Large Language Models (LLMs) for adaptive learning using pharmacology MCQs. Examines accuracy, validity, and limitations in higher-order thinking, advising cautious adoption.

0
10 views

Abstract

Large Language Models (LLMs) have gained prominence as adaptive learning tools. Although effective in assessing multiple-choice questions (MCQs), their accuracy and feedback validity remain uncertain. This feasibility study examines the accuracy, validity, and scientific substantiation of LLMs’ responses to expert-generated pharmacology MCQs and assesses the automated generation of MCQs to elucidate their potential roles and limitations in adaptive learning. Fifty MCQs, designed and validated by academic pharmacists, were used to test the accuracy of four LLMs. Each question was classified according to Bloom’s Taxonomy. Direct prompting was applied to generate responses for each question. Responses were analysed for accuracy, validity of rationale, and the existence and relevance of supporting references. Chi-Square test and Fisher-Freeman-Halton Test were used to evaluate quantitative findings. Among the four LLMs, ChatGPT-4o achieved the highest accuracy (84%), followed by Google Gemini 1.5 (Gemini) (80%), Microsoft Copilot (Copilot) (72%), and Claude Sonnet 3.5 (Claude) (68%). An answer-rationale discordance and a decline in performance with an increased cognitive complexity, stratified through Bloom’s Taxonomy, were noted across the LLMs. Artificial hallucinations were observed in the study. These limitations underline the challenges of using LLMs in complex, evidence-driven disciplines like pharmacology. LLMs, as an educational tool, may provide value to adaptive learning. However, limitations in logical reasoning, scientific support, and higher-order thinking highlight the need for cautious adoption. Continuous efforts to validate LLMs with larger, more diverse question banks are also necessary to fully investigate their potential in adaptive learning.


Review

This feasibility study provides a timely and critical examination of Large Language Models (LLMs) in supporting adaptive learning, specifically their application in assessing pharmacology multiple-choice questions (MCQs). Given the growing prominence of LLMs as educational tools, the authors rightly address the uncertainties surrounding their accuracy, feedback validity, and scientific substantiation. The study's objective to assess LLM responses to expert-generated MCQs and their ability to generate MCQs themselves offers a valuable dual perspective on their potential roles and limitations within adaptive learning environments, particularly in complex, evidence-driven disciplines. The methodology employed involved testing four prominent LLMs (ChatGPT-4o, Google Gemini 1.5, Microsoft Copilot, and Claude Sonnet 3.5) against a robust set of 50 expert-designed and validated pharmacology MCQs, categorized by Bloom's Taxonomy. The findings reveal a varying but generally high level of accuracy among the LLMs, with ChatGPT-4o performing best at 84%, followed by Gemini (80%), Copilot (72%), and Claude (68%). Crucially, the study uncovered significant qualitative limitations, including answer-rationale discordance, a noticeable decline in performance as cognitive complexity (stratified by Bloom’s Taxonomy) increased, and the presence of artificial hallucinations. These observations underscore fundamental challenges in LLM reliability beyond mere factual recall. In conclusion, while the study suggests that LLMs may offer value to adaptive learning due to their demonstrable accuracy in certain contexts, it robustly highlights significant limitations that necessitate cautious adoption. The observed deficiencies in logical reasoning, scientific support, and higher-order thinking, particularly in a domain like pharmacology, raise concerns about their suitability for complex educational tasks. The authors' call for continuous validation with larger and more diverse question banks is paramount to thoroughly investigate their potential and mitigate the risks associated with their current shortcomings. This work serves as an important reminder that the integration of LLMs into education must be accompanied by rigorous evaluation and a clear understanding of their inherent boundaries.


Full Text

You need to be logged in to view the full text and Download file of this article - A Feasibility Study on the Use of Large Language Models in Supporting Adaptive Learning from Asian Journal of the Scholarship of Teaching and Learning .

Login to View Full Text And Download

Comments


You need to be logged in to post a comment.