Evaluating Human vs. LLM Assessments Through Character Assembling
ABSTRACT. The advancement in artificial intelligence (AI) gives rise to large language models (LLMs) capable of producing text. LLM is coherent and appropriate to the context provided by input prompts. These emerging technologies are used in educational settings. However, there remains a lack of research on how human evaluations of storytelling compare with those made by LLMs, particularly character design and visual components. This study examines the effectiveness of LLMs in evaluating children’s storytelling abilities and explores the impact of paper-based character assembly on enhancing their creativity and storytelling skills. Specifically, it assesses the performance of open- and closed-source LLMs in a Traditional Chinese educational context, aiming to understand how these LLMs compare with human evaluations. The study involves 11 primary school students who participate in storytelling sessions before and after engaging in paper-based character assembly activities. We utilize a triangulation methodology that includes evaluations of story performance, time consumption, and analysis of character assembly. Both human raters and LLMs, including the novel Deepseek-r1-distill and other models, such as GPT-3.5-turbo and Llama33-70b, are employed to evaluate the narratives. The results demonstrate that the Deepseek-r1-distill model displays high stability and strong correlation with human ratings, outperforming closed source models, showing lower consistency and weaker correlations. Additionally, engaging students in the assembly of paper-based characters noticeably enhances their creativity and storytelling abilities, as evidenced by the development of more complex and personalized narratives. The findings highlight the potential of integrating LLMs into educational settings to enhance children’s storytelling skills. The study also reveals the beneficial effects of active learning approaches in fostering creativity and engagement in storytelling. These insights provide valuable directions for educators seeking to incorporate innovative assessment and learning strategies.
From Machine Translation Use to AI Literacy: A Mixed-Methods Study of Learner Profiles and Developmental Trajectories in EFL Learning
ABSTRACT. This study examines how Japanese English as a Foreign Language (EFL) learners engage with machine translation (MT) in academic writing through an eight-week instructional intervention incorporating pre-editing and post-editing activities. Using a profile-informed sequential explanatory mixed-methods design, the study identifies learner profiles based on perceptions and usage patterns and explores their developmental trajectories over time. Quantitative analysis revealed five distinct learner profiles, while qualitative findings showed that MT use develops differently depending on initial profiles. Specifically, some learners remain constrained by anxiety and limited strategies, whereas others develop strategic and self-regulated use through active engagement. The findings suggest that effective AI use does not emerge automatically from access to tools but requires instructional support, particularly through scaffolding and collaborative learning. The study proposes a developmental model of MT use and highlights its implications for fostering AI literacy in language learning. In addition, the findings suggest that the skills developed through MT use, such as critical evaluation and revision of generated output, constitute foundational competencies for engaging with generative AI, suggesting the need for instructional design that supports learners’ adaptive use of generative AI tools.
GenAI Literacy and Technology Acceptance Among EFL Learners: A Mixed-Methods Study
ABSTRACT. This study explores the relationship between generative AI (GenAI) literacy and Technology Acceptance Model (TAM) constructs among English as a Foreign Language (EFL) learners using a sequential explanatory mixed-methods design. In the quantitative phase, survey data were collected from 225 university students in Japan and analyzed using partial least squares structural equation modeling (PLS-SEM) and importance–performance map analysis (IPMA). In the qualitative phase, semi-structured interviews were conducted with six participants to provide in-depth insights into the quantitative findings. The results revealed that GenAI literacy had significant indirect effects on attitude, behavioral intention, and actual use, with perceived usefulness serving as the primary mediator. Among the GenAI literacy components, perceived benefits emerged as the most influential predictor of behavioral intention, whereas ethical awareness showed high performance but limited impact. Qualitative findings further demonstrated that although learners recognized the benefits of GenAI, they used it selectively as a supplementary tool while maintaining critical and ethical awareness. In addition, learners’ behavioral intention was found to be multi-layered, involving decisions not only about whether to use GenAI, but also when, where, and how to use it. These findings suggest that learners’ acceptance of GenAI is more complex than assumed in traditional TAM. The study highlights the need to extend TAM by incorporating multi-layered attitudes, behavioral intention, and GenAI literacy to better capture learners’ engagement with GenAI.
Cognitive Engagement in AI-Assisted versus Independent Problem Solving: Evidence from an EEG Study
ABSTRACT. As generative artificial intelligence (AI) becomes increasingly integrated into education, understanding its impact on students’ cognitive engagement during problem solving is critical. While behavioral and self-report measures provide some insight, they cannot capture the real-time neural dynamics of AI-supported learning. This study employed electroencephalography (EEG) to investigate cognitive engagement during problem-solving tasks with and without AI support. Using a within-subjects design, participants completed structured and ill-structured problems under human-only and AI-assisted conditions. EEG power spectral density in the frontal, occipital, and temporal regions was analyzed across preparation and solution phases. Linear mixed-effects models showed that AI-assisted problem solving was associated with reduced alpha and theta power, especially during the preparation phase. These results provide evidence of how AI modulates cognitive engagement and demonstrate the utility of EEG for examining the cognitive impact of generative AI in educational contexts.
Research on Teaching Strategies for Promoting Chinese Language Classroom Interaction in Smart Classroom Environments
ABSTRACT. As a representative form of smart learning environment, the smart classroom provides robust technological support for classroom interaction. However, existing research has predominantly focused on developing generic analytical frameworks while paying insufficient attention to the distinctive pedagogical demands of specific subjects, making it difficult to yield discipline-specific instructional implications. Drawing on social interaction theory, this study proposes four teaching strategies for promoting classroom interaction in primary school Chinese language instruction within smart classroom environments: constructing interactive instructional resources, cultivating a harmonious classroom climate, implementing heterogeneous collaborative learning, and providing timely evaluation and feedback. A quasi-experimental study was conducted with two fifth-grade classes from a primary school. Classroom videos were coded using an interaction coding scheme tailored to the Chinese language discipline and were subsequently analyzed through behavioral frequency analysis and Lag Sequential Analysis (LSA). The results indicate that the quality of classroom interaction in the experimental class was significantly higher than that in the control class, suggesting that the proposed strategies can effectively enhance classroom interaction in primary school Chinese language teaching.
ABSTRACT. Predicting student dropout is a pivotal concern for educational institutions, as timely identification of at-risk students enables them to implement early interventions aimed at reducing attrition. This paper proposes an ensemble deep learning model that combines the application of convolutional (CNN) and long short-term memory (LSTM) networks to analyze both spatial feature patterns and temporal dependencies in student academic data. To address the class imbalance inherent in dropout datasets, we apply the SMOTETomek Link technique, which improves the model's sensitivity to the minority class. The proposed CNN-LSTM ensemble aggregates the outputs of five independently trained models to enhance generalization and reduce overfitting. Our model supports both tabular and time-series input formats by transforming student semester-wise data into sequential representations. Experimental evaluation using a real-world dataset from the Polytechnic Institute of Portalegre demonstrates that the ensemble CNN-LSTM model outperforms independent CNN and LSTM model architectures. The best-performing configuration demonstrated an accuracy of 90.38%, a recall of 91.83%, and an F1-score of 90.31%. These results highlight the efficacy of the proposed approach for robust and early dropout prediction, with strong potential to support data-driven decision-making in educational institutions.
ABSTRACT. While multiple-document integration is vital for knowledge construction, how readers dynamically allocate cognitive resources during digital reading remains under- explored. This study aims to combine eye-tracking technology with machine learning algorithms to identify key reading behaviors that predict integration performance. The primary goal is to provide an empirical foundation for developing real-time feedback and navigation support systems in digital learning environments. The dataset consisted of 122 observations from 62 participants, each of whom completed two reading topics under a consistent task instruction (either summary or argumentation). For each topic, participants read four conflicting texts interconnected with hyperlinks, with each text structured into three functional paragraphs: source information, evidence, and conclusion. To capture the complexity of the reading process, we extracted indicators including first-pass reading time and rereading time for each paragraph, saccade counts between paragraphs, and hyperlink clicks. Data analysis first utilized K-means clustering to categorize readers into high- and low-integration groups based on three writing indicators: concept coverage, information transformation, and the use of connective words. Subsequently, Decision Tree and XGBoost algorithms were applied to predict group membership based on process data and task instruction. The results revealed that evidence rereading time and task instruction were the most critical predictors of high-quality integration. Specifically, while the summary instruction generally increased the probability of high-integration classification, the decision tree revealed that this advantage was conditional: readers under the summary instruction still fell into the low-integration group if they demonstrated insufficient initial processing or low navigation engagement. Conversely, sufficient rereading time on evidence paragraphs emerged as a prerequisite for success across both tasks. These findings suggest that digital reading platforms should leverage real-time behavioral data and task demands to provide personalized scaffolding, thereby optimizing learners' performance in multiple-document integration.
Beyond Positive Attitudes Toward Inquiry: Structural and Psychological Predictors of Inquiry-Based Learning Implementation
ABSTRACT. This study examined the psychological and structural factors associated with the implementation of inquiry-based learning (IBL) among 650 in-service teachers in Japan. An exploratory factor analysis of items related to the school environment identified five dimensions: “School Infrastructure,” “Implementation Barriers,” “Educational Value,” “Support from External Stakeholders,” and “Leadership Support.” Hierarchical regression analysis revealed that school infrastructure, teacher self-efficacy for inquiry, and professional learning orientation were significant predictors of IBL practice, even when psychological and structural variables were simultaneously considered. In contrast, positive attitudes toward IBL were not significant when structural and capability-related variables were controlled. These findings suggest that sustainable IBL implementation depends not only on teachers’ personal beliefs, but also on the availability of adequate organizational infrastructure and opportunities for continuous professional development. This is particularly relevant in the context of technology-enhanced learning environments, where ICT infrastructure forms a central component of school readiness for inquiry-based practice.
Teachers’ Integration of LLMs in Place-Based Education in Aotearoa
ABSTRACT. This study investigated Aotearoa New Zealand (NZ) teachers' perceptions and integration of Large Language Models (LLMs) in teaching, with particular attention to place-based education (PBE) contexts. Data were collected from 182 in-service NZ teachers via a survey combining scales and open-ended questions. Using an embedded mixed methods design, quantitative data were analysed through hierarchical regression and mediation analysis; qualitative data analysis through deductive thematic analysis guided by quantitative findings. Perceived usefulness emerged as the dominant predictor of AI integration, suggesting that attitude and belief function as proximal cognitive mediators through which utility perceptions are translated into adoption decisions. Despite moderately positive attitudes, trust was lower than midpoint, shaped primarily by professional culture, teaching level, and discipline identity. PBE experience did not predict AI integration significantly, suggesting that PBE educators neither simply resist nor embrace AI. Teachers navigated the tension between LLM and PBE through selective adoption. Using AI for generic procedural tasks while maintaining PBE practices for culturally grounded content involving Māori history including local pūrākau (narratives) and oral histories. These findings suggest mainstream LLMs' structural limitations for place-based and Indigenous educational contexts where knowledge is relational, oral, and community-embedded, highlighting the need for culturally responsive AI tools.
The Effects of Supportive and Reflective Scaffolding on Argumentation Skills in Scientific Argumentation
ABSTRACT. Multiple studies have indicated that scaffolded argumentation instruction is crucial for developing students' scientific argumentation skills. However, existing research lacks empirical examination of the differential effects of supportive versus reflective scaffolding in middle school scientific argumentation. This quasi-experimental study investigated how both scaffold types affect students' argumentation skills. Participants included 148 eighth-grade students assigned to a supportive scaffolding group (n = 49), a reflective scaffolding group (n = 49), and a no-scaffolding control group (n = 50). The supportive group received four argumentation sessions with supportive scaffolding, while the reflective group received four sessions with reflective scaffolding. The control group completed argumentation tasks without scaffolds. Pretests confirmed baseline equivalence across groups. After the 4-week intervention, results revealed: (1) For argumentation skills, the reflective scaffolding group produced significantly more advanced elements (reasoning,evidence and rebuttals); (2) The reflective group exhibited a higher proportion of advanced arguers and faster progression rates during later interventions. These findings suggest implementing context-appropriate scaffolding strategies at different argumentation stages. This study provides empirical evidence for middle school science argumentation pedagogy.
Making Verification Visible in AI-Supported Engineering Education: A Mixed-Methods Design Case in a UAV Perception Course
ABSTRACT. Artificial intelligence (AI)-assisted tools can help students enter complex engineering workflows, but fluent outputs may hide wrong commands, fabricated values, and weak engineering assumptions. This paper reports a practice-driven mixed-methods design case from the first offering of a project-based postgraduate course on low-altitude unmanned aerial vehicle (UAV) perception. The course treated AI output as a hypothesis to be checked against documentation, code, calibration files, logs, metrics, visual outputs, and repositories. Students worked with DockerHub demos, Docker and Robot Operating System (ROS), ORB-SLAM3, OpenSplat, U-Net-style segmentation, public UAV datasets, and a group final project. Data included non-matched beginning- and end-of-semester surveys, 31 reflection reports, and GitHub traces from two required individual assignments, with final project reports, slides, optional repositories, and leaderboards analysed as contextual group evidence. The post-course cohort reported higher AI Literacy, Critical Thinking, and Problem Solving scores. Reflections showed that students used AI to enter difficult workflows, met unreliable outputs, and recovered through documentation checks, metric comparison, visual inspection, one-variable tests, basic consistency checks, and manual debugging. The study contributes a verification-oriented accountability cycle and three practice principles for responsible AI-supported engineering education.
Reorganizing Collaborative Problem Solving through Anthropomorphic Affective-Cognitive Support: Evidence from a Virtual Pedagogical Agent in CSCL
ABSTRACT. This study examines how an anthropomorphic virtual pedagogical agent (VPA) reorganizes students’ collaborative problem solving (CPS) processes through affective-cognitive support. Rather than treating pedagogical agents only as task tutors, we conceptualized the VPA as a CSCL support mechanism that couples cognitive regulation with socio-emotional cues. A quasi-experimental study was conducted with 126 undergraduate students assigned to three conditions: an anthropomorphic VPA integrating cognitive and affective regulation, a cognitive-only pedagogical agent (PA), and a no-agent condition (NA). The intervention lasted eight weeks and involved collaborative learning tasks aligned with the PISA 2015 CPS framework. CPS outcomes were measured using a 24-item questionnaire, and group discourse was analyzed through metacognitive coding, CPS coding, epistemic network analysis (ENA), and lag sequential analysis (LSA). Results showed that both agent-supported conditions improved overall CPS, but the VPA produced stronger gains than the PA, especially in representation and monitoring. Process analyses further indicated that the VPA strengthened connections among task understanding, strategy selection, planning, monitoring, and emotion/conflict handling, and supported more integrated cross-dimensional CPS sequences. The findings suggest that anthropomorphic features can function as pedagogically meaningful affective-cognitive regulation mechanisms when tightly coupled with task guidance, role coordination, and process monitoring. This study contributes to AI-supported CSCL by explaining not only whether an agent improves CPS, but how anthropomorphic affective-cognitive support reshapes collaborative regulation.
Connecting Collaborative Argumentation and Individual Writing: A Macro-Scripted, Socratic AI Workflow
ABSTRACT. This study examines the feasibility of a macro-scripted instructional workflow designed to bridge the gap between collaborative argumentation and individual writing in a university English as a Foreign Language (EFL) course. The research addresses a critical internalization challenge in Computer-Supported Collaborative Learning (CSCL): the transition where ideas generated through peer discourse do not automatically manifest as coherent individual reasoning unless learners are supported in reorganizing collective resources into personal writing plans. We developed the Collaborative Argumentation and Writing System (CAWS) to orchestrate this transition through three aligned stages: Knowledge Forum-based collaborative argumentation, non-generative Socratic AI planning, and individual essay drafting. Fifty-eight first-year undergraduates participated in a six-week intervention. Pre- and post-intervention essays were assessed using a multi-dimensional analytic rubric, alongside a questionnaire measuring perceived usefulness, ease of use, trust, and learning gains. Results showed significant gains in overall writing quality, with substantial improvements in reasoning development and counterargument/refutation, whereas evidence use showed limited growth. While students reported positive perceptions, these subjective measures did not significantly correlate with objective writing gains. The findings suggest that the CAWS workflow is a feasible, authorship-preserving model for linking social knowledge building to individual composition, highlighting scaffolded internalization as a key target for future longitudinal research in CSCL
Knowledge-State-Based Peer Selection for Slight Upward Social Comparison in K-12 Dashboards
ABSTRACT. This study examines how slight upward social comparison can be designed for K-12 student-facing dashboards without relying on grades or final outcomes. Although upward comparison has been used in learning analytics dashboards to promote engagement, score-based comparison may intensify competition and obscure students’ actual learning needs. To address this issue, the study focuses on knowledge states estimated through an Open Knowledge and Learner Model (OKLM), which links learning records and test results to knowledge components in a textbook-based knowledge map. Data were collected from 120 first-year junior high school students learning mathematics. Using concept-level mastery values in OKLM, time-series clustering was conducted to identify patterns of change in students’ knowledge states. Three transition patterns were identified: students struggling with applications, students with consistently high understanding, and students struggling with basics. A one-way ANOVA using final test scores showed significant differences among the clusters, but post-hoc comparisons indicated no significant difference between the two struggling groups. These results suggest that students with similar final scores may have different underlying knowledge-state trajectories. Based on these findings, the study proposes knowledge-state-based peer selection as a way to support attainable upward comparison, enabling learners to refer to peers who share similar difficulties but show slightly higher mastery.
Developing a Diagnostic Framework for Teacher–AI Collaborative Decision-Making Processes
ABSTRACT. As generative artificial intelligence (GenAI) becomes increasingly embedded in classroom practice, understanding how teachers and AI jointly construct teaching decisions has emerged as a critical yet underexplored challenge. Existing research largely focuses on outcomes or tool effectiveness, offering limited insight into the underlying decision-making processes. To address this gap, this study developed a process-sensitive multimodal diagnostic framework for teacher–AI collaborative decision-making. Using a mixed inductive–consensus design, Phase A applied procedural grounded theory to 30 authentic classroom interaction recordings (1,282 coded instances) to generate candidate indicators. Phase B refined the framework through a two-round Delphi study with 18 experts, and Phase C employed the Analytic Hierarchy Process to derive normalized indicator weights. The results yielded a three-stage decision model—Information Sharing, Strategy Internalization and Generation, and Collaborative Construction—comprising three primary and nine secondary indicators. Weighting results highlighted Collaborative Construction as the most influential stage, with Feedback Regulation identified as critical. The framework provides a replicable foundation for diagnosing, comparing, and optimizing teacher–AI decision processes in instructional contexts.
When Can Simulated Learners Be Trusted? Decision-Centered Validity for Learner Simulators
ABSTRACT. Educational technology platforms that personalize instruction routinely test proposed changes to their decision rules by running those rules against simulated student data before applying any change to real students. As generative and agentic AI enter learning ecologies, large-language-model-based learner simulators are increasingly proposed for this offline testing. A learner simulator takes a student's recorded interaction history and generates a synthetic sequence of future actions that stands in for the student's actual future behavior. Current evaluations focus on whether simulated sequences look realistic in timing patterns, action types, and response rates. This focus leaves a consequential gap: a simulator can produce realistic-looking sequences while still misidentifying which students need support, which follow-up best serves a student who answered incorrectly, or which students are ready to advance. This paper introduces decision-centered validity, a reusable deployment-readiness framework that tests whether replacing actual student data with simulated data changes the decisions a platform would make about individual students. The audit evaluates four decisions platforms commonly make: selecting which students receive limited instructional support, assigning follow-up activities after incorrect responses, identifying students who have gone inactive, and judging which students are ready to advance. Results are organized through explicit pass-or-fail tests and three readiness tiers indicating whether a simulator is appropriate for a given decision type. We apply the audit to three publicly available datasets: the Open University Learning Analytics Dataset (OULAD), EdNet KT3, and ASSISTments. Together they cover 35,873 students and 13.57 million recorded interactions across nine benchmarked simulator families. No simulator passes all applicable decision tests, and every family fails the support-allocation ranking test; this verdict holds for frontier and open-weight generative large language models and under an independent multi-model auditor panel. Because support allocation governs which active learners receive scarce resources, a simulator that cannot preserve this ranking raises an equity-of-access concern: a platform may deploy rules that fail to reach the students who need help most.
AI Agents in Immersive Learning Environments: Types, Techniques, and Implementation Challenges
ABSTRACT. AI agents are increasingly deployed as interactive partners in immersive learning environments (ILEs), yet existing reviews provide limited synthesis of the educational roles these agents perform, the techniques that enable them, and the challenges reported in empirical implementation. This paper reports selected findings from a systematic review of 22 studies, guided by two research questions on agent types/techniques and implementation challenges. Across the reviewed studies, AI agents serve four educational roles: scaffolding and instructional mediation, orchestration and process adaptation, assessment and diagnosis, and companion and persona. Their technical implementations have shifted toward multimodal hybrid architectures integrated with ILE platforms, combining generative language and conversational pipelines, learner modelling and adaptive control, structured knowledge and rule-based reasoning, and embodied multimodal integration. Reported challenges are clustered into pedagogical issues, including support versus learner agency, personalisation versus appropriateness of support and presence versus instructional coherence, as well as infrastructural constraints, including interaction latency, physical and attentional load, and complex real-time system integration. Overall, AI agents in ILEs are best understood as contributors and collaborators in a learning process rather than as fixed platform features. The review also suggests aligning agent roles with task phases and learner proficiency, reporting evidence of interaction processes alongside learning outcomes, and exploring emerging capabilities such as agent memory, skills, and autonomy.
How Strict Should Mastery Be? Architecture-Aware Threshold Calibration Across Knowledge Tracing Models
ABSTRACT. When a student uses an adaptive tutoring system, the system must decide whether the student has learned a skill well enough to advance or should keep practicing. This decision is controlled by a mastery threshold, a numerical cutoff applied to the student model's estimate of mastery: when the estimate reaches the cutoff, the system advances the student; otherwise practice continues. As intelligent learning environments increasingly swap classical student models for neural ones, many systems reuse the same threshold, assuming that the same cutoff has the same practical meaning across model types. This paper tests that assumption by evaluating six knowledge tracing (KT) models, computational methods that estimate student mastery from practice records, across four public educational datasets. Using a shared, reproducible evaluation framework, we compare post-mastery hit rate, the fraction of correct responses in the first five practice items after advancement; learner coverage, the fraction of students who ever reach the threshold; practice burden, the number of attempts required before advancement; and a transparent instructional-utility score across 14 thresholds from 0.50 to 0.99. Bayesian Knowledge Tracing is relatively insensitive to threshold increases, while neural models show sharp trade-offs between reported post-mastery performance and learner coverage. Higher thresholds are positively associated with reported post-mastery performance, negligibly for BKT and moderately for neural models, but for neural models this association is about half explained by which students are selected to advance rather than by improved readiness for all who practiced. Reusing a Bayesian Knowledge Tracing-calibrated threshold with a neural model produces the largest instructional utility loss among all transfer directions tested. Stricter thresholds also do not close performance gaps between students with stronger and weaker prior records and can widen disparities in which students reach advancement. The paper provides an architecture-aware calibration framework and practical, equity-conscious guidance for selecting or recalibrating mastery thresholds when tutoring systems change student models or instructional goals.
AI-Supported Formative Feedback for Multimodal Student Reports: Evidence from a Real Classroom Deployment
ABSTRACT. This study investigates the impact of AI-supported formative feedback on student report revision in higher education. Although prior research has primarily focused on text-based assignments, reports used in authentic classroom settings often include multimodal elements such as slides, figures, and visual explanations. The effectiveness of AI-supported feedback for such multimodal reports remains underexplored. To address this gap, we developed an LLM-based system that analyzes multimodal reports through a unified prompt design and generates formative feedback to support students’ self-reflection before submission. The system integrates structure detection, content summarization, and suggestion generation within a single multimodal large language model. We conducted a comparative study in a graduate-level course with more than 100 students, comparing two academic years: one without AI support (2024) and one with the AI system (2025). Multiple data sources were analyzed, including peer evaluations, instructor evaluations, report artifacts, system usage logs, and student surveys. The results show that report quality was significantly higher after the introduction of AI-supported feedback (p < .001), with particularly strong gains in higher-order components such as discussion and conclusion. Students also increased the amount of content in their reports, suggesting more constructive revision behavior. Survey results further indicate generally positive perceptions of the system. These findings provide empirical evidence that LLM-based formative feedback can support the revision of multimodal reports in authentic classroom settings and highlight its potential as a scalable support tool in higher education.
A Preliminary Analysis of the Relationship between Novice Programmers’ Feedback Use and Performance within the RunCode Learning Platform
ABSTRACT. We investigate the impact of LLM-generated feedback on first-year students’
programming activities by analyzing interaction logs from the Run-Code platform. We
collected log data from 37 students across 22 programming exercises. Data were
collected asynchronously over five weeks as part of an introductory programming
course on Java in a university in the Philippines. In this preliminary analysis of the data, we attempted to answer two research questions: What was the relationship between
students’ feedback usage and the student performance on the next problem? What
was the relationship between students’ use of feedback and future programming
performance? Our results show that feedback usage was not a statistically significant
predictor of next-task performance, next-task error count, or future mean task scores,
but it was a statistically significant positive predictor of future mean error count. These findings suggest that feedback access alone does not translate into measurable
improvements in programming performance, and the positive association with future
error counts likely reflects underlying learner difficulty rather than a detrimental effect of feedback.
PEARL: A Fail-Closed Audit Benchmark for Differentially Private Synthetic Educational Data
ABSTRACT. Universities and online learning platforms generate student records used for research on dropout prediction, knowledge tracing, and adaptive tutoring. Direct sharing of these records is restricted by privacy regulations, including FERPA (Family Educational Rights and Privacy Act) and GDPR (General Data Protection Regulation). Differentially private (DP) synthetic data, which replaces real records with generated substitutes that bound each individual's influence on the release, offers one path to compliant sharing. However, existing evaluations report privacy risk and utility as separate scores, missing the practical release question: a dataset may pass a privacy check while losing the information needed for the intended task. We introduce PEARL (Privacy-Equivalence Audit and Release Ledger), a fail-closed audit benchmark for DP synthetic educational data. PEARL rejects a candidate release unless it passes sequential checks for validity, privacy, utility, and task-specific educational use, and records the reason each rejected release fails. Applied to 96 configurations across tabular student records, MOOC clickstreams, and knowledge-tracing datasets, PEARL certifies 12 releases. These include nine tabular releases from MST or DP-CTGAN and three ACT-MOOC releases from DP-Markov (membership inference attack AUC = 0.491, retention = 0.922 at ε = 1.0). Many privacy-passing releases fail due to missing target classes, low utility, label artifacts, or loss of ordered skill-transition structure. A subgroup fairness gate additionally rejects tabular releases that distort at-risk recall across protected groups, and stronger shadow-model and likelihood-ratio attacks leave every certification unchanged. Deep Knowledge Tracing (DKT) and Self-Attentive Knowledge Tracing (SAKT) achieve exactly 0.000 F1 on synthetic data across all tested settings (five seeds, every privacy budget, every synthesizer). This is a reproducible sign that sequence structure is destroyed rather than an implementation artifact. PEARL shows that synthetic educational data should be judged by whether each release still supports its intended educational use, such as knowledge tracing, dropout prediction, and adaptive tutoring, not by privacy alone. Audited this way, responsible data sharing and learner data rights advance, rather than undercut, research in intelligent learning environments.
When LLMs Disagree with Human Annotations: A Manual-Grounded Analysis of TalkMoves Coding
ABSTRACT. Automated coding of classroom discourse typically treats human-annotated labels as a reliable gold standard. This paper questions that assumption. We apply a manual-grounded prompt---the official TalkMoves coding manual, verbatim and unmodified---to two LLMs (DeepSeek V4 Pro and Claude Opus 4.7) on 400 teacher utterances. Both models exceed the prior best reported kappa of 0.58, confirming the strategy's effectiveness. We then treat the 104 AI-human disagreements not as model errors, but as cases for audit. Manual review finds that 42% reflect annotation inconsistencies in the gold standard, not AI mistakes. This finding is supported by cross-model evidence: cases where both models agree against the gold standard carry a 54% inconsistency rate versus 13% for single-model disagreements, with 89% of cross-model cases involving both models predicting the exact same alternative label. An independent second rater further agrees with the AI in 56.7% of disputed cases versus 30.0% for the original labels. These results suggest that manual-grounded LLM coding can serve as a practical annotation quality audit tool for educational discourse datasets.
Information Retrieval in Learning Materials: Which Language Models Work Best and How to Improve Them
ABSTRACT. Accurate information retrieval from learning materials is a prerequisite for trustworthiness in many Artificial Intelligence in Education (AIED) applications, including Retrieval-Augmented Generation (RAG) and personalized learning frameworks. However, the effectiveness of general retrieval models in capturing complex pedagogical relationships remains largely unexplored, hindering the selection of effective models for real-world learning environments. This paper presents a technical contribution to this educational and domain-specific retrieval by introducing a benchmark of 1,470 annotated slide pairs designed to evaluate the retrieval of pedagogically meaningful relationships. In addition, we propose two methods to adapt language models to this educational setting. Our findings provide practical guidance for model selection, showing that OpenAI’s embeddings perform strongly among general models and demonstrate that our adaptation methodologies lead to an improvement of retrieval by up to 43% for RoBERTa and 6% for OpenAI embeddings (PR-AUC). Finally, our approach enhances interpretability by providing clearer retrieval rationales, supporting transparency and trust in AI-driven educational systems.
Teacher-Orchestrated MyGPT for EFL Speaking Rehearsal and Writing Revision: A Classroom-Based Acceptance Study
ABSTRACT. Generative artificial intelligence (AI) and large language models offer new opportunities for technology-enhanced language learning (TELL), but classroom use remains challenging when AI tools are treated as generic chatbots without pedagogical structure. This exploratory classroom-based study examined a teacher-orchestrated MyGPT designed to support productive English as a Foreign Language (EFL) practice. The instructor configured two task-specific modules: Voice UP for short speaking rehearsal and Reading Summary Checker for paragraph-level academic summary revision. Fifty-three first-year university EFL students used the MyGPT during a three-week classroom implementation consisting of three one-hour sessions and completed a 15-item post-use questionnaire measuring perceived usefulness, perceived ease of use, behavioral intention, and satisfaction. Results showed consistently positive acceptance across all four constructs, with no significant differences by gender, self-rated English proficiency, or prior AI tool experience. Students’ general attitude toward AI in education was positively associated with all acceptance outcomes and explained additional variance after background characteristics were considered. The findings suggest that teacher-orchestrated MyGPT designs can serve as configurable pedagogical infrastructures for AI-supported speaking and writing practice, offering a manageable entry point for integrating generative AI into EFL classrooms.
From Comments to Actions: Multi-task Analysis of Vietnamese Student Feedback for Educational Quality Assurance
ABSTRACT. Open-ended student feedback is an important source of data for improving teaching quality, curriculum design, and learning services. However, these comments are usually written as free-form text, making them difficult to analyze manually at scale. Existing studies on Vietnamese student feedback mainly focus on sentiment analysis or topic classification, while less attention has been paid to identifying issues that can be turned into concrete improvement actions. This paper proposes an action-oriented approach to student feedback analysis with four tasks: sentiment classification, topic classification, actionability detection, and issue type classification. Based on the existing sentiment and topic labels, we construct two additional label groups, namely Actionability and Issue Type, using a rule-based labeling strategy followed by manual review. Using a dataset of more than 16,000 Vietnamese student comments, we compare traditional machine learning models, single-task Transformer models, and multi-task Transformer models. The experimental results show that XLM-R-MTL achieves F1-scores of 0.9413 for sentiment classification, 0.8927 for topic classification, 0.9639 for actionability detection, and 0.9045 for issue type classification. These results suggest that multi-task learning is an effective approach for jointly capturing descriptive and action-oriented signals in student feedback, thereby supporting educational quality assurance in higher education.
Harnessing LLM as Partner and Mentor in Human Pragmatic Competency Training
ABSTRACT. Conversational skills go beyond the ability to decipher text and use words together in a proper grammatical structure to comprehend and convey an intended message. It entails pragmatic competence or the ability to use language effectively in a social context, manage turn-taking, adapt to topic shifts, adhere to cultural norms, and repair communication breakdowns to achieve communicative goal. Pragmatic competence training can be facilitated through role-playing where learners practice language and navigate contextual nuances by acting out social scenarios. In this paper, we describe the design and implementation of Socius, a conversational agent architecture that leverages a Large Language Model to provide pedagogical scaffolding. Socius offers a platform where learners can engage in role-playing conversation with an AI Partner and receive immediate, rubric-based feedback from an AI Mentor. The system was evaluated over a 5-day usage period by young adults aged 18-24 years old (N=23). Quantitative results from a 5-point Likert scale indicate high user acceptance across three metrics: performance (M = 4.542), humanity (M = 4.468), and affect (M = 4.519). Qualitative feedback indicates that students perceived the agent as increasingly human-like over time, with emotional engagement correlating closely with the perceived quality of the AI’s role-playing and feedback. The AI Mentor scored participants’ pragmatic competency as very good at logical intent, proficient in achieving communicative goals, but struggle with politeness and emotional tone. These findings demonstrate that LLM-driven dual-role agents can provide adaptive, context-sensitive environments essential for mastering pragmatic competence while balancing universal pragmatic rules with the informal reality of local dialects and personal identity.
ABSTRACT. Large Language Model (LLM)-driven Socratic dialogue may deepen cognitive engagement in Robot-Assisted Language Learning (RALL), yet its usability cost across distinct facets remains underexplored. This exploratory study deployed three AI agent designs—Direct Feedback, Socratic, and Adaptive Socratic—on a Robot-and-Tablet (R&T) platform with 32 ninth-grade English as a Foreign Language (EFL) students in Taiwan. A multidimensional usability analysis revealed significant group differences across all three dimensions—Usefulness, Easiness, and Satisfaction—with the Socratic agent consistently scoring the lowest; the deficit was most consistent on Usefulness, the only dimension on which both pairwise contrasts remained significant following a Bonferroni correction. Decomposing Easiness revealed an asymmetric pattern: the deficit was largely confined to Ease of Use (p = .012), with no significant difference observed in Ease of Learning (p = .163); because the platform interface was held constant across conditions, this asymmetry locates the cost in the Socratic interaction pattern rather than the shared interface. The Adaptive Socratic agent (which directed challenges to unmastered items via cumulative error-tracking) did not differ significantly from the Direct Feedback agent across any usability dimensions; however, this null pattern cannot establish equivalence at the present sample size, and direct contrasts between the Socratic and Adaptive Socratic agents were only marginally significant after statistical correction. Adaptive targeting is therefore advanced as a design hypothesis for mitigation rather than as an established empirical effect. Overall, students' English learning attitudes improved significantly (p = .001), with no agent-specific differences detected. Together, these findings reframe Socratic difficulty as a dimension-specific usability cost in AI-tutored RALL, selectively reducing Ease of Use with no detectable effect on Ease of Learning, and motivate further investigation of adaptive content targeting as a candidate mitigation.
Exploring Heterogeneous Treatment Effects with Educational Real-World Data: Promise and Scaling Challenge
ABSTRACT. Educational findings often lack generalisability due to substantial learner heterogeneity, shifting focus from “what works?” to “who benefits?” in recent years. However, attempts to estimate heterogeneity of treatment effects (HTE) remain scarce in education, and prior work relies on costly large-scale experiments, which also face ethical challenges. To address this issue, this paper explores the potential of their estimation by the secondary use of real-world data (RWD), an approach that remains underexplored in the literature. Our case study uses RWD of secondary-school students collected in the LEAF (Learning and Evidence Analytics Framework) system, including log data of reading behaviour and learning outcomes (𝑁=357), and focuses on the effect of bookmark revisiting behaviour on learning outcomes. Because learning is complex and dynamic, we use a non-parametric machine learning method for estimating individual conditional treatment effects. We find that, despite no significant average treatment effect, students with frequent revisits and many sessions tend to benefit from bookmark revisiting behaviour. This suggests the potential impact of encouraging them to effectively use bookmarks. However, the model fails to provide robust statistical evidence for HTE, likely due to the limited sample size and quality of covariates. Overall, the case study demonstrates the potential of the secondary use of educational RWD, while highlighting the clear need for scaling up RWD sharing in terms of volume and variety. To promote data sharing, we release the dataset used in this study to the research community using ReLEAF, a recently proposed trustworthy data sharing platform.
Using Eye-Tracking Learning Analytics to Examine How Student Activity Sheets Support Micro-Course Learning
ABSTRACT. Micro-course learning offers flexible and personalized learning opportunities, yet learners often experience divided attention and limited engagement. This study examines how Student Activity Sheets (SAS) support attention and learning in micro-course environments through eye-tracking-based learning analytics. Seventy-seven elementary school students were randomly assigned to a worksheet group, a graphic organizer group, or a control group without SAS. Eye-tracking indicators, including first fixation time, number of long fixations, and total fixation duration, were used to capture learners’ attentional processes, while retention and transfer tests assessed learning outcomes. Results showed that both SAS formats significantly improved learners’ attention and learning performance compared with the control condition, with no significant differences between the two SAS formats. Mediation analysis further revealed that attention, particularly total fixation duration, mediated the relationship between SAS use and learning outcomes. These findings provide process-level evidence for the role of SAS in scaffolded micro-course learning and demonstrate the value of integrating eye-tracking analytics with learning assessment to understand attentional mechanisms in technology-enhanced learning environments.
From generation to revision: Integrating peer assessment into generative AI-supported instructional design for pre-service teachers
ABSTRACT. The increasing use of generative artificial intelligence (GAI) in education raises questions about how AI tools can be pedagogically orchestrated to support meaningful learning rather than automate task production. In teacher education, pre-service teachers must not only use AI but also critically revise and refine instructional designs. This study examined a peer-assessment-enhanced GPT-supported condition (PA-GPT) to investigate its effects on pre-service teachers’ instructional design learning. A quasi-experimental study involving 82 pre-service teachers compared PA-GPT (n = 43) with a comparison ChatGPT-supported condition (C-GPT, n = 39). Results indicated that PA-GPT was associated with significantly higher performance in completeness, innovation, and interaction, as well as higher critical thinking tendency and self-efficacy, and lower cognitive load. Interview findings further suggested that peer feedback and comparison prompted learners to reconsider design rationales and compare alternative design choices. The findings highlight the added value of structured peer assessment within a generative AI- and concept map-supported instructional design environment.
Team-based Learning to Teach AI in Large Classes: Case Studies of an Undergraduate Interdisciplinary Course and a Postgraduate Engineering Course
ABSTRACT. This paper examines the deployment of Team-Based Learning (TBL) in teaching Artificial Intelligence (AI) through two case studies in large university classes: an interdisciplinary undergraduate course and a technical postgraduate engineering course. Utilizing TBL as the core pedagogical framework, both courses leveraged readiness assessments, team-based missions, and a purpose-built collaborative environment. By analyzing scores and feedback from students and instructors, we evaluate the efficacy of TBL in these contrasting contexts. Our findings demonstrate that TBL scales effectively for large cohorts while fostering interdisciplinary dialogue, integrating technical and societal perspectives, and cultivating essential soft skills. This work contributes to the field of AI education by illustrating how TBL transcends traditional lecture-based delivery to promote collaborative, ethical, and practice-oriented learning.
Designing Tablet-Based Augmented Reality Resources to Support Mobile and Authentic Inquiry in Primary Science: A Design Case Study on Simple Machines
ABSTRACT. Tablet-based augmented reality (AR) provides a new pathway for developing resources that support authentic problem-based inquiry in primary science classrooms. To address the difficulty of sustaining variable manipulation, evidence collection, and collaborative explanation in conventional classrooms, this study adopted a design case approach and developed a set of tablet-based AR resources for primary science inquiry, using the Grade 5 topic of “simple machines” as an example. Drawing on scientific inquiry, situated learning, and multimedia learning theories, the resources were designed around four consecutive components: problem, operation, feedback, and explanation. Three scenarios, namely the inclined plane, lever, and wheel and axle, were developed through a mapping relationship among authentic problems, scientific variables, AR interactions, feedback evidence, and explanation tasks. The resources were implemented using Unity and Vuforia, with marker cards used to trigger tabletop AR scenarios. Students could manipulate variables in groups, observe visual and numerical feedback, and construct explanations. Expert and teacher evaluations, a student technology acceptance pre-experiment, and a formative classroom trial indicated that the resources were feasible for use in primary science classrooms. They also revealed key design issues, including variable visibility, operation-feedback coupling, recognition stability, and device sharing. This study further extracts preliminary design principles for tablet-based AR science inquiry resources, providing transferable design references for the development of mobile AR resources oriented toward authentic problem contexts.
When, What, and How Long to Study: Capturing the Domain-Specific Dynamics of Learning Efficiency
ABSTRACT. Using behavioral logs from 313 Japanese high school students over an eight-month period in mobile learning across mathematics, chemistry, Japanese, English, and physics, this study examines learning efficiency variations by subject. Linear mixed-effects models show that optimal study times are subject-dependent rather than universal. Session duration reveals a U-shaped curve for all subjects—where efficiency initially drops before recovering in longer sessions. The effect of subject sequencing is domain specific. While switching into English from a different subject helps maintain efficiency, repeated engagement in English undermines efficiency. These findings offer empirical parameters for optimizing study schedules tailored to subject-specific cognitive demands and temporal rhythms, informing adaptive learning systems and self-regulated learning strategies to enhance learning efficiency and academic outcomes in mobile learning contexts. Limitations include missing psychological variables and fixed time-slot categorization, suggesting future research should integrate mixed methods and data-driven temporal segmentation. Overall, this study advances the understanding of multi-subject learning dynamics in ubiquitous environments, providing actionable insights for educators, learners, and developers to foster effective time management in K-12 education.
CodeRunner Agent: An Open-Source Moodle Plugin for LLM-Powered Self-Regulated Learning Support in Programming Practice
ABSTRACT. Large Language Models (LLMs) have created new opportunities for scalable programming support, but many existing tools prioritize direct coding assistance and remain external to institutional learning environments. These limitations make it difficult to provide context-aware feedback, cultivate Self-Regulated Learning (SRL), and give instructors visibility into how students interact with AI assistance. Building on our prior conceptual design and interview study, this paper presents the implementation, public release, and preliminary evaluation of CodeRunner Agent, an open-source LLM-powered SRL scaffolding system integrated with the Moodle CodeRunner plugin. The system combines an embedded student support panel, SRL-oriented AI scaffolding, automatic SRL phase detection, code guardrails, and an instructor dashboard for real-time learning analytics. The student interface supports four operational phases: Planning, Program Creation, Error Correction, and Self-Reflection. We conducted a preliminary evaluation with 12 students and 8 instructors using interaction logs, questionnaires, and open-ended feedback. Log analysis showed that students used the system across multiple SRL phases, especially for error correction. Students and instructors reported generally positive perceptions of the system’s usability, context-aware support, and hint-based feedback. These findings provide design implications for embedding SRL-oriented LLM support into existing programming learning environments.
Implicit Adaptive AI Matching for Nonviolent Communication Education: A Quasi-Experimental Study with Middle School Students
ABSTRACT. Nonviolent Communication (NVC) offers a structured framework for empathetic conflict resolution, yet internalizing its principles into spontaneous behavior remains a formidable challenge for adolescents. This study examined whether an "Implicit Adaptive Matching" mechanism — in which a custom AI dialogue system covertly diagnoses a learner's immediate psychological need (emotional support vs. problem-solving) and routes them to a matched AI persona (Warm vs. Competent) — produces superior NVC learning outcomes compared to a standardized baseline AI. Using a nonequivalent pretest-posttest quasi-experimental design, 86 valid samples of eighth-grade students were retained for final analysis after a two-week intervention. A manipulation check confirmed that students perceived the intended AI persona differences (t(54) = 2.754, p = .008). An ANCOVA controlling for pre-test scores revealed a marginally significant advantage for the Adaptive AI group on the Nonviolent Communication Behavior Scale (NVCBS) post-test (F(1,83) = 3.632, p = .060, ηp² = .042). Qualitative analysis of chat logs identified the "Jackal language cognitive loop" as the core learning barrier, and demonstrated how the adaptive mechanism provided precisely tailored affective and cognitive scaffolding to overcome it. System Usability Scale (SUS) scores confirmed that interface usability was uniformly high across all conditions (F(2,108) = 0.645, p = .527), ruling out technology as a confound. These findings advance knowledge of AI design grounded in the Stereotype Content Model (SCM) for social-emotional learning (SEL) and provide a replicable blueprint for adaptive affective computing in middle school counseling contexts.
Disengagement-Aware Student Simulators: Auditing Learner States for Pre-Deployment Tutor Evaluation
ABSTRACT. Intelligent tutoring systems are routinely tested before classroom deployment with simulated students, yet most simulators model only cooperative learners. This leaves a systematic evaluation gap, because real tutoring logs also contain students who guess rapidly, stall on a single skill, or disengage for long periods. We present Disengagement-Aware Student Simulators (DAS²), a reproducible pre-deployment framework for auditing whether LLM-based tutors stay trustworthy across the full range of learner engagement. DAS² generates simulated students in five learner states (engaged, gaming, wheel-spinning, off-task, and mixed) and tests whether tutor scores and rankings are sensitive to those states. In a 100-session blinded audit of ASSISTments09, two independent coders reached substantial agreement (κ = 0.78; 84% agreement), and human-agreed labels aligned with rule-assigned labels (κ = 0.75), supporting their use as evaluation groups. Conditioning simulation on these labels reduced the simulated-to-real gamer correctness gap from 0.54 to 0.20, and a post-hoc calibration step narrowed the gaming and wheel-spinning correctness gaps to at most 0.05. Across 10-turn and 20-turn Arena evaluations of five LLM tutors, rankings stayed stable under the primary judge (Spearman ρ = 1.00), yet absolute judge scores varied meaningfully across learner states, and a weak-tutor injection recovered the expected order, placing GPT-4o last on engagement with non-overlapping confidence intervals. A transparent construct-validity audit lifted the weak relevance judge from ρ = 0.33 to ρ = 0.61, though composite agreement reached ρ = 0.82, below the pre-registered 0.85 target. We treat this score-versus-rank dissociation as the central finding: a stable leaderboard can mask qualitatively different levels of tutor support for disengaged students. DAS² offers an auditable, reproducible way to inspect learner-state coverage before deployment on a single dataset and judge; it does not measure classroom learning gains.
ABSTRACT. he Chinese Proficiency Test (HSK) is one of the most widely used assessments in international Chinese education. This study examines whether a large language model (LLM) can support reliable item generation for HSK Level 4 reading comprehension. We built an integrated pipeline with three linked stages: deriving item-writing principles, generating items with prompt engineering, and validating the results with real test data. Drawing on the HSK syllabus, the national grading standard (GF 0025-2021), a review of prior research, and interviews with three subject-matter experts, we distilled 35 item-writing principles and turned them into a structured prompt. We then ran three studies using GPT-5.4. Study 1 generated items from a fixed corpus and asked seven experienced HSK item writers to review them. Study 2 used the model to screen reading texts and compared its decisions against an expert gold standard. Study 3 administered an automatically generated test form and an authentic HSK form to 120 learners and compared the two forms. The expert review showed good agreement (ICC = 0.82) and an average quality score of 2.32 out of 3. The model screened texts with an F1 score of 89.26 percent against the expert gold standard. In the field test, the generated form and the authentic form did not differ in a meaningful way (Cohen's d = 0.20), reached similar reliability (alpha = 0.78 versus 0.81), and showed comparable item difficulty and discrimination. About 78 percent of learners rated their overall test experience as positive. The findings suggest that an LLM-based pipeline can produce HSK Level 4 reading items of usable quality, while human review is still needed for distractor design and difficulty control.
Temporal Dynamics of Disengagement in E-Learning: Self-Regulated Reflections, Collapse Timing, and Learner Profiles
ABSTRACT. Although self-regulated learning (SRL) is generally assumed to promote persistence and achievement, highly self-regulated learners do not always maintain stable engagement in online learning environments. This study investigated the temporal relationship between SRL, perceived difficulty, and engagement collapse using longitudinal weekly learning records from university EFL learners (417 learners, 2,970 learner-week observations). Linear mixed-effects modeling revealed a significant negative relationship between SRL and completion rates. This relationship was initially moderated by perceived difficulty, such that higher SRL predicted lower completion more strongly under lower perceived difficulty conditions. However, robustness analyses excluding complete non-participation episodes (0% completion) showed that the interaction disappeared, suggesting that the observed effect primarily reflected emerging disengagement rather than active learning behavior itself. To examine temporal disengagement processes, survival analyses were conducted using individualized baseline engagement levels. Cox proportional hazards modeling indicated that higher SRL significantly predicted earlier collapse timing (HR = 1.39, p < .001), while learner clusters demonstrated distinct collapse risks and engagement trajectories. These findings suggest that SRL-related reflective activity may function as a context-sensitive and temporally reactive signal associated with engagement instability and impending disengagement in online learning environments.
A Cognitive-Model-Based Workflow for AI-Generated Visual Resources Supporting Literary Reading
ABSTRACT. This study proposes a cognitive-model-based workflow for generating AI-based visual resources to support literary reading. Existing educational applications of generative AI have often emphasized practical efficiency rather than the cognitive processes involved in literary comprehension. To address this issue, the study introduces “visual resources” as visual information supporting readers’ construction of meaning during reading. The proposed workflow integrates situation model theory, Dual Coding Theory, and human review to connect literary analysis, drawing policy generation, and image generation. The workflow emphasizes balancing cognitive support and interpretive openness so that generated images assist emotional inference without excessively constraining interpretation. Through an exploratory application to Japanese literary texts and a lexical analysis of teachers’ impressions, the study discusses the educational possibilities and limitations of AI-generated visual resources in literary education.
Developing K–12 Teacher Readiness for Data-Informed Teaching Through Professional Development
ABSTRACT. School teachers are expected to utilise data in instructional decision-making, but many teachers report facing limited preparation. This study evaluated a professional development (PD) programme on teacher data literacy in Singapore using an explanatory sequential mixed-methods design. Quantitative data was collected from N = 175 teachers, with matched pre-post data obtained from n = 29 participants across three PD cohorts. Semi-structured interviews with n = 10 teachers were subsequently conducted to contextualise quantitative findings. Quantitative analyses revealed statistically significant improvements in teacher self-efficacy across all three cohorts, alongside significant gains in data knowledge for two cohorts. Changes in perceived value of data use were more variable, with only one cohort demonstrating significant improvement. Qualitative findings indicated that hands-on analysis of authentic classroom datasets and collaborative discussion supported confidence development and engagement with data use. However, teachers continued to report difficulties translating analytical findings into concrete instructional strategies. Participants also identified structural barriers including limited time, competing responsibilities, and insufficient mentoring support. Findings suggest that PD may support proximal knowledge development more readily than deeper pedagogical or dispositional shifts.
ABSTRACT. Federated learning enables privacy-preserving analysis of educational data across multiple institutions, but its performance is often degraded by non-independent and identically distributed (non-IID) data across clients. While many existing studies have addressed this issue through model aggregation strategies, the role of feature representation has received limited attention. This study investigates differential features, which represent pairwise differences in learning-log features and grade-related labels within each client, as a representation-level approach to reducing distributional discrepancies in federated learning. Using learning log data from more than 1,400 students across 17 courses in an electronic textbook system, we analyze how differential features affect three types of data heterogeneity: label skew, feature skew, and quantity skew. The experimental results show that differential features substantially reduce inter-client discrepancies in pairwise label-difference distributions, as measured by JSD from 0.37 to 0.25 and HD from 0.39 to 0.26, and in feature distributions, as measured by JSD from 0.63 to 0.40 and HD from 0.59 to 0.36. Differential features also improve prediction performance in a ranking-based at-risk student prediction task, increasing average nDCG from 0.80 to 0.85 and PR-AUC from 0.66 to 0.72 on held-out clients. However, they can amplify quantity skew due to the quadratic growth of pairwise samples, indicating a trade-off among different types of data heterogeneity. These findings suggest that differential features provide a promising representation-level approach for mitigating non-IID effects in federated learning for at-risk student prediction.
Nonlinear Associations Between Primary–Auxiliary Screen Time Allocation and Mathematics Performance in Digital Assessment
ABSTRACT. Digital mathematics assessments provide screen-level process data that capture how students regulate attention and coordinate information across functionally distinct task screens during mathematical problem solving. Using TIMSS 2023 Grade 4 digital mathematics data from nine countries, this study conceptualized screen-level time allocation as an observable indicator of representational coordination within mathematics task blocks. The analytic sample included 84,811 student-by-block records from 42,528 students. Screens within each block were classified into primary cognitive screens and auxiliary representation screens. Cognitive–representation allocation was operationalized as the proportion of eligible block time spent on auxiliary representation screens. The study examined overall allocation patterns, compared linear and quadratic allocation models for two performance outcomes, and tested whether first-pass processing and revisiting moderated these associations. Students spent most block time on primary cognitive screens, while auxiliary representation screens accounted for a smaller but meaningful share. Quadratic models fit better than linear models for both outcomes, indicating an inverted-U association between allocation and performance. Predicted performance was highest when students maintained a moderate balance between primary task processing and auxiliary representation use. Moderation evidence was limited and mainly appeared for revisit burden in positive-score performance. Heavier revisiting made higher auxiliary-screen allocation less favorable among positive-score observations. These findings suggest that screen-level time allocation reflects not only how long students spend on tasks, but also how they coordinate functionally distinct representations during digital mathematical problem solving.
Professional Identity Expectations and Behavioral Intention Toward AI-Assisted Teaching: Testing a Dual-Mediation Model of Self-Efficacy and Well-Being in Pre-service Teachers
ABSTRACT. This study examines how conventional identity expectations (CIE) and digital identity expectations (DIE) differentially influence pre-service teachers' behavioral intention (BI) toward AI-assisted teaching, mediated by professional self-efficacy (PSE) and professional well-being (PWB). Drawing on identity theory, social cognitive theory, and UTAUT, two studies were conducted with pre-service teachers in eastern China. Study 1 used PLS-SEM (N=425) to test the hypothesized path model. Study 2 evaluated the Teacher Identity Reflection and Expectation System (TIRES), a generative AI-powered intervention, through a quasi-experimental design (n=85, eight weeks). CIE unexpectedly and positively predicted both PSE and PWB, functioning as a psychological resource rather than a constraint, while DIE exerted a significant direct effect on BI. PSE emerged as a robust mediator for both identity dimensions, whereas PWB did not reach significance as a mediator. TIRES produced sustained gains in PSE and PWB and a protective effect against decline in the control group. These findings position professional identity expectations as theoretically necessary upstream antecedents in AI adoption models.
Integrating AI in Subject Teaching for Pedagogical and Metacognitive Growth: A Location and Direction Writing Chatbot for In-Service English Language Teachers
ABSTRACT. As AI rapidly permeates K–12 education, the transition from basic tool acquisition to effective pedagogical integration remains a significant challenge. This mixed-methods study evaluates the effectiveness of a targeted teacher development (TD) workshop designed to enhance the subject-specific TPACK and metacognitive capacity of 116 in-service primary English teachers in Hong Kong. The four-hour intervention utilised a customised, curriculum-aligned AI chatbot specifically designed to scaffold the teaching of writing about location and direction. Data were collected through pre- and post-workshop surveys and written qualitative reflections. Wilcoxon signed-rank tests revealed statistically significant improvements (p < .001) across all TPACK sub-domains—CK, PCK, TCK, and TPACK—and all sub-domains of metacognitive capacity. Thematic analysis of teacher reflections corroborated these findings, highlighting empowerment to adopt new pedagogical approaches and a recognised shift from traditional instruction towards fostering student metacognition and autonomy. This research demonstrates that customised, subject-specific TD is essential for moving beyond generic AI use. Future research should investigate the long-term effects of such interventions on classroom practice and student outcomes.
Evaluation of a foundational artificial intelligence course on machine learning for executive professionals
ABSTRACT. Establishing a foundation of artificial intelligence (AI) literacy for all has become a critical necessity in societies where AI has fully permeated the workplace. It is vital to demonstrate practices that foster AI literacy, guiding participants from an introductory understanding to a deeper understanding of machine learning (ML). This study evaluates a 39-hour foundational AI course offered within a university programme for executive professionals, examining its impact on their understanding of ML. Employing a mixed-methods approach, data from 32 participants were analysed through a 27-item ML test and reflective writing tasks, administered both pre- and post-course, to assess changes in their ML understanding. Quantitative data were analysed using the Wilcoxon signed-rank test, while qualitative reflections were subjected to in-depth textual analysis. The quantitative results revealed a statistically significant improvement in ML test scores, with a large effect size (p < .001, r = .72). Qualitative findings indicated a notable progression in participants' ML literacy, evolving from a general understanding of ML concepts to a structured technical taxonomy of ML. Specifically, participants advanced from recognising basic concepts such as supervised, unsupervised, and reinforcement learning to demonstrating a more sophisticated understanding of ML and its practical applications in real-world contexts.
From Organizational Innovation to Cognitive Mediation: A Cross-Case Analysis of AI in Education
ABSTRACT. This study constituted a qualitative multiple-case analysis of how Artificial Intelligence (AI) is integrated across two educational contexts: institutional-pedagogical implementation in secondary schools and cognitive-mediated learning in higher education. Grounded in a multiple-case design, the study combined within-case and cross-case analysis to explore how AI functions across these different educational environments. The first case study examined 12 Arab teachers and 12 Arab principals’ perceptions of Artificial Intelligence in Education (AIED) integration into Arab secondary schools in Israel. The second case study explored the experiences of 27 graduate students using AI-based pedagogical chatbots and scenario-based simulators in self-directed learning environments. Findings show AI was perceived differently across contexts: in secondary schools, AI emerged as an organizational and pedagogical innovation requiring leadership, infrastructure, and ethical regulation, whereas in higher education, it functioned as a cognitive mediator supporting reflection, scaffolding, self-regulation, and autonomous learning. The findings suggest AIED represents a multidimensional educational transformation operating simultaneously at organizational, pedagogical, and cognitive levels.
Evaluating Game-Based Learning in Asynchronous AI Literacy: Impacts on Cognitive, Behavioral, Affective, and Ethical Development
ABSTRACT. The widespread integration of AI literacy in higher education remains challenging, particularly when it comes to helping beginners grasp abstract AI related-concepts. Although game-based learning (GBL) has been explored in prior research, its application in compulsory, asynchronous AI courses for diverse undergraduate populations has received limited attention. This study examines a GBL-based AI literacy course by analyzing pre- and post-survey data from 311 first-year undergraduates, focusing on cognitive, behavioral, affective, and ethical learning outcomes. The results indicate significant improvements in cognitive, behavioral, and affective dimensions, but no statistically significant change in ethical learning. By providing large-scale empirical evidence, this study contributes to the literature and proposes instructional design for integrating AI literacy into higher education contexts.
ABSTRACT. The current paper aims to investigating the role of incorporating a humanistic approach into virtual reality (VR)-based humanities lessons by means of gamification. Gamification is described within the paper by analyzing three literary masterpieces, such as *To Kill a Mockingbird*, *The Kite Runner*, and *The Stranger*. These examples helped to determine the most important elements of gamified learning, such as narrative plot, role-playing, branching pathway, and trial-and-error approach. Gamified learning contributes to the development of a personalized approach to learning, empathy, and reflection; thus, it promotes a more efficient learning process. In addition, it was proven that VR-based learning in the humanities can be considered a gamified system contributing to the development of individuality, error handling, and identity formation. The described techniques promote the achievement of Sustainable Development Goals (SDG 4 and SDG 10). This framework shows how VR-based humanities education can be viewed as gamified education, where storytelling, acting, and problem-solving are used to encourage students and help them learn deeply.
Understanding How Gamification Guides and Sustains Student Engagement in a Quantified Self Dashboard through Trace-based Analysis
ABSTRACT. Learning Analytics Dashboards (LADs) hold promise for supporting learner self-monitoring and data-driven decision-making, yet learners frequently neglect dashboard access and associated data behaviors. Gamification has been proposed as a means to address this challenge, but how gamification elements relate to intended behaviors at the level of fine-grained interaction remains largely unexplored. This study analyzes 55,603 interaction logs generated by 110 K8 students over a 15-day gamified Quantified Self activity, alongside pre-survey data on data use efficacy and awareness. Using Transition Network Analysis (TNA) and mixed-effects logistic regression, we investigate how gamification elements are associated with expected behaviors on the dashboard and which engagement factors within a gamified dashboard predict daily dashboard access. TNA results show that the gamification node achieved the highest betweenness centrality, functioning as a recurring hub through which students navigated toward learning records, data analysis, and inquiry recording behaviors. Regression results identify mission completion rate and consecutive access days as positive predictors of next-day access. In contrast, cumulative dashboard access days showed a negative association, suggesting that novelty-driven engagement in the early stages of the program gave way to gradual disengagement over time. These findings offer actionable insights for gamification design oriented toward sustaining intended learning behaviors in student-facing LADs.
Slide-Based Lecture Video Generation with Semantic Gestures from a Single Teacher Picture
ABSTRACT. With the rapid expansion of online education, there is an increasing demand for video content in which teachers provide explanations alongside lecture slides. However, the production of such videos requires high costs and intensive labor. In this paper, we propose a novel method for automatically generating lecture videos with a teacher avatar performing semantic gestures synchronized with the spoken content. The system requires only lecture slides and a single full-body picture as input. This research contributes to the field of online education by lowering the cost to lecture video production and enhancing the social presence of instructors due to the automated generation of expressive, context-aware movements.
The proposed system first utilizes a Large Language Model (LLM) to generate a lecture script based on the provided slides. This script then is used as the basis for generating skeletal motions that performs contextually appropriate gestures. Finally, by generating the avatar using the skeletal motion and the single full-body picture, the system produces a video in which the teacher avatar delivers a lecture with natural, expressive movements.
Subjective evaluations with 20 university students were conducted to assess the effectiveness of the proposed method. The results demonstrated that the lecture video with the teacher avatar significantly improved "Social Presence," "Understandability," and "Desire to Watch" compared to videos without an avatar, suggesting the pedagogical value of semantic gestures in automated lecture video production.
Exploring the Role of Personality Traits in Student Engagement with Generative AI
ABSTRACT. Generative artificial intelligence (GenAI) learning environments can both engage and disengage students, with personality traits playing a key role in shaping these outcomes. This study examines how the Big Five personality traits influence students’ behavioral, cognitive, and emotional engagement in GenAI-supported learning contexts. Data were collected from 486 university students in Hong Kong and Indonesia. The results indicate that personality traits significantly predict all three dimensions of student engagement, with varying strengths across dimensions. Personality traits exert the strongest influence on cognitive engagement, suggesting that individual differences shape how students process information and apply learning strategies when using GenAI tools. Behavioral engagement is also strongly influenced by personality, indicating that traits affect students’ participation and learning effort. In addition, personality traits show a significant positive effect on emotional engagement, although the effect is relatively weaker compared to cognitive and behavioral dimensions. Overall, the findings highlight the importance of considering personality differences when designing GenAI-supported learning environments. Tailored instructional approaches that account for diverse personality profiles may enhance student engagement and improve learning outcomes.
ABSTRACT. Educational AI models are increasingly used to support academic prediction and early intervention, yet model evaluation is often dominated by average predictive accuracy alone. Using seven waves of nationally representative China Family Panel Studies (CFPS) data from 2010 to 2022, this study examined academic prediction through a responsible educational AI framework emphasizing temporal transferability, urban–rural fairness, worst-group robustness, and explainability. Adjacent-wave prediction tasks were constructed to predict later academic performance from child, family, home learning, and prior achievement variables. Models were trained on five historical windows from 2010→2012 to 2018→2020 and evaluated on a fully held-out future window from 2020→2022. Conventional machine-learning models, deep learning baselines, a temporal domain-adversarial neural network (DANN), and the proposed FairTransEdu-GDRO framework were compared. Results showed that boosting-based models achieved the strongest overall predictive accuracy, whereas Ridge regression demonstrated the smallest urban–rural error disparity and best worst-group performance. Compared with DANN, FairTransEdu-GDRO reduced urban–rural prediction error gaps and improved worst-group robustness while maintaining competitive predictive performance. SHAP analysis further indicated that prior academic performance was the dominant predictive feature, followed by age, family educational expenditure, and learning-related self-regulation. The findings suggest that educational AI models should be evaluated not only by accuracy, but also by temporal reliability, subgroup fairness, robustness, and interpretability.
Equity-Aware Adaptive Tutoring for Open Educational Resources: A Systematic Review and Conceptual Design Framework
ABSTRACT. Open Educational Resources (OER) hold genuine promise for educational equity, yet a growing empirical record reveals that nominal access to open content does not automatically produce equitable learning outcomes—learners facing limited bandwidth, language barriers, or device restrictions remain systematically underserved even when materials are free. This paper makes two contributions. First, it presents a PRISMA-adapted systematic review of 61 peer-reviewed publications (2021–2026) spanning six thematic axes—OER equity, intelligent tutoring systems, AI fairness, reinforcement learning, user profiling, and explainable AI—and identifies three convergent research gaps: the absence of access-constraint modeling in adaptive tutoring agents; the conflation of recommendation with tutoring; and the lack of fairness-aware bandit formulations applied to tutoring strategy selection. Second, building on these gaps, it derives RL-Tutor (Reinforcement Learning Tutor), a conceptual design framework for an equity-aware tutoring agent that selects guidance strategies based on each learner's real-world access profile using a fairness-penalized contextual bandit algorithm. RL-Tutor is presented as a research blueprint, not an evaluated system; no experimental results are reported, and empirical validation is the explicit objective of future work.
ABSTRACT. Artificial intelligence is rapidly expanding into physical environments through Vision-Language-Action (VLA) systems. However, current primary robotics education remains largely rule-based, while existing AI education often emphasizes AI use rather than understanding AI reasoning. To address this gap, we developed PIE BRIDGE, a VLA-inspired educational web application for primary students. PIE BRIDGE serves as a pedagogical bridge between physical computing and physical AI education by simulating an inspectable VLA workflow through visible stages of Vision, Language, and Action. The system supports human–AI co-piloting by allowing students to review and revise AI decisions throughout the process. To accommodate classroom technical constraints and current vision model limitations, the system utilizes a discrete workflow rather than a continuous loop and restricts robot movements to a grid-board environment. A two-round modified Delphi study with ten experts yielded consensus across all domains, validating PIE BRIDGE as a pedagogically robust VLA-inspired learning environment for primary school students.
An Exploratory Comparison of Architectural and Parameter-Level Strategies for Pedagogical Remediation
ABSTRACT. Intelligent tutoring systems promise scalable personalized learning, yet reliably detecting and correcting student misconceptions through conversational AI remains fundamentally challenging. Effective tutoring requires confronting conceptual errors in real time, but language models aligned for conversational helpfulness exhibit systematic failure modes: they often affirm incorrect reasoning, avoid corrective intervention, and generate hallucinated instructional guidance. These behaviors undermine pedagogical fidelity, revealing a mismatch between general conversational alignment and the demands of education.
We investigate whether system design, rather than model scale alone, can improve tutoring behavior. Comparing zero-shot prompting, parameter adaptation, and a task-decomposed conversational pipeline that separates misconception identification from instructional response generation, we find that structured interaction design yields modest but consistent improvements, particularly for smaller models. Despite persistent limitations in robustness and naturalness, our findings suggest that effective conversational tutoring may depend less on larger models and more on pedagogically grounded system architectures.
Inverse Student Simulation for Auditing Structured Student State in Tutoring Dialogues
ABSTRACT. Large language model tutoring systems can generate student-state summaries that appear in teacher dashboards, influence problem selection, and shape feedback. Such a summary is a form of educational hallucination when it derives from the problem statement or a repeated annotation template rather than from what the student actually said. We present inverse student simulation (ISS), a trustworthy-AI audit framework for this risk. ISS estimates a structured student-state record from tutoring dialogue and then checks whether that record is recoverable from dialogue evidence and whether it changes a forward simulator’s predictions. The record, Z, encodes mastery probabilities for 30 knowledge components (KCs), 70 misconception-code probabilities, and four exploratory metacognitive scalars. Using MathDial dialogues, we find that broad human labeling is reliable but model recovery is uneven. Two coders labeled all 394 silver-labeled test dialogues and reached substantial to almost-perfect agreement (κ = 0.701–0.847). The fine-tuned ISS inverter shows only weak within-dialogue KC ordering (Spearman ρ = 0.274; top-3 overlap = 40.7%), and a label-free prefix heuristic achieves lower calibration error than the full-dialogue ISS model (Brier 0.028 vs. 0.038). Misconception scores are dominated by a degenerate four-label pattern shared across all 394 test dialogues; the inverter emits a near-constant misconception vector (across-dialogue σ ≈ 0.002) that a zero-parameter label-prior baseline matches without reading any dialogue. Metacognitive recovery is unreliable, with most scalar correlations near zero. Forward replay shows the full-context model predicts responses with little dependence on Z (mean NLL difference = 0.006), while Z-only probes confirm Z carries signal when context is removed. These are negative results, not a working model: under degenerate silver supervision the inverter collapses to a label prior rather than recovering student state, so ISS's contribution is the reusable, model-agnostic audit framework and this quantified failure-mode diagnosis. Labels must pass audit before reaching dashboards, problem selection, or feedback.
A Systematic Review on Spherical Video-Based Virtual Reality in Foreign Language Education Based on SAMR Model
ABSTRACT. Spherical video-based virtual reality (SVVR) has attracted increasing attention in second or foreign language (FL) education, but the extent of its pedagogical integration remains unclear. This study reports a systematic review of 23 empirical studies on SVVR in FL education retrieved from the Web of Science Core Collection in accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020. Drawing on the Substitution, Augmentation, Modification, and Redefinition (SAMR) model, the review examines both the research profile of the field and the levels at which SVVR has been integrated into pedagogical practice. The results show that SVVR research has expanded notably since 2020, with most studies conducted in East Asia and in higher education contexts. Mixed-methods and quasi-experimental designs were predominant, and speaking was the most frequently targeted language skill. In terms of SAMR classification, most studies were situated at the Augmentation and Modification levels, fewer reached Redefinition, and only one was classified as Substitution. These findings indicate that SVVR is commonly used to enhance or redesign language learning tasks, but more limited evidence exists for fully transformative uses. The study contributes an up-to-date synthesis of SVVR research in FL education and provides implications for the design of more pedagogically meaningful immersive learning activities.
Cognitive Offloading in AI-Supported EFL Writing: A Longitudinal Trajectory Analysis
ABSTRACT. Generative AI writing feedback systems provide immediate support for EFL learners, yet concerns remain regarding cognitive offloading, where learners delegate cognitive work to AI rather than actively engaging their own strategies. The study examined how 111 Japanese university students engaged with an AI writing feedback system over a 14-week English writing course. To assess changes in self-regulated learning and cognitive strategy application, engagement patterns were analyzed longitudinally at two levels: across six writing topics to track adaptation to the system, and within each topic to trace error correction trajectories across four linguistic categories. Results showed significant improvement in writing performance, but an overall decline in cognitive strategy use. Patterns diverged markedly by learning outcome groups. High-outcome students sustained cognitive strategy use and engaged in deep-level processing of syntactic and lexical errors, progressively reducing newly introduced errors. In contrast, low-outcome students showed narrowing engagement focused on surface-level corrections and increasing passive dependence on AI feedback. These findings suggest that AI writing feedback effectiveness depends critically on learners' regulatory engagement patterns, with prolonged, excessive cognitive offloading among less successful learners being associated with a decline in their self-reported general cognitive strategy use.
Designing AI Feedback to Visualize Learners' Strategy Achievement in a Strategy-Focused Speaking Practice System and Its Classroom Evaluation
ABSTRACT. Strategic competence, which is the ability to convey intended meaning through alternative expressions when appropriate vocabulary is unavailable, is essential for EFL learners but is often insufficiently supported in conventional speaking instruction. To promote three expressive achievement strategies (approximation, circumlocution, and restructuring), we developed HanaStrategy WebApp (HanaS), a browser-based platform implementing a strategy-focused speaking task: learners describe an illustration without using pre-specified "restricted expressions," a constraint that pushes them to search for and deploy the strategies as solutions. Between two attempts per task, HanaS provides AI feedback that visualizes how the learner's speech would be interpreted by a third party (rendered as an illustration), identifies the strategies used, and offers alternative model answers. The visualization is intended to make conveyance gaps salient and to elicit content-oriented reflection on what was insufficient. We conducted a two-month classroom implementation with 18 first-year undergraduates (CEFR B1-B2), comparing word count in the pre/post-speech tests, the number of strategies used in the first and last strategy-focused speaking tasks, and qualitatively coded reflections. Three findings emerged. First, word counts increased significantly from pre- to post-test. Second, circumlocution spread across more learners. Third, learners whose word count improved had named specific strategies less frequently in their improvement plans; content-oriented reflections were associated with greater gains, whereas naming strategies was associated with smaller gains. These findings suggest that AI feedback eliciting content-oriented re-expression, rather than mere strategy labeling, can scaffold strategic competence development.
Understanding the Dynamics of Social Comparison in K-12 Students’ Learning During Summer Vacation
ABSTRACT. This study investigated how social comparison functioned in a student-facing dashboard during K-12 students’ summer vacation learning. Although social comparison is widely used in dashboards, it remains unclear when students engage in it during learning, whom they compare themselves with, and whether such comparison leads to metacognitive awareness. To address these issues, we analyzed the learning logs of 236 Japanese seventh- and eighth-grade students during their summer vacation assignment, including both logs from the e-book system and logs of their use of a dashboard that allowed them to compare their learning processes with those of others. Transition network analysis was applied to each learning actions, and memo data were coded using eight components of metacognitive awareness. The results showed that social comparison was embedded in the learning flow and was repeatedly used in an exploratory manner. Most comparisons were made with classmates, while prior-cohort comparisons were less frequent and often limited to all or high-performing seniors. Low-performing students were more likely than high-performing students to compare themselves with others at the start of learning, suggesting reliance on others’ progress as an external reference. However, only a small number of students wrote reflective memos, and no clear difference was found between high- and low-performing students in the amount or type of metacognitive awareness expressed. These findings suggest that social comparison can support learning engagement and decision making, but dashboards should also provide support that helps K-12 students interpret comparison information and turn it into metacognitive awareness.
Developing an AI-Powered Video Analytic Tool for Monitoring Online Learner Emotional Engagement
ABSTRACT. Blended synchronous learning (BSL) is a highly practical instructional modality in higher education. However, it presents instructors with the challenge of monitoring the emotional engagement of online learners in real time. This paper presents an AI-powered Video Analytic Tool (AVAT), a system designed to monitor and analyze learner facial expressions during live Zoom sessions and provide instructors with real-time engagement data via a dashboard. The tool captures and classifies learning-relevant emotions at one-second intervals, distinguishing positive states (e.g., enjoyment, excitement) from negative ones (e.g., frustration, boredom, anxiety). A web-based prototype has been developed and integrated with the Zoom SDK and AWS cloud infrastructure. This paper presents the conceptual framework, system architecture, key features, and current prototype, and a plan for future validation and deployment in authentic BSL settings.
Causal Discovery from Real-World Educational Data via LLM-Informed Methods
ABSTRACT. While expectations for Evidence-Based Education continue to grow, significant challenges remain in the systematic extraction of evidence from Real-World Educational Data (RWED). Among various methodologies, causal discovery has emerged as a promising focus for the automated extraction of evidence from such data. However, conventional algorithms often face practical limitations, such as the necessity for large-scale datasets to achieve stability and an inherent vulnerability to unobserved confounding factors, which are common in educational settings. To address these limitations, this study proposes a novel methodology that integrates the vast relational knowledge of Large Language Models (LLMs) into the causal discovery process. Specifically, we incorporate causal relationships by an LLM as prior knowledge into LiNGAM-MMI, which is a causal discovery algorithm distinguished by its robustness against confounding and its effectiveness with small-scale datasets. We conducted comprehensive evaluation experiments using both standard causal analysis benchmarks and RWED. The results demonstrate that the proposed method achieves superior accuracy compared to existing baseline techniques on benchmark data and exhibits high compatibility when applied to complex educational datasets. These findings suggest that the integration of LLM-based prior knowledge with LiNGAM-MMI is a highly promising approach for the extraction of Real-World Evidence (RWE) from Real-World Data (RWD). The outcomes of this research are expected to facilitate the routine and continuous collection of RWE directly from educational practice, thereby accelerating the practical and social adoption of evidence-based education.
Development of a Platform for Sharing Experiential Knowledge Based on Thought Experiences in Research Activities
ABSTRACT. In research organizations, members acquire experiential knowledge through daily research activities. However, such knowledge is often embedded in individual thought processes and remains tacit, making it difficult for subsequent members to understand how their predecessors interpreted problems, made decisions, revised artifacts, and derived reusable knowledge. We develop an experiential knowledge sharing platform based on thought experiences. Designed with reference to the Dual-Loop Model, the platform supports the externalization and accumulation of thought experiences, as well as the stepwise formalization and sharing of experiential knowledge derived from them. The platform treats evolving research artifacts as thought products and records their changes as versions and histories. It then uses accumulated thought experiences as cues for formalizing experiential knowledge, which is shared within the organization together with its formation context. A practice in a university laboratory suggested that the platform can support the construction and sharing of experiential knowledge based on individual thought experiences.
Towards a Pedagogical Sandbox for TPD: A Tri-Level Alignment Study of a Data-Grounded LLM Simulator in CSCL
ABSTRACT. Proactive Teacher Professional Development (TPD) requires reliable environments for pedagogical rehearsal; however, orchestrating complex group dynamics in Computer-Supported Collaborative Learning (CSCL) imposes severe cognitive burdens on educators. Existing Large Language Model (LLM) simulations fail to provide such reliable rehearsals due to an “alignment tax”, where models auto-correct seeded student misconceptions, causing “logic hallucinations.” To bridge this fidelity gap, we propose a Data-Grounded Cognitive (DGC) architecture that anchors multi-agent behaviors in authentic, multimodal student profiles (cognitive anchors, metacognitive confidence, and social personality traits). We validated the system through a Tri-Level Alignment Study, benchmarking in-silico logs against empirical ground-truth from a remote CSCL session. Results demonstrate high individual and group alignment in foundational tasks; however, complex tasks reveal a decline in process alignment, empirically exposing the LLM alignment tax. Crucially, Lag Sequential Analysis confirms profound structural isomorphism between the simulated socio-metacognitive flow and authentic discourse hegemony. The DGC architecture effectively provides a high-fidelity “Human-on-the-loop” (HOTL) sandbox, advancing evidence-based TPD.
Developing Computational Thinking through Gamified Learning: A Mixed-Methods Intervention Study among Teacher Trainees
ABSTRACT. The development of Computational Thinking has become a key priority in contemporary education, particularly in teacher education where future educators are expected to design meaningful and problem-based learning experiences. Although gamification has been widely recognised for enhancing engagement, its role in supporting higher-order cognitive development remains insufficiently understood. This study investigates how CT develops within a gamified STEM learning environment among teacher trainees. An explanatory sequential mixed-methods design was employed involving 87 Bachelor of Education students enrolled in an Educational Gamification course. Quantitative data were collected using a 30-item performance-based multiple-choice instrument administered before and after a 14-week intervention. Qualitative data were obtained through semi-structured interviews to explore participants’ learning experiences and cognitive processes. The quantitative results indicated an improvement in CT performance, with mean scores increasing from 59.32 (SD = 11.55) to 75.48 (SD = 8.47), accompanied by a large effect size (d = 1.22). The qualitative findings identified four key themes: gamification as a cognitive structuring mechanism, the embedding of CT within game mechanics, the role of authentic STEM problem-solving contexts, and the transformation of teacher trainees’ pedagogical perspectives. The integration of findings suggests that the effectiveness of gamification extends beyond motivation, functioning as a structured environment that shapes how learners engage in problem-solving processes. By aligning game mechanics with CT components and situating learning within authentic STEM contexts, the intervention enabled meaningful and transferable cognitive development. This study contributes to the literature by providing empirical evidence on the mechanisms through which gamified learning environments support computational thinking and offers practical implications for instructional design in teacher education.
“Vesper”: An Automatic Evaluation System for Visual Elements in Statistical Posters Using CNN
ABSTRACT. This paper presents Vesper, an automatic evaluation system for visual
elements in statistical posters using convolutional neural networks (CNNs). In high
school Informatics education in Japan, the creation of statistical posters is sometimes included in the curriculum. Such posters are evaluated from both data utilization and information design perspectives; however, this system focuses specifically on the visual elements of information design. We construct CNN-based models to evaluate eight visual elements and develop a web-based system that assesses student-uploaded posters and provides feedback with confidence scores. A classroom-based field experiment showed that 97.6% of students rated the system as useful, indicating its effectiveness as a learning support tool.
Exploring the Initial Level of Prompt Engineering Among Primary School Students: A Participatory Observation Approach
ABSTRACT. With the increasing integration of Generative Artificial Intelligence (GenAI) into K-12 education, the initial level of primary school students' prompt engineering remains insufficiently explored through systematic empirical research. Using a participatory observation methodology, this article analyzes primary school students' self-directed human-AI interaction processes in the context of learning tasks from both behavioral and cognitive perspectives, examining their prompt engineering and exploring the relationship between different modes of student-generated prompt and the quality of AI-generated outputs. This article reveal that primary school students have already possessed a foundational level of prompt engineering and demonstrate distinctive cognitive and behavioral characteristics. Crucially, students who express prompts in a written style mode tend to obtain higher-quality generated outputs than those using spoken-style modes, which better supports the completion of learning tasks. Finally, attempting to break through generic prompt frameworks,this article proposes a two-dimensional “elements–features” framework based on the existing “starting point” of primary school students' prompt engineering. This framework provides theoretical and practical references for enhancing students' prompt engineering and for promoting GenAI applications from a student-centered perspective in K-12 education.
ABSTRACT. Collaborative learning depends not only on how much students speak, but also on how they respond to, build on, and influence one another's ideas. However, feedback for group discussions often relies on quantitative activity indicators, such as the number of utterances or posts, which may overlook students who speak less frequently but play important interactional roles. This paper proposes an LLM-based network analysis method for visualizing interaction-based student contributions in group discussions. The method segments discussion transcripts into topic units, extracts explicit relations among utterances with a large language model, and constructs a directed student network in which edges represent response relations such as agreement, disagreement, clarification requests, and evidence-based critiques. The HITS algorithm is then applied to calculate two indicators: Social Impact, which represents the extent to which a student's utterances elicit responses from others, and Responsivity, which represents the extent to which a student responds to others' utterances. We applied the method to six discussions conducted by 12 undergraduate and graduate students and compared the results with expert ratings, self-ratings, peer ratings, LLM zero-shot ratings, and utterance counts. The results showed that LLM zero-shot ratings corresponded more closely with expert ratings, whereas the proposed method showed lower correlations with utterance counts. A questionnaire further showed that LLM zero-shot feedback was easier for students to understand, while the proposed network-based feedback helped some students notice response relations among members that were difficult to recognize during discussion. These findings suggest that the proposed method should be positioned not as a replacement for expert or LLM holistic ratings, but as complementary structural feedback for supporting students' reflection on collaborative interactions.
Characterization of Learners at Different Levels of Collaborative Task Complexity: A Three-dimensional Perspective Based on Cognition, Social Interaction, and Emotion
ABSTRACT. Collaborative task complexity plays a critical role in shaping learner engagement, yet its effects on cognitive, social, and emotional characteristics remain underexplored. This study employs a multimodal learning analytics approach to examine how learners respond to tasks of varying complexity. Fifty-four undergraduates in triads collaboratively solved both simple and complex algorithmic problems. Data were collected across three modalities: heart rate variability (for emotion), self-reported cognitive load, and transcribed verbal behaviours. Using a two-level coding framework, cognitive and social behaviours were categorized alongside collaborative problem-solving (CPS) stages, followed by lag-sequential analysis to identify discourse patterns. Results showed that complex tasks induced higher cognitive load and more intricate transitions among cognitive behaviours, particularly Seeking Information and Clarifying. In contrast, simple tasks fostered more frequent and reciprocal social behaviours. Surprisingly, complex tasks also elicited more positive and fewer negative emotional responses, suggesting challenge may enhance affective engagement. These findings highlight the need for adaptive task design that integrates cognitive challenge, social scaffolding, and emotional regulation.
When AI Joins the Group: Sequential Discourse Dynamics in Speech-Based AI-Mediated Collaborative Learning
ABSTRACT. Although generative artificial intelligence (AI) is increasingly used in collaborative learning, existing research has mainly examined text-based systems or AI as a scaffold, leaving limited understanding of how speech-based AI shapes real-time group discourse. This study examines sequential discourse dynamics in triadic collaborative discussions in which two students interact with a speech-based AI positioned as a third group member. Grounded in hybrid intelligence and human–AI shared regulation perspectives, the study conceptualizes AI participation as part of a shared interactional system in which student and AI discourse moves may mutually shape one another. Data were collected from 66 online collaborative discussions involving 132 undergraduate students. Sessions were video- and audio-recorded, transcribed, segmented into turns, and coded for communicative function. Descriptive statistics were used to examine discourse distributions, and lag sequential analysis was conducted to identify significant transition patterns between student and AI moves. Results showed a broadly balanced but asymmetric interactional structure. Student questions strongly elicited AI answers, while AI questions opened space for student answers, further questions, and information-sharing. AI answers and information frequently led to subsequent AI questions, suggesting a provision–elicitation cycle. Socio-emotional exchanges formed a reciprocal discourse channel, and student information-sharing supported continued knowledge-building. These findings suggest that speech-based AI functioned not only as an information provider but also as a conversational participant that shaped the temporal organization of collaborative discourse. The study informs the design of AI agents that balance provision, elicitation, restraint, and socio-emotional responsiveness while preserving student agency.
ABSTRACT. Productive failure pedagogy intentionally places learners in challenging problem‑solving situations prior to formal instruction, making emotional experience and regulation central to learning. In a secondary school classroom study (N = 159), I examined which emotions students experience during the initial collaborative problem-solving phase, how they naturally regulate emotions, and how experimentally guiding savoring or dampening of the most strongly felt emotion affects regulation uptake and learning. Results show that students’ emotional experiences are functionally organized rather than valence‑based, with epistemic emotions like interest and confusion playing central roles. Most students do not naturally engage in emotion‑aware problem-solving, instead relying on emotion‑minimizing or neglecting strategies. Although dampening negative emotions was easier to interpret than savoring, intuitive regulation strategies did not translate into delayed learning benefits. Critically, posttest learning showed a significant interaction between emotional valence and regulation condition: dampening supported learning when emotions were negative, whereas no regulation was the most effective when emotions were positive. Findings demonstrate that intuitive emotion regulation is not necessarily productive during failure‑driven learning and highlight the need for emotion‑aware scaffolds aligning regulation with epistemic goals.