ICCE 2026: THE 34TH INTERNATIONAL CONFERENCE ON COMPUTERS IN EDUCATION 2026
PROGRAM FOR FRIDAY, DECEMBER 4TH
Days:
previous day
all days

View: session overviewtalk overview

09:00-10:00 Session 31: Keynote Speaker 4

Keynote Speaker (C6: TELL Chun Lai)

Location: Savoy Ballroom
10:00-10:20Coffee Break
10:20-11:20 Session 32B: C1 Session K
Location: Savoy 2
10:20
Development and Evaluation of a Reflection Support System for Learning from Mathematical Errors
PRESENTER: Yuna Yamaguchi

ABSTRACT. Lesson induction, a learning strategy that enables learners to learn from their errors in problem solving, has been shown to be effective. However, deriving high-quality lessons independently remains challenging for many learners. To address this issue, this study proposes and develops an interactive lesson induction support system using LLMs. The system recognizes learners’ handwritten answers and guides their reflection through a four-stage dialogue based on cognitive counseling. A practical evaluation in mathematics with junior high school students showed that the quality of the lessons significantly improved among learners who completed the dialogue through the final stage.

10:35
Breaking Passivity in Video Learning: A System Integrating Physical Social Robots and Time-Adaptive RAG
PRESENTER: Kota Hashiyada

ABSTRACT. On-demand video lectures frequently induce a passive learning state, leading to a loss of concentration and an "illusion of competence". Furthermore, a significant problem exists in that learners rarely take the initiative to seek help—such as by conducting independent research or asking questions—even when they encounter difficulties during the learning process. To solve this problem, interactive systems utilizing virtual agents and Large Language Models have been proposed. However, they lack the physical embodiment often considered necessary to assist in re-engaging distracted learners. Furthermore, these digital systems struggle to chronologically synchronize the LLM's knowledge base with the learner's real-time progress through the video material. To address these challenges, this study proposes a novel active learning support system that integrates a physical social robot (Kebbi) with a "Time-Adaptive Retrieval-Augmented Generation" approach. The system automatically pauses the video at educationally significant moments and proactively intervenes using the robot's physical gestures (e.g., head tilting) to elicit spontaneous questions. Furthermore, Time-Adaptive RAG dynamically restricts the LLM's reference knowledge to the transcript of the video watched up to that exact moment, enabling contextually synchronized dialogue. A preliminary feasibility study with eight university students was conducted to evaluate the system. The quantitative and qualitative results provided initial support for the feasibility and acceptability of the proposed system. First, the dialogue logs suggested that Time-Adaptive RAG generated contextually appropriate responses aligned with the learner’s current video progress (H1). Second, the proactive interventions appeared to encourage spontaneous questioning, while head pose data showed no significant increase in postural collapse (H2). Third, the combination of automatic video pausing, the robot's subtle physical affordances, and seamless progression control based on LLM intention estimation was generally accepted by participants without obvious signs of fatigue or visual deviation (H3). In conclusion, integrating a physical robot with temporally synchronized LLM dialogue suggests a feasible and acceptable approach to transforming passive video viewing into an engaging, active learning experience.

10:50
Problem Recommendation Based on Skill–Challenge Balance: Predicting Learner Confidence and Challenge with LLMs and Learner Profiles

ABSTRACT. This study defines an individually optimal problem as one in which a learner’s perceived confidence and challenge are balanced and preliminarily evaluates a method for recommending such problems. This approach is grounded in flow theory, which suggests that learners become most engaged when challenge and skill are well balanced. Thus, recommending individually optimal problems not only requires considering a problem’s objective difficulty, but also learners’ subjective perceptions of difficulty. To address this issue, we propose a recommendation method based on Large Language Models (LLMs) that constructs learner profiles from a small set of diagnostic problems and uses these profiles to predict learners’ confidence and perceived challenge for unseen problems. Based on flow theory, we define Gap as Confidence − Challenge and recommend problems with a small |Gap|. A preliminary evaluation was conducted using problems targeting a single knowledge component in linear algebra, specifically elementary row operations and the degrees of freedom of solutions. The results indicate that the LLM can predict learner confidence with moderate accuracy and that the proposed prompt incorporating learner profiles improves the prediction accuracy of both Confidence and Gap compared with a baseline prompt that does not use learner profiles. In contrast, challenge prediction accuracy remained limited because perceived challenge strongly depended on problem-specific characteristics and because a mismatch existed between the procedural load assumed in the prompt and the subjective ratings provided by participants, who did not actually solve the unseen problems.

11:05
NLP-based Auto-coding and Sequential Analysis of Cognitive Engagement in Elementary Science Classes

ABSTRACT. This study developed and validated a Natural Language Processing (NLP) framework for the automated assessment of students' cognitive engagement in offline elementary science classes. A coding scheme was first constructed by integrating the ICAP framework with Bloom’s Taxonomy and refined through the Delphi expert method. Subsequently, speech data from 22 classroom videos (totaling 850 minutes) was transcribed and segmented. A fine-tuned Llama-2 large language model was then employed to automatically classify the 5,241 derived text segments according to the coding scheme.

Validation against human coding demonstrated satisfactory reliability, with an accuracy of 80.89% and a high intraclass correlation coefficient (ICC) of 0.85. The analysis of the automatically coded data yielded three main findings: (1) "Active" engagement was most frequent (41.02%), while higher-order "Constructive" engagement was relatively low (10.59%); (2) Lag sequential analysis revealed significant cognitive-state transitions, such as "Receiving → Discussing" and "Elaborating → Questioning"; (3) A clear developmental trajectory was observed, with engagement shifting from passive reception in lower grades towards more active application and social discussion in higher grades.

The results confirm the feasibility and reliability of using advanced NLP for automated, fine-grained analysis of classroom dialogue. This work provides a practical tool and empirical insights for evidence-based classroom evaluation and the enhancement of instructional practices in science education.

10:20-11:20 Session 32C: C7 Session G
Location: Savoy 3
10:20
Reconsidering On-the-Job Training (OJT) as Onboarding Support: A Job Crafting Perspective

ABSTRACT. HHistorically, Japanese companies have relied heavily on On-the-Job Training (OJT) for human resource development. However, recent demographic and social changes have contributed to rising turnover among early-career employees, creating a growing need to reconsider OJT from a broader developmental perspective. This study reconceptualizes OJT as a form of onboarding support and explores its potential from a job crafting perspective. An exploratory case study was conducted in a Japanese small and medium-sized enterprise (SME). The findings suggest that OJT can support workplace adaptation not only by transferring job-related knowledge and skills but also by fostering relationship building, meaning-making, and proactive learning. These findings highlight the potential of understanding OJT as an onboarding practice that supports employee development and adaptation.

10:35
Who is responsible when GenAI enters teaching? Teachers' responsibility boundary work in GenAI integration

ABSTRACT. Generative AI (GenAI) is entering lesson preparation, feedback, assessment support, and student learning tasks, making teachers' work not only more efficient but also more accountable. This study investigates how teachers draw responsibility boundaries when GenAI becomes part of everyday teaching. Based on interviews with 27 teachers from primary, junior secondary, senior secondary, secondary vocational, and university contexts, we used constructivist grounded theory to analyze critical incidents, responsibility talk, and situated judgment. The analysis generated a model of responsibility boundary work with six interrelated components: task sorting, conditional trust, responsibility backstopping, process evidence, hidden labor, and institutional negotiation. Teachers did not simply adopt or resist GenAI. They allowed it to enter low-risk, reversible, and checkable tasks while retaining human judgment in assessment, safety, value-laden decisions, and student development. They treated GenAI output as a draft, reference, or conjecture, and used process evidence to repair assessment fairness when final products no longer reliably represented student capability. The model shows that responsible GenAI integration requires not only individual teacher AI literacy but also school-level rules, disciplinary review routines, student declaration mechanisms, and professional development focused on situated responsibility judgment. The study contributes to practice-driven research by linking GenAI classroom use with teacher professional judgment and institutional support.

10:50
Toward an Ideal Teacher-Researcher Relationship in Learning Analytics: Policy Delphi for Developing Practical Research Policies

ABSTRACT. This study examines the collaborative relationship between Learning Analytics (LA) researchers and in-service teachers in the field of learning analytics research. Although this relationship has long been assumed to exist, it has never been formally defined. Although collaboration between researchers and teachers is widely recognized as essential in LA practice research, its implementation has relied more on implicit expectations and ad hoc arrangements than on formal agreements. To address this issue, we conducted a Policy Delphi study consisting of three rounds of structured dialogues with four LA researchers and six K-12 (kindergarten through high school) in-service teachers. Based on 10 research topics covering the entire cycle of LA practice research, we analyzed points of agreement and disagreement between the two stakeholder groups. The results revealed that while collaboration was generally recognized as desirable, there were significant initial discrepancies in expectations regarding roles, authority, and workload. While the Policy Delphi process contributed to strong consensus on four topics, it exacerbated conflicts on others, particularly those related to methodological authority and the right to withdraw from the research. Based on these findings, we derived three design principles for LA practice research policy: the clear articulation of conditions for collaboration, the differentiation of structured roles, and the institutionalization of structurally contentious issues. These principles offer a practical foundation for formalizing researcher-teacher relationships in LA practice research and highlight the need for concrete policy instruments rather than reliance on goodwill alone.

11:05
Validating a Five-Stage Implementation Framework for Competence-Oriented Interdisciplinary Thematic Instruction in High School Information Technology

ABSTRACT. To address the marginalization of the Information Technology (IT) discipline and superficial knowledge integration in current high school interdisciplinary instruction, this study constructs and empirically validates a competence-oriented, five-stage implementation framework encompassing theme selection, conceptual linkage and problem chain organization, competence-based goals, 5E inquiry processes, and comprehensive assessment. Through a one-semester, three-cycle action research involving 52 tenth-grade students exploring the theme "My 'Superpower' Persona," mixed-methods data demonstrated that the framework significantly enhanced students' computational thinking, information awareness, digital learning and innovation, and interdisciplinary competencies. Furthermore, classroom engagement and artifact quality improved substantially. Although the growth in information social responsibility was non-significant—indicating the need for longer-term cultivation—the proposed framework effectively safeguards the core educational value of the IT discipline. It successfully overcomes the "knowledge patchwork" dilemma, providing a replicable and systematic paradigm for high school IT interdisciplinary instruction.

10:20-11:20 Session 32D: C3 Session J
Location: Savoy 4
10:20
Towards Credential Evolution to Competence Authentication: A Taxonomy-Based Analysis of Japanese University Micro-Credentials

ABSTRACT. This study presents the first large-scale, taxonomy-based analysis of micro-credential (MC) metadata characteristics across a national higher education ecosystem, analyzing 233 digital badges stratified-sampled from a national registry of 2,858 publicly available MCs across nine institutional segments. Using the IACET Badging Taxonomy as an international benchmark, two independent coders evaluated each MC against three criteria: Granularity, Self-containment, and Alignment. The results revealed three features of the Japanese MC ecosystem. First, MCs skewed heavily toward IACET Levels 3–4 (course completion), with only 13.7% (n = 32) reaching Level 5-6, suggesting that MCs function as digitized certificates rather than as evidence of competence. Second, a significant characteristic imbalance was observed (Friedman χ²(2) = 154.25, p < .001) on MCs. Third—and most critically—we uncover a "standardization paradox": Data Science MCs exhibited the highest alignment (Dunn's test, p < .001) yet a lower granularity than Specialized MCs, as voluntary benchmark alignment replaced granular learning outcomes with generic institutional boilerplate. We argue that MCs must evolve toward learner-centric, competence-driven design, and we offer concrete metadata-level guidelines for instructional designers and digital credentialing platforms within the ICCE research community.

10:35
Prompting Independence: Preliminary Findings on Large Language Model–Assisted Portfolio of Evidence Development among Alternative Learning System Adult Learners

ABSTRACT. This study examines the role of large language model (LLM)–powered tools in supporting adult learners in the Philippine Alternative Learning System (ALS) during the development of their Portfolio of Evidence (PoE). Grounded in Self-Determination Theory (SDT), the study investigates how LLM-assisted writing environments may influence learners’ autonomy, technical competence, relatedness, confidence, and motivation in non-formal education settings. A sequential explanatory mixed-methods design is employed, consisting of quantitative and qualitative phases. The quantitative phase utilizes a structured survey questionnaire analyzed using Partial Least Squares Structural Equation Modeling (PLS-SEM), while the qualitative phase uses semi-structured interviews to explore learners’ experiences with LLM-assisted PoE writing. Prior to the main implementation, pilot testing was conducted among 30 ALS adult learners with prior experience using LLM-powered tools such as ChatGPT, Gemini, Microsoft Copilot, Claude, and Grammarly. Preliminary pilot testing results demonstrated acceptable reliability and validity. Cronbach’s Alpha and Composite Reliability values exceeded the recommended threshold of 0.70, while Average Variance Extracted (AVE) values surpassed the acceptable level of 0.50. Discriminant validity was established using the Heterotrait-Monotrait Ratio (HTMT) and Fornell-Larcker Criterion. Low-performing indicators were removed to refine the survey instrument for full-scale implementation. The study contributes to emerging discussions on AI-assisted learning in non-formal education and provides preliminary evidence supporting the use of SDT as a framework for examining LLM-assisted writing among ALS adult learners.

10:50
Knowing Isn’t What You Think: Designing a Dashboard for Supporting Knowledge Monitoring
PRESENTER: Li Chen

ABSTRACT. Knowledge monitoring (KM), the ability to judge what one knows and does not know, is a core metacognitive process in self-regulated learning. Although the Knowledge Monitoring Assessment (KMA) framework provides a theoretically grounded approach for assessing KM, limited work has examined how to translate KMA-based information into learner-facing tools for metacognitive support. This study presents the design and formative evaluation of MetaCompass, an AI-based dashboard grounded in KMA framework. MetaCompass is designed through three core principles: dual information visualization of performance and monitoring accuracy, interpretable learner profiles, and profile-aligned adaptive feedback. A formative evaluation was conducted with undergraduate students and instructors to examine the validity of the feedback design, usability, and learning usefulness of MetaCompass. The results indicated that while MetaCompass was perceived as usable, perceived learning usefulness varied among students. The correlation analysis showed that usability was positively correlated with perceived learning usefulness only among students with high metacognitive awareness. These findings indicated that usability alone is insufficient to ensure learning value and highlight the need for metacognitive scaffolding to help learners interpret dashboard data, providing design-based insights into the conditions under which learner-facing dashboards can support KM.

10:20-11:20 Session 32E: C6 Session F
Location: Windsor
10:20
Artificial Intelligence and Child L2 Speaking: A Systematic Review and Meta-Analysis

ABSTRACT. This meta-analysis examined whether artificial intelligence (AI) supports second language (L2) speaking in children aged 3–12 years. Many AI tools, often using machine learning based natural language processing, can simulate conversational partners, deliver immediate feedback, and adapt tasks to learners, which may benefit child L2 pronunciation, fluency, and oral production. We searched nine databases and conducted citation tracing. After double-blind screening of 1,339 records, 11 peer-reviewed studies met the inclusion criteria. Random-effects meta-analysis with robust variance estimation showed a significant positive overall effect (effect size = 1.16, standard error = 0.36, t(18) = 3.20, p = .005), indicating that AI–based interventions improved child L2 speaking, with particularly significant effects on productive vocabulary and content. Moderator analyses suggested that simpler feedback and conversational formats were associated with larger gains. Overall, the findings support the potential of AI to address practical gaps in early L2 speaking instruction, including providing individualized and timely feedback that may be difficult to achieve in typical classrooms or in settings with limited opportunities for input and output. At the same time, conclusions are tempered by the small evidence base and variability in intervention designs and outcome measures, underscoring the need for more rigorous and well-powered empirical studies.

10:35
Unveiling the Dialogic Loops: An Exploratory Lag Sequential Analysis of Learner-AI Interaction in GenAI-Mediated English Debate Training

ABSTRACT. The increasing integration of Generative Artificial Intelligence (GenAI) in language education presents unprecedented opportunities for personalized feedback. However, existing research has primarily evaluated the static quality of AI-generated feedback, leaving the underlying, dynamic interactional processes underexamined. This study employs Lag Sequential Analysis (LSA) to investigate the behavioral patterns within a five-week corpus of learner-AI interactions (80 turns) collected during an undergraduate English debate training program. Interaction logs between an intermediate-level English as Foreign Language (EFL) learner and a large language model debate coach were coded according to learner behaviors (initiation and advancement) and AI feedback functions (evaluation, diagnosis, scaffolding, and elicitation). LSA revealed seven statistically significant behavioral transitions that form a continuous dialogic cycle. These pathways highlight a three-stage process: learner-driven triggering, AI pedagogical processing, and a learner response loop. The findings indicate that learner-AI interaction is organized around recurring dialogic feedback cycles, in which evaluation, diagnosis, scaffolding, and elicitation collectively support continued interaction. Crucially, targeted AI elicitation served as a necessary catalyst for learner advancement, effectively pushing the learner toward higher-order metacognitive negotiation. These findings suggest that GenAI functions not merely as a feedback provider, but as a dynamic co-regulatory partner that facilitates reflective and iterative learning processes.

10:50
Understanding the Roles of Critical Thinking and Learner-AI Interaction in EFL Reading Comprehension

ABSTRACT. This study investigates how learner-AI interaction patterns differ according to learners’ critical thinking skills in generative AI-assisted EFL reading and how these interactions relate to summary writing performance. An experiment was conducted with 40 adult EFL learners. Participants first completed a critical thinking assessment and received prompt guidelines. They then interacted with ChatGPT while reading and comprehending an English passage, followed by a summary writing task completed without AI assistance. Learner-AI interaction patterns were analyzed using Epistemic Network Analysis (ENA), and the effects of interaction on task performance were examined through moderated regression analysis. The results revealed that learners with higher critical thinking skills generated significantly more Critique and Deep Reasoning prompts than those with lower critical thinking skills. ENA further showed that the interaction patterns of the high critical thinking group were characterized by stronger connections involving Deep Reasoning, whereas the low critical thinking groups’ interactions were centered on Comprehension-related prompts. In addition, Feedback Request emerged as a significant predictor of summary writing performance regardless of learners’ critical thinking levels. These findings suggest that critical thinking is reflected in the ways learners interact with generative AI during reading activities. Furthermore, interaction behaviors that encourage monitoring and reflection, such as Feedback Request, may contribute to improved reading comprehension outcomes. The study highlights the importance of incorporating reflecting interaction strategies into the design of AI-assisted foreign language learning environments.

11:05
Tamil Language Learning Platform through NLP Techniques and Gamification

ABSTRACT. Natural Language Processing (NLP) techniques are rapidly integrating into daily life in the form of mobile and web applications, mainly for educational purposes. Educational sectors are making use of this to improve language learning methods. However, there is a lack of such applications for Tamil language and efforts in that domain remain low, despite a huge Tamil population globally. This paper explores the implementation of different NLP techniques based on various Tamil datasets and their integration into a gamified Tamil language learning platform, targeted at Primary School students to build their foundation in a meaningful and engaging manner. Our results, obtained through various models and methodologies, highlight the high potential of this platform along with key limitations of current Tamil datasets. This study provides practical insights on the creation of a gamified Tamil language learning platform using NLP techniques, contributing toward more interactive and AI-assisted Tamil educational applications for younger learners.

10:20-11:20 Session 32F: C4 Session B
Location: Clarendon
10:20
Are Online Learning Self-Efficacy Scale Reliability Measured? A Reliability Generalisation Meta-Analysis with a Small-Sample Caution

ABSTRACT. Measuring online learning self-efficacy is important for understanding learner success in today’s ubiquitous digital environments. A wide variety of measurements have been developed over the decades, yet the reliability of these assessments is often assumed rather than empirically verified, and their use remains scattered. Adhering to the PRISMA statement (Moher et al., 2009; Page et al., 2021) and the REGEMA checklist (Sánchez-Meca et al., 2021), this reliability generalisation (RG) meta-analysis synthesised Cronbach’s alpha coefficients from studies using three prominent online learning self-efficacy measures: Miltiadou and Yu (2000), Tsai et al. (2020), and Zimmerman and Kulikowich (2016). We screened 20,318 records from Scopus and Web of Science (2000-2025), yielding only 15 eligible articles reporting usable reliability data (k = 18, N = 4,581). The pooled overall reliability was excellent (α = .9395, 95% CI [.9156, .9566]). However, extreme heterogeneity was also observed (I² = 97.54%). Scale-specific estimates ranged from α = .8395 (Tsai et al.) to α = .9775 (Zimmerman & Kulikowich), with substantial between-study variance (τ² up to 0.8283). Moderator analyses revealed that reliability was not an invariant property. In particular, gender composition, geographical region, Likert format, scale modifications, and reporting of factor analyses significantly influenced estimates. Importantly, publication bias tests suggested no systematic suppression of low reliability studies. However, as an important finding, this meta-analysis is constrained by a very small number of primary studies (k = 18 across all scales), limiting statistical power for moderator analyses and precluding meta-regression with multiple simultaneous covariates. Most eligible studies failed to report McDonald’s omega or test-retest reliability, and few employed probability sampling. Hence, we suggest researchers should (a) routinely report score reliability for their own sample rather than citing past values, (b) conduct and report confirmatory factor analysis, (c) include diverse samples beyond convenience-based university students, (d) choose suitable and validated measures for particular samples with full explanations provided.

10:35
How Technology-Supported Outdoor Learning Promotes Self-Directed Learning in Citizenship Education

ABSTRACT. Compared with conventional classroom settings, outdoor learning environments are typically more open and less structured, requiring students to engage more actively in planning, monitoring, and regulating their own learning. Despite growing evidence that digital outdoor learning can promote self-directed learning (SDL), the mechanisms underlying this effect remain underexplored. This study investigated how technology-supported outdoor learning shapes high-school students’ SDL in the Citizenship and Social Development (CSD) curriculum by examining the relationships among cognitive load, motivation, learning engagement, self-efficacy, and SDL. Data were collected through an SDL questionnaire-based instrument from 396 Hong Kong senior secondary students after a two-hour CSD-related outdoor learning activity supported by an outdoor learning platform, EduVenture. The results showed that motivation, learning engagement, and self-efficacy significantly predicted SDL. Cognitive load did not directly affect SDL; however, it indirectly influenced SDL through motivation, learning engagement, and self-efficacy. Furthermore, cognitive load significantly moderated the relationship between motivation and learning engagement, suggesting that excessive cognitive load may weaken the positive effect of motivation on learning engagement. These findings suggest that effective outdoor learning design should balance motivational support with manageable cognitive demands. Well-designed scaffolding strategies, such as staged task release, context-aware prompts, and embedded reflection cues, may help maintain optimal cognitive load while strengthening students’ learning engagement, self-efficacy, and SDL capability. The study contributes to a deeper theoretical understanding of SDL in technology-supported outdoor learning and offers practical implications for the design of digitally mediated experiential learning in citizenship education.

10:50
Measuring Emotional Engagement in XR-Based Intercultural Virtual Exchange: A Pilot Study Using Physiological and Self-Report Indicators

ABSTRACT. This pilot study explores how physiological and self-report measures can be combined to assess emotional engagement in immersive learning. Seven university students participated in a within-subject experiment comparing head-mounted display (HMD)-based learning materials with slide-based materials in the context of a multinational Virtual Exchange (VE). Heart rate and skin conductance were measured, along with self-reported emotional states using the General Affect Scale. The analysis revealed significant interactions between the learning format and the measurement period for both skin conductance and for positive and negative affect. Specifically, the HMD-based materials were shown to elicit higher emotional arousal during the task compared to the slide-based materials. These preliminary findings suggest the utility of an integrated approach, combining physiological signals with traditional measures, to deepen the understanding of emotional engagement in XR-based education.

11:05
AllerColler: An LLM-Empowered Reflection Support System Bridging Classroom and Field-Based Learning

ABSTRACT. Field-based learning, such as fieldwork, plays a vital role in bridging the gap between theory and practice. Specifically, reflection is essential for enabling learners to restructure their understanding of prior classroom learning. However, learners often struggle to relate their practical experiences to theoretical knowledge. Despite advancements in technology-enhanced learning approaches, few studies have integrated both field-based and classroom learning contexts to address this difficulty. To fill this gap, this study proposes a context-aware learning environment called AllerColler. It is designed to provide personalized feedback that facilitates sense-making across both learning settings during reflection. The proposed system leverages a large language model (LLM) to estimate relevant classroom learning content based on learner inputs. To ensure the feasibility of the proposed system in practical educational settings, semi-structured interviews were conducted with five experts in field-based learning in Japan. A thematic analysis revealed that the proposed system could navigate field-based learning where the learning objectives are predetermined or guided by educators. Additionally, the content of personalized feedback could facilitate reflection, although a potential risk of overreliance on intelligent support was identified. These findings clarify future directions for this approach to promote the thoughtful mobility of learning across classrooms and real-world settings.

11:20-12:30 Session 33A: C1 Session L
Location: Savoy West
11:20
When Plausible Becomes Believed: Possibility-to-Actuality Hallucination and Distributed Confabulation as Challenges for AI Literacy in Learning Ecologies

ABSTRACT. As generative AI systems become embedded in learning ecologies, learners are increasingly exposed to AI-generated content whose epistemic status, whether it represents verified fact, plausible inference, or fabrication, is difficult to discern. This paper examines the educational implications of two recently characterized AI failure modes: possibility-to-actuality hallucination (P2A hallucination), wherein AI systems treat mathematically possible but evidentially unsupported claims as factual, and distributed confabulation, wherein human–AI interaction creates a feedback loop that materializes and reinforces unsupported inferences. We argue that these phenomena pose a distinctive challenge for AI literacy education that current curricula are not equipped to address, because they focus on teaching learners that AI can hallucinate but not how to recognize the subtle, plausible hallucinations that are most resistant to detection. Drawing on constructivist learning theory, epistemic cognition research, and AI literacy frameworks, we propose a Pedagogical Framework for Evidential Reasoning with AI (PF-ERA) that integrates modal reasoning competencies into existing literacy curricula. We present five classroom-ready pedagogical scenarios, map these onto Bloom’s revised taxonomy, and analyze how P2A hallucination propagates through learning ecologies when AI-generated content circulates among learners, teachers, and institutional knowledge systems. Our framework addresses a critical gap in reimagining learning ecologies for the age of intelligent technologies: developing learners’ capacity to evaluate not just whether AI output is true, but how well-supported it is by available evidence.

11:35
Comparing One-Layer and Two-Layer AI Feedback with Human Feedback: The Moderating Role of Feedback Literacy

ABSTRACT. Artificial Intelligence (AI) Technologies (e.g., large language models) make it possible to deliver immediate, adaptive feedback at scale, offering a potential substitute for the costly authoring effort required for human-designed feedback. Yet it remains unclear whether different structures of AI feedback yield learning gains comparable to human feedback, and how these effects vary across learners who differ in prior knowledge and feedback literacy. We conducted a randomized study with 302 college students assigned to one of three feedback conditions: (a) human feedback (business-as-usual), (b) AI one-layer feedback (provide correct answer paired with elaboration and suggestions in all attempts; non-layered), or (c) AI two-layer feedback (provide encouragement, hints, and reflective prompts in the first attempt, correct answer in a later attempt; layered). Learning was assessed through pre- and post-tests, and participants additionally reported their feedback literacy (the understandings, capacities, and dispositions needed to make sense of and act on feedback), and perceptions of the benefits of the feedback they received. Results show that, first, both AI feedback conditions produced post-test gains comparable to human feedback overall, supporting the viability of AI feedback as a scalable alternative. Second, feedback literacy (FL) but not prior knowledge moderated the effect: low-FL learners benefited more from AI feedback (both one-layer and two-layer) than from human feedback, while high-FL learners performed comparably across conditions. Low-FL students were more likely to perceive AI feedback as promoting self-regulated learning benefits (i.e., positive affect and motivation as well as self-reflection and improvement). These findings suggest that AI feedback of different structures can match the learning value of human feedback and may be particularly valuable for learners who enter the task with lower feedback literacy, which also informs further AI system design to adapt to learner characteristics.

11:50
Generative AI as a Cognitive Scaffold in Project-Based Programming Learning: Effects on Higher-Order Thinking and Learning Performance

ABSTRACT. Project-based learning (PBL) is widely recognized for fostering higher-order thinking in programming education, yet its effectiveness is often constrained by insufficient scaffolding and feedback. With the emergence of generative artificial intelligence (GenAI), new opportunities arise to support learners during complex project tasks. This study proposes a Generative AI-supported Project-Based Programming Learning (Gen-PPBL) model and investigates its effectiveness through a quasi-experimental study involving 81 high school students. The experimental group received GenAI-supported instruction, while the control group followed conventional PBL. The results showed that the experimental group significantly outperformed the control group in decision-making, creative thinking, critical thinking, and problem-solving abilities. Students also demonstrated improvements in project innovation, learning engagement, and programming confidence. The findings suggest that GenAI functions as an effective cognitive scaffold, enhancing student engagement and supporting complex problem-solving in programming education.

12:05
Investigating Oddness Annotation Ambiguity in LLM-Generated Stories for Elementary School Kanji Learning

ABSTRACT. With the rapid advancement of Large Language Models (LLMs), their integration into educational contexts has accelerated, including tools that generate personalized stories to support kanji (Chinese characters) learning for elementary school students. However, the educational appropriateness and linguistic naturalness of these AI-generated narratives require rigorous validation, as conventional automated metrics often fail to capture subtle textual nuances. This paper investigates the feasibility of utilizing LLMs to automatically detect and classify specific types of textual "oddness" in LLM-generated stories, with two objectives: (1) evaluating how well LLM-based classifiers reproduce human-defined oddness judgments. (2) identifying ambiguities in annotation criteria through analysis of LLM-generated rationales. We defined a typology of seven oddness categories organized into three hierarchical levels—local, contextual, and educational oddness—along with a "Natural" category. A dataset of 480 LLM-generated stories was annotated by human raters. and inter-annotator agreement was analyzed using Cohen's kappa. Moderate agreement was obtained for the Natural category, whereas agreement for individual labels was at or below chance level. indicating that the boundaries between labels remain ambiguous. LLM-based classification was evaluated using a Multi-Label Confusion Matrix (MLCM); the highest F1 score in the eight-label setting was 0.237, whereas binary detection of oddness presence achieved an F1 score of 0.69. These findings suggest that annotation ambiguity, rather than model capability alone, constitutes the primary barrier to reliable automatic evaluation of AI-generated educational narratives. This study provides foundational insights for developing benchmark tools to evaluate the quality of AI-generated educational content.

11:20-12:30 Session 33B: C1 Session M
Location: Savoy 2
11:20
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics

ABSTRACT. Large language models (LLMs) are increasingly entering students’ learning practices, but their educational value depends on whether they are used to support reasoning or to complete tasks without engaging in the underlying reasoning. This study examines guided LLM use in an undergraduate Probability and Statistics course, focusing on the distinction between assigned LLM access and the quality of students’ actual interaction with the model. In a four-week quasi-experimental summer program, students were organized into three balanced conditions: no LLM access, unrestricted LLM access, and guided LLM access. The guided condition used the same LLM platform as the unrestricted condition, but students received explicit training and rules intended to promote reasoning-focused help-seeking, stepwise hints, verification, and ethical use. All quizzes and the delayed final exam were completed without LLM or external assistance, allowing us to separate AI-supported practice performance from independent learning. Results show that guided use was associated with a clearer learning-oriented interaction pattern than unrestricted access, especially in prioritizing reasoning over final answers and requesting stepwise support. Guided-LLM students showed a promising pattern of stronger no-help quiz performance in the intervention phase, while unrestricted access appeared more useful for assisted practice completion than for consistently improving independent performance. Available time measures did not support a simple duration-based explanation, and self-assessment calibration suggested better alignment between perceived and demonstrated understanding in Guided-LLM. Overall, the findings suggest that LLM access alone is an incomplete educational intervention. For Artificial Intelligence in Education (AIED), the central design challenge is to scaffold how students use LLMs so that these systems function as partners in reasoning rather than answer-getting tools.

11:45
Adversarial Thinking Training System with Situation Recommendation to Overcome Limited Perspectives

ABSTRACT. Adversarial thinking is to adopt an intruder's perspective to plan a break-in. In crime prevention, this approach helps residents generate security measures from diverse viewpoints without being constrained by limited perspectives. However, most existing training systems for adversarial thinking only provide environments to encourage thinking like an intruder. They do not support the derivation of concrete security measures. The objective of this study is to develop an adversarial thinking training system which identifies a learner's limited perspectives based on their proposed security measures and recommends specific situations to help them overcome these perspectives. The system uses a 3D house model as a target to set security measures. First, the learner sets security measures from a resident's perspective. Next, the learner plans a break-in from an intruder's perspective. In this step, the system analyzes the learner's initial security measures to identify limited perspectives and recommends tools for breaking into the house to prompt perspectives the learner missed. Last, the learner generates new security measures from different viewpoints based on the break-in planned with these tools.

12:00
What Do Reviews Tell Us About Generative AI and Emotional Engagement in Education? An Umbrella Review

ABSTRACT. Generative artificial intelligence (GenAI) is increasingly shaping learners’ cognitive, behavioral, and emotional engagement in education. This umbrella review examined what literature reviews have been conducted and what results have been reported on GenAI and emotional engagement. Fourteen eligible review articles published between 2023 and 2026 were identified and analyzed. Findings show that relevant evidence is distributed across three strands: reviews on multidimensional learner engagement, reviews on adjacent emotional constructs, and reviews on emotion-sensitive AI capabilities. Emotional engagement was most directly discussed as one dimension of learner engagement, but its indicators varied across contexts, including interest, enjoyment, confidence, value, relevance, reduced anxiety. Existing reviews suggest that GenAI may support emotional engagement-related outcomes through feedback, personalization, conversational interaction, and task support. However, gaps remain in construct clarity, mechanism synthesis, and measurement.

11:20-12:30 Session 33C: C7 Session H
Location: Savoy 3
11:20
Effects of a Generative-AI-Based Instructional Material Development Course on Pre-service Teachers' TPACK and AI Teaching Efficacy

ABSTRACT. As generative artificial intelligence (AI) becomes embedded in everyday teaching, pre-service teachers must be prepared not only to use such tools but to integrate them pedagogically. This study examined the effects of a course in which pre-service teachers developed their own instructional materials using generative AI (a "vibe coding" approach) on their Technological Pedagogical Content Knowledge (TPACK) and AI teaching efficacy. A single-group pre-post design was used; of the participants, 20 who completed both the pre- and post-tests were matched and analysed using paired-samples t-tests, Wilcoxon signed-rank tests, and Cohen's d. AI teaching efficacy improved significantly overall (t(19) = −8.83, p < .001, d = 1.97), with the largest gains in AI instructional-design competence (PATE, d = 1.83). TPACK also rose significantly across all sub-domains (overall d = 1.74), with the largest effects in the technology-integrated domains (TPACK, d = 1.57; TPK, d = 1.26; TCK, d = 0.87). General attitudes toward AI (ATSE) and interaction with AI (IWAI) also improved significantly, but with notably smaller effects (both d = 0.55). Findings suggest that authentic AI-based material-development tasks strongly strengthen pre-service teachers' capacity to integrate AI into teaching, whereas broader dispositions toward AI shift to a more modest degree within a single-semester intervention.

11:45
Who Can Afford Responsible AI in Education? Public Investment, Institutional Capacity, and Inequality in AI-Enabled Learning Systems

ABSTRACT. Artificial intelligence (AI) in schooling is often framed as a technical or ethical problem, but responsible implementation also depends on public investment, teacher capacity, and Information and Communication Technology (ICT) policy. This study examines whether education systems with higher levels of public education spending are better positioned to support equitable access to AI-enabled learning. Because cross-national data on AI-specific education spending are not yet available, the study uses digital readiness and access to learning technologies as proxies for the infrastructure needed for responsible AI deployment. The analysis draws on student-level data from PISA 2022, covering 37 OECD countries and 295,157 students (OECD, 2023). It also uses country-level finance and macroeconomic indicators from the World Bank (2024) and the Government AI Readiness Index from Oxford Insights (2022). The study estimates country-level ordinary least squares (OLS) models to examine the relationship between public education spending, digital readiness, and socioeconomic gaps in access. It also uses weighted student-level OLS models with country fixed effects to examine whether returns to digital access differ by SES. Results show that public spending is positively associated with national digital readiness, income inequality is the strongest predictor of within-country access gaps, and low-SES students gain smaller learning returns from digital access. The findings position responsible AI as a practice-driven and policy challenge for ICT in education.

12:00
Reimagining AI Learning Ecologies: Government School Children’s Perceptions and Explanations of Artificial Intelligence in India

ABSTRACT. As AI education expands across the Global South, a critical question remains underexplored: what do children already believe about AI before formal instruction begins, and how do those beliefs shape their readiness to learn? This paper examines the perceptions and emerging explanatory models of AI held by 120 government school students in Grades 5–8 across four municipal and Zilla Parishad schools in Pune, India. Using a purpose-designed Marathi-language survey combining quantitative items, scenario-based reasoning tasks, and drawing prompts, we identify three recurring sense-making orientations: device-bound, anthropomorphic or cognitive, and process-oriented. We show that device-bound perceptions generate a hardware misconception - the belief that AI cannot be learned without devices - functioning as a conceptual access barrier in low-resource school contexts. Crucially, nearly half of hardware-bound students simultaneously demonstrated latent machine-learning intuitions, suggesting conceptual capacity is present but pedagogically unreached. An AI recognition task further reveals a YouTube paradox: the platform through which 73.5% of students first learned about AI is the one fewest can recognise as AI. We argue that equitable AI education must begin by mapping sense-making orientations learners already bring, requiring neither devices nor internet.

12:15
Youth Questioning, Trust, and AI Boundaries: A Co-Design Study Informing the Design of AI Guidance Systems in Vocational Education

ABSTRACT. AI guidance systems for vocational learners are being deployed at scale across India and comparable contexts, yet their design is driven by what the technology can do rather than what learners need, feel, and will accept. This paper addresses that gap through a co-design study with 28 young people enrolled in Industrial and Vocational Training Institutes across eight Indian states. Participants took part in a six-hour structured workshop comprising three activities: mapping the questions they carry and the emotions those questions produce, tracing their journeys across information sources to understand how trust forms and breaks, and co-designing the behavioural boundaries they consider non-negotiable for an AI guidance system. The study was conducted to inform the design of AskAbhi, an AI chatbot developed by Quest Alliance for the MyQuest learning application. Three findings emerge. (1) Learners bring emotionally charged, high-stakes questions, not information requests. Six thematic domains emerged; career and employment questions accounted for 28% of all questions, and AI and technological uncertainty for 17%. Confusion was the most frequent emotional response across all domains. Fear characterised 50% of AI-related questions, and 67% of decision ambiguity questions were associated with confusion. (2) Learners do not trust AI as a primary authority. Help-seeking followed sequential, multi-source paths: AI was used alongside teachers, family, peers, and search engines, and trust formed only when multiple sources aligned. Breakdown moments arose when learners needed help choosing between options they already had, not when they lacked information. (3) When learners defined what they wanted from AI, they led with what it must never do. Across 183 co-designed boundary statements, the dominant categories were answer style, accuracy, and privacy, each framed as prohibitions. Structured clarity, accuracy, and privacy emerged as preconditions for trust, not desirable features. Six design principles derived from these findings offer a grounded framework for AI guidance systems in high-pressure, low-resource vocational education contexts.

11:20-12:30 Session 33D: C3 Session K
Location: Savoy 4
11:20
Letting Learners Shape What to Reflect On: Design Principles for an LLM-Based Skill Analytics System

ABSTRACT. Conventional Learning Analytics (LA) has delegated the design of assessment criteria to institutions and researchers and has relied on Learning Management System (LMS) logs, thereby relegating learners to passive consumers of analytical results and excluding practical experience data—such as internships and project-based activities—from analysis. This study proposes an LA system that leverages Large Language Models (LLMs) to enable learners to design assessment criteria, jointly analyze formal learning data and practical experience data, and derive design principles from an evaluation study (N=12) using a prototype implementation. The findings indicate that (1) institutionally collected course data serve as an initial reference frame for self-understanding, but assessing interpersonal, self-management, and problem-solving skills requires the addition of practical experience data; and (2) generic skill axes and custom skill axes complement one another—the former supporting retrospective reflection of past experiences and the latter supporting future-oriented reflection. This work contributes design principles for designing LLM-supported, learner-centered LA systems in which learners participate in constructing both the data and the assessment axes used for reflection.

11:45
Seductive Details in Video Lectures: Examining the Effects of Placement on Learner Engagement

ABSTRACT. The seductive details effect occurs when learners’ attention is diverted by interesting but irrelevant information, potentially impairing comprehension and retention. While prior research in text-based instruction shows that seductive details placed early hinder learner performance more than when placed later, it is unclear how such placement affects engagement in video lectures. This experimental study investigated how the presence and placement of seductive details relate to learner performance and engagement across four video conditions: seductive details-only, no seductive details condition, and two mixed versions with seductive details appearing in the first or the second half of the video. The results indicated no statistically significant differences in test performance across conditions. However, placement was associated with differences in engagement measures. Specifically, the mixed conditions were associated with reduced task-unrelated thoughts and decreased difficulty disengaging from one’s own thoughts compared to the other conditions. These findings suggest that incorporating seductive details in the first half of a video may help support learner engagement without compromising learning outcomes, offering implications for the design of online instructional materials.

12:00
MetaGuru: Computer-Based Learning Environment to Foster Metacognitive Strategies in Engineering Problem-Solving

ABSTRACT. Computer-based learning environments (CBLEs) place substantial metacognitive demands on learners who must independently plan, monitor, and evaluate their own learning while navigating non-linear, multimedia-rich content. Although metacognitive prompts have been shown to address this challenge in reading and writing contexts, their design and empirical evaluation in structured engineering problem-solving remain largely unexplored. To address this gap, we have designed and developed a learning environment, MetaGuru, embedded five metacognitive prompts: orientation, planning, monitoring, evaluation, and reflection to foster metacognitive strategies during circuit analysis problem-solving. This paper introduces the design and development of MetaGuru. For the development of MetaGuru, we used the ADDIE model of iterative instructional design, which provides a principled framework for translating theoretical design decisions into empirically validated instructional interventions through successive cycles of formative evaluation. Following the development of MetaGuru, a study was conducted with first-year engineering students (n = 45) to evaluate the system's effectiveness in supporting learning in circuit analysis. A pre-test and post-test design were used to assess learners' conceptual knowledge before and after interacting with MetaGuru. Post-test scores were significantly higher than pre-test scores (pre-test M = 9.64, SD = 3.31; post-test M = 11.91, SD = 2.79; W = 31.5, p < 0.001, Cohen's d = 0.93), with a mean normalized learning gain of g = 0.19, suggesting that MetaGuru supports initial conceptual learning in circuit analysis problem-solving.

11:20-12:30 Session 33E: C6 Session G
Location: Windsor
11:20
WPM Dashboard for Time-Aware Reading Practice in EFL Test Preparation

ABSTRACT. This study investigated how dashboard-based reading-speed feedback influences self-monitoring during time-aware reading practice for university entrance examination preparation. Effective time management is an important component of reading performance in examination settings; however, learners often receive limited support for monitoring their reading pace. To address this issue, a Words Per Minute (WPM) dashboard was developed as part of a Time-Aware Reading Practice (TARP) framework designed to support both reading practice and time management. Dashboard logs and reading performance data from Japanese high school English classes were analyzed. The results show that the students actively engaged in self-monitoring, accessing the dashboard a median of nine times across ten reading tasks. Even when the timer was hidden, 34% of students accessed the dashboard, indicating heightened awareness of reading pace. Differences in self-monitoring behavior were observed across learner groups, suggesting that learners engaged with dashboard feedback in different ways depending on their reading proficiency. These findings suggest that temporal feedback combined with social comparison information can encourage reflective reading behaviors and promote time-aware learning strategies in examination-oriented contexts. The study contributes to the field of learning analytics (LA) by demonstrating how an LA-supported TARP framework can facilitate both reading fluency development and metacognitive regulation. The results also highlight the potential of dashboard-based feedback to support learners with different levels of reading proficiency through adaptive self-monitoring opportunities.

11:35
Direct Output vs. Socratic Questioning: How AI Feedback Modes Differentially Activate Self-Regulated Learning Behaviors in University Writing Revision

ABSTRACT. Abstract: Generative AI writing assistants are widely used in universities, yet how different feedback designs shape students' self-regulated learning (SRL) behaviors during revision remains underexplored. This between-subjects experiment compared two AI feedback modes (direct output and Socratic questioning) as twenty Chinese EFL university students revised English argumentative essays. Using turn-level coding grounded in Pintrich's (2004) SRL framework (κ = 0.885) and Epistemic Network Analysis, we analyzed 165 student turns. Results showed that direct output engaged a Control–Reflection cluster while Socratic questioning engaged a Monitoring–Help cluster (Cohen's d = 1.51). Neither mode triggered complete P→M→C→R microcycles; knowledge-rule extraction was absent in both groups. The behavior–outcome relationship reversed across conditions: SRL behavior was unrelated to score gain in the direct-output group (ρ = −0.11) but positively associated in the Socratic group (ρ = +0.62). These results suggest that adaptive mixed-mode scaffolding, adjusting to students' current regulatory state, is needed.

11:50
An EFL Writing Support System Using AI-Generated Image Feedback to Inspire Writing
PRESENTER: Kalai Wong

ABSTRACT. Automated writing evaluation systems commonly provide textual feedback for second language writing, but such feedback often focuses learners’ attention on surface-level errors rather than content development. Drawing on the Noticing Hypothesis, this study examines whether AI-generated images can serve as an alternative feedback modality that redirects attention toward meaning construction. It investigates how image-based feedback influences learners’ interactions with an AI-assisted writing system and whether it supports content development during revision. We developed Avery, an AI-assisted writing support system that provides either textual or image-based feedback on descriptive writing tasks. Sixty-seven Japanese university students were randomly assigned to an image-feedback group or a text-feedback group. Participants completed writing tasks, revised their work after receiving feedback, and resubmitted their texts. Interaction logs and revision records were analyzed using established revision frameworks and Transition Network Analysis. Analysis of 95 writing sessions and 1,630 interaction events showed that the image-feedback group exhibited a denser transition network, indicating more varied interaction patterns. These learners more frequently progressed to subsequent writing tasks, viewed feedback, and checked the leaderboard, while showing fewer behaviors associated with leaving and returning to the system. This suggests higher engagement and sustained participation. Revision analysis revealed similar rates of content-level revisions across groups, but substantially fewer surface-level corrections in the image-feedback condition. While textual feedback encouraged linguistic correction, image-based feedback promoted engagement, task continuation, and attention to meaning. These findings highlight the potential of AI-generated images as an alternative feedback modality for supporting L2 writing development.

12:05
Artificial Intelligence Agents in Language Learning: A Scoping Review of SSCI-Indexed Web of Science Research (2016-2026)

ABSTRACT. Artificial intelligence agents are increasingly shaping technology-enhanced language learning by supporting dialogue, feedback, assessment, scaffolding, and personalized practice. However, the field remains conceptually fragmented because related systems are often described using overlapping terms such as chatbots, conversational agents, pedagogical agents, intelligent tutoring systems, dialogue systems, virtual agents, and generative AI tools. This scoping review maps SSCI-indexed Web of Science research on AI agents in language learning to clarify publication trends, agent types, language-skill coverage, learning outcomes, research designs, and future directions. Guided by PRISMA-ScR and JBI-informed principles, the review synthesizes literature on AI-agent-supported language learning across educational technology, applied linguistics, and computer-assisted language learning. The findings show that research has expanded rapidly with the emergence of generative AI and large language models, shifting the field from earlier dialogue-based and tutoring systems toward more flexible conversational, feedback-oriented, and learner-support agents. AI agents are commonly used as speaking partners, writing assistants, feedback providers, personalized tutors, pronunciation-support systems, and self-regulated learning companions. Despite these developments, the field still needs clearer conceptual definitions, stronger longitudinal evidence, deeper attention to ethical AI literacy, more explicit teacher-AI collaboration models, and broader research in multilingual and mobile-first EFL contexts. The review concludes that AI agents should be understood not only as technological tools but also as pedagogical systems that require careful instructional design, learner training, teacher mediation, and responsible integration.

11:20-12:30 Session 33F: C2 Session D
Location: Clarendon
11:20
Exploring the Relationships Between Students’ Perceptions of Generative AI and Participation in Collaborative Learning

ABSTRACT. This study investigates how university students’ perceptions of generative AI relate to their engagement in collaborative learning and to subsequent changes in those perceptions. The study was conducted in a first-year data science class at a Japanese university, involving 554 first-year students from the faculties of Medicine, Dentistry, Pharmaceutical Sciences, and Law. Students participated in a group-based AI literacy activity using generative AI tools, with discussions conducted through a chat-based system. Participation was measured using log data, including the number of posts and total character counts. The analysis consisted of pre–post comparisons using Wilcoxon signed-rank tests and regression models examining relationships between perceptions and participation. The results showed significant increases in institutional expectations toward AI, perceived choice in learning about AI, and interest in AI learning. Exploratory regression analyses indicated that students’ initial AI ethics-related perceptions were positively associated with posting frequency, whereas anxiety-related factors were not significantly related to participation. Furthermore, a higher number of posts was associated with a smaller increase in AI ethics-related perceptions, suggesting that active engagement in critical discussion may foster more nuanced and reflective understandings of AI. These findings suggest that the relationship between perceptions and participation is selective and dynamic, highlighting the value of structured collaborative activities for promoting critical and ethical reflection. The study contributes to learning analytics research by demonstrating how behavioral log data and perception measures can be integrated to better understand AI-supported collaborative learning.

11:35
From Storyboard to Situated Validation: Metaboard for Repairing Learning Discontinuities in Generative-AI-Supported Project-Based Learning

ABSTRACT. Generative AI (GenAI) can help students rapidly create persuasive storyboards, but these outputs may also allow learners to self-confirm imagined designs without testing their fit with physical, contextual, and client needs. This study proposes Metaboard, a generative-AI-supported cyber-physical learning design that transforms storyboards into situated validation spaces. In Metaboard, students enter LLM-generated professional scenarios, define product-context interfaces, synchronize physical prototypes with digital twins, revise prompts when cyber-physical mismatches occur, and publicly demonstrate their designs as future professionals. We conducted a quasi-experimental study in an artificial intelligence of things (AIOT) and robot project course and implemented the same Metaboard principles in a hospitality design course to examine transferability across domains. Results suggest that Metaboard improved learning motivation, design thinking, situated cognition, psychological ownership, and individual project-based learning (PBL) performance compared with GenAI-supported scenario demonstration. Qualitative evidence further suggests that Metaboard shifted students from self-demonstrable functionality toward client-oriented, interface-connected, and publicly accountable design reasoning. The hospitality course further suggested transferability to service design. This study contributes a design principle for AI-supported PBL: AI-generated scenarios become educationally meaningful when they are made answerable to physical constraints, contextual interfaces, professional identity, and public validation through epistemic friction.

12:00
Acting Upon Peer Feedback Under Different AI-Support Conditions: An Exploratory Sequential Analysis

ABSTRACT. Peer feedback is widely used in educational contexts, but students often find it difficult to interpret feedback and use it to improve their work. The recent use of artificial intelligence (AI) offers new ways to support this process, yet little is known about how different forms of AI support influence students’ behaviours when acting upon peer feedback. Drawing on the feedback literacy framework, this exploratory study examined students’ sequential patterns of acting upon peer feedback and the timing of judgement making under three conditions: No-AI, AI-assisted, and AI-directed. In the AI-assisted condition, a chatbot provided clarification and prompts without directly rewriting text. In the AI-directed condition, a chatbot generated revised text that students could adopt, edit, or reject. Thirty-nine graduate students were randomly assigned to the three conditions in an argumentative writing revision task. Screen recordings yielded 223 behavioural episodes from 34 valid participants, which were coded into three dimensions: attending to feedback, making judgements, and taking action. Sequential pattern analysis suggested different behavioural patterns across conditions. Students in the No-AI condition mainly engaged in repeated self-authored rewriting after attending to peer feedback. Students in the AI-assisted condition showed more varied behaviours, often comparing peer feedback, AI responses, and their own drafts before integrating or rewriting suggestions. Students in the AI-directed condition frequently adopt AI-generated revisions, followed by later evaluation, editing, or rejection. The distribution of specific feedback literacy codes differed significantly across conditions, χ² = 43.79, p < .001, and descriptively judgement making occurred earlier in the AI-supported conditions. These findings add process-level evidence to feedback literacy research by showing how students work with peer comments, AI-generated input, and their own task goals when acting upon peer feedback. They also suggest the need to design AI support that guides students to interpret, compare, and justify their responses to feedback.

12:15
Characterising Persistence in Collaborative Electronic Making

ABSTRACT. Makerspaces are failure-rich environments where persistence is central to how learners continue engaging with open-ended projects. Research on persistence in maker education documents that students spend substantial time on projects, return across sessions, and describe their experiences as ones of overcoming difficulty. However, existing approaches treat challenges or obstacles learners face as equivalent across instances. In collaborative electronic making, this assumption does not hold. Challenges vary fundamentally and make qualitatively different persistence demands. To address this gap, we conducted a three-week workshop on collaborative electronic making (N=20). A total of 152 trouble episodes were identified and analysed from approximately 51 hours of video documentation using interaction analysis. Episodes were coded inductively for trouble type, and persistence was characterised through attempt count and episode duration. Inductive coding produced six trouble types and revealed that digital and hardware troubles dominated the trouble landscape. The number of attempts per episode established a distinction between bounded troubles and open troubles. Plotting episodes against attempt count and mean duration per attempt revealed qualitatively different persistence profiles. A wheel-spinning-like profile appeared almost exclusively in open troubles and was associated with a 60% unresolved rate. A deliberate, sustained engagement (productive persistence) profile was associated with a substantially lower unresolved rate (33.3%). These findings suggest that trouble type is a necessary context for interpreting persistence metrics in open-ended electronic making. The trouble typology developed here offers researchers a systematic basis for studying persistence and provides a conceptual framework to inform facilitator intervention decisions.

12:30-13:30Lunch Break
13:30-14:30 Session 35B: C1 Session N
Location: Savoy 2
13:30
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring

ABSTRACT. Recently, as Large Language Models (LLMs) have fundamentally impacted various fields, the methods for incorporating up-to-date information into LLMs or adding external knowledge to construct domain-specific models have garnered wide attention. Retrieval-Augmented Generation (RAG), serving as an inference-time scaling method, is notable for its low cost and minimal effort for parameter tuning. However, due to heterogeneous training data and model architecture, the variant embedding models used in RAG exhibit different benefits across various areas, often leading to different similarity calculation results and, consequently, varying response quality from LLMs. To address this problem, we propose and examine two novel approaches that combine the benefits of multiple embedding models, named Mixture-Embedding RAG and Confident RAG. Mixture-Embedding RAG simply sorts and selects retrievals from multiple embedding models based on standardized similarity; however, it does not outperform vanilla RAG. In contrast, Confident RAG generates responses multiple times using different embedding models and then selects the responses with the highest confidence level, demonstrating average improvements of approximately 10% and 5% over vanilla LLMs and RAG, respectively. The consistent results across different LLMs and embedding models indicate that Confident RAG is an efficient plug-and-play solution for mathematics question answering. The code of our paper is provided at https://github.com/RS2002/Confident-RAG.

13:55
Agentic Learner–LLM Interaction Framework: Toward a Human-Centered Perspective on LLM-Supported Learning

ABSTRACT. Abstract. Large language models (LLMs) are increasingly used in education to support writing, problem solving, feedback, and personalized learning. Yet much of the current discussion frames their value in terms of efficiency, responsiveness, and task performance. This paper argues that such a framing is insufficient for educational analysis because it overlooks how learner-LLM interaction restructures who is doing the thinking, judging, and monitoring during learning. To address this gap, the paper proposes a human-centered framework for analyzing and designing learner-LLM interaction through three dimensions: cognitive responsibility distribution, metacognitive engagement demand, and productive resistance. The framework takes subjectification as its normative orientation and treats learner agency as a central educational concern rather than a byproduct of system effectiveness. It is then applied to contrasting learner-facing systems to show that the key difference is not simply how much help AI provides, but how that help is organized in interaction. By shifting attention from assistance alone to the conditions under which learners remain active participants in meaning making, the paper offers a conceptual contribution for future research and a design-oriented lens for more human-centered uses of LLMs in education.

13:30-14:30 Session 35C: C3 Session L
Location: Savoy 4
13:30
Prompt-Controlled Generative AI as Visual Scaffolding for Sustainable Design Education

ABSTRACT. Generative artificial intelligence (GenAI) has increasingly entered design education, yet its educational value should not be understood merely as the production of visually attractive design images. In sustainable design education, novice learners often need to connect abstract environmental concepts, such as passive ventilation, daylighting, material reduction, thermal buffering, and biophilic integration, with observable spatial and material cues. This short paper examines how prompt-controlled AI-generated images can function as teacher-mediated visual scaffolds for sustainable design learning. Rather than evaluating image generation only as a technical representation problem, the study focuses on how learners interpret unannotated visual materials and how pedagogical cue clarity may support conceptual understanding. The proposed framework includes feature classification, visual cue operationalization, prompt formulation, image screening, and learner evaluation. Fifty interior-level low-carbon design features were organized into five categories and generated under controlled scene and rendering conditions using Stable Diffusion. Eighty-four design-background participants evaluated six randomly assigned images without textual labels. The preliminary results show that objective recognition accuracy was moderate, whereas perceived instructional usefulness was relatively positive. Visual realism alone was not significantly associated with objective recognition, suggesting that pedagogical cue clarity may be more important than photorealistic quality and that cognitive load should be considered when unannotated AI-generated images are used with novice learners. The paper contributes to computers in education by reframing text-to-image GenAI as a teacher-mediated visual scaffolding process rather than a purely creative production process. It also provides methodological groundwork for future X+AI cross-domain design research, especially craft-vocabulary-guided interior scene generation and mixed-reality co-creation learning.

13:45
Beyond Action Sequences: Temporal Dynamics of Real-Time Learner Interaction Behavior in Immersive VR Learning Environments

ABSTRACT. Virtual Reality Learning Environments (VRLEs) offer strong potential for inquiry-based science learning, yet the temporal dynamics of learner interaction, such as behavioral onset, duration, and evolving transition structures, remain underexplored. This exploratory study analyzes interaction traces of learners engaging with a VR module on Electromagnetic Induction (EMI), comparing high-performing (HP) and low-performing (LP) groups (n=10 per group, post-test tertiles). Raincloud plots and Transition Network Analysis (TNA) with five successive 200-second sliding windows reveal distinct temporal signatures. HP learners exhibited efficient onboarding, early establishment of a tightly coupled Inquiry Diamond loop (parameter manipulation → positioning → experimentation → evaluation), and evaluation-led task closure. LP learners showed prolonged onboarding with premature evaluation, and persistent experimentation without convergence (wheel-spinning). A complete temporal inversion in dominant behaviors across temporal phases is identified between HP and LP. These findings suggest that performance in VRLEs is driven less by time-on-task and more by the structural systematicity of inquiry processes, with implications for precisely timed adaptive scaffolding that targets the right intervention at the right moment in each learner's session trajectory.

14:00
A Residual-Based DIF Tree Method for Fairness Analysis in IRT-Based Assessment

ABSTRACT. This study proposes the RDIF_R tree, which extends RDIF_R, a statistic used for detecting uniform differential item functioning (DIF) within the Residual-Based DIF framework, into a tree-based structure to identify DIF that may threaten test fairness and measurement validity. The RDIF_R tree explores possible split points based on examinee covariates without requiring groups to be defined in advance, thereby enabling the exploratory identification of subgroups in which DIF occurs and interaction structures among covariates. In this study, simulation data were generated based on the three-parameter logistic model (3PLM), and the detection performance of the RDIF_R tree was examined under conditions varying in the proportion of DIF items, DIF magnitude, Sample size and covariate structure. The results showed that the unpurified RDIF_R exhibited a substantial increase in Type I error rate as the proportion of DIF items increased, whereas the RDIF_R tree controlled the Type I error rate stably at a level comparable to the purified RDIF_R. In addition, the RDIF_R tree maintained high power and demonstrated its usefulness in exploring subgroups exhibiting DIF and interaction structures when compared with the Item-Focused tree. These findings suggest that the RDIF_R tree can be used as a method for exploring DIF associated with group characteristics that are not specified in advance, while preserving the computational efficiency and interpretability of the conventional RDIF_R.

14:15
Using Log Data to Interpret Gender DIF in a Performance-Based Digital Literacy Assessment

ABSTRACT. This study examined how gender-related differential item functioning (DIF) in a performance-based digital literacy assessment can be interpreted using log data. Log data from two items flagged for gender DIF in the 2024 National Digital Literacy Assessment were analyzed: Item 2-8 (a functional task) and Item 3-5 (a production task). For Item 2-8, behavioral sequence data were constructed and analyzed using optimal matching and agglomerative hierarchical clustering. For Item 3-5, time- and action-based variables were derived and analyzed using K-medoids clustering with Manhattan distance. The results showed that the identified behavioral clusters were associated with both performance outcomes and gender distributions. For Item 2-8, the DIF favoring males was associated with differences in file-selection strategies and program execution processes. In Item 3-5, the DIF favoring females was associated with differences in checking and applying the required conditions. These findings suggest that log data can provide process-based evidence for interpreting DIF and examining the fairness and validity of performance-based assessments.

14:30-17:10 Session 36A: C1 Session O
Location: Savoy West
14:30
A Multimodal Evaluation Framework for Classroom Instructional Support: Development and Validation

ABSTRACT. With the advancement of educational reform, instructional support has become a critical indicator for evaluating teaching effectiveness. However, existing evaluation frameworks for instructional support remain largely subjective and difficult to quantify. To address these limitations, this study integrates multimodal classroom data and artificial intelligence (AI) techniques to develop a multimodal evaluation framework for classroom instructional support. Grounded in the Classroom Assessment Scoring System (CLASS) framework and established educational theories, we developed an initial evaluation framework through a systematic literature review. We then refined and optimized the framework through expert interviews and a two-round Delphi method, and quantified indicator weights at all levels using the Analytic Hierarchy Process (AHP). A complete framework with 3 primary indicators and 7 secondary indicators was finally established. We verified the framework’s validity using 30 classroom videos analyzed by a customized Generative Pre-trained Transformers (GPTs) annotation model, which supports a zero-code automated annotation workflow. This study offers a theoretical basis and practical tool for the standardized evaluation and precise diagnosis of multimodal classroom instructional support.

14:45
Human–AI Collaboration in Higher Education: Tri-Rater Assessment and AI Learning Scaffolds

ABSTRACT. This study examines the dual role of generative AI (GAI) in a cross-disciplinary higher education course. On the instructor side, GAI was investigated as a calibration reference within a tri-rater framework involving Teacher A, Teacher B, and an AI Rater. On the student side, it was examined as a learning scaffold within a human–AI collaboration context. Adopting a repeated cross-sectional field study design, the broader study spans three consecutive semesters and is expected to include 297 students in total; however, the current submission reports analyses based on the two completed tracks available at the time of submission. Bland–Altman analysis and regression models were used to examine rating consistency and students’ technology acceptance. The findings indicate significant scoring discrepancies between the two human teachers, with both the direction and magnitude of these discrepancies varying across semesters. This pattern suggests the presence of systematic bias and rater drift, highlighting the difficulty of ensuring scoring consistency through one or two human raters alone. In the tri-rater analysis, the AI Rater did not function as a fully neutral scorer; however, its scores often provided a useful reference point for examining disagreement among raters, particularly when divergence between the two teachers was more pronounced. At the same time, the AI Rater showed larger discrepancies in some humanities-related tasks, suggesting that its performance remained sensitive to task characteristics. Overall, the magnitude of AI-related discrepancies did not appear to exceed the range of disagreement observed between human raters, supporting its potential value as a supplementary assessment tool. Further analysis showed that perceived ease of use predicted perceived usefulness, whereas system satisfaction depended primarily on perceived usefulness rather than ease of use alone. These findings suggest that the educational value of AI lies not in replacing teachers, but in supporting human–AI collaborative assessment as a practical calibration reference while also functioning as a learning scaffold for students.

15:00
Diagnosing Pedagogical Reward Hacking in AI Tutoring: An Evaluation-Time Framework

ABSTRACT. AI tutors can appear successful under proxy objectives while failing to support learning. A response may be fluent, confident, or correctly formatted, yet still skip reasoning, accept a misconception, or provide invalid intermediate steps. This paper proposes an evaluation-time diagnostic framework for identifying pedagogi- cal reward-hacking risks in AI tutor outputs. The framework combines adversarial prompt stress tests, automatic behavioral detectors, human pedagogical ratings, and multi-objective scoring. Using Claude Haiku 4.5 as the common foundation model, we evaluate five prompt-level conditions on a balanced GSM8K sample: base, reward shaping, RLHF-style preference prompting, CPO-style constraints, and the proposed diagnostic condition. No model weights are updated, so the study should be read as diagnostic evaluation rather than training-time mitiga- tion. Results show that the proposed condition improves over prompt-simulated diagnostic baselines on final-answer accuracy and receives the strongest human pedagogical ratings, while base Claude remains strongest on raw answer accuracy.

15:15
Unexpected Information with Cognitive Conflict and Surprise for Reactivating Web-Based Investigative Learning

ABSTRACT. Web-based investigative learning is a significant form of self-directed information exploration, in which learners construct knowledge by investigating Web resources. However, learners’ interest often declines as investigation progresses, which also cause an engagement decrease and premature termination of investigation. To address this issue, this paper proposes a method that uses generative AI to present unexpected information tailored to each learner. The proposed method is designed to elicit both cognitive conflict and surprise by presenting information that is difficult for the learner to explain based on what they have learned so far. We have implemented this method in a cognitive tool called iLSB (interactive Learning Scenario Builder), where learning scenarios created by learners and related logs are analyzed to estimate learners’ knowledge learned and to generate unexpected information adaptively. A case study was conducted to evaluate the proposed method. The results suggest that the presented unexpected information can induce surprise and cognitive conflict, enhance learners’ engagement, and promote subsequent Web-based investigative learning. These findings indicate the potential of generative AI not only as a tool for providing answers but also as a means of reactivating information investigation.

15:30
Investigating Explanations for Presentation Skills Learning in Active Video Watching

ABSTRACT. Explanations of artificial intelligence features in AI in Education (AIED) systems potentially can deepen learning and engagement. Explanations may clarify an AI’s decision-making process or help users deepen their understanding. However, explanations may also increase cognitive load or confuse students. Given this, it is important to evaluate the effectiveness of explanations and their impact on students. Previous research on explanations in Active Video Watching (AVW) showed that they have a positive impact on students’ learning of empathy skills. We enhanced AVW-Space, an AVW platform, by adding explanations of how the quality of comments students write on videos is determined. We present a study examining the impact of explanations in AVW on presentation skills. We compared data collected from a study performed in two consecutive years in the same first-year introductory engineering course at the University of Canterbury. In 2024, students only received a quality score for their comments. Meanwhile, in 2025, in addition to the quality indicator, students could also obtain the explanation of how the quality was determined. The results show an increase in the number of comments and their overall quality, shifting to more high-quality comments. This also translated to constructive behavior among those who accessed explanations. The results of causal modeling reveal the impact of explanations on the conceptual knowledge scores of students. This study verifies the impact of explanations in AVW and its applicability in different contexts. This study contributes to the widening research on explainability in VBL and AIED.

15:45
The impact of a personalized Multi-Agent System on secondary students’ design thinking and GenAI critical knowledge

ABSTRACT. Design thinking is essential for students in modern society, and integrated STEM education serves as a vital approach to cultivating this capacity. However, implementing STEM education poses significant challenges for teachers due to time constraints and a lack of multidisciplinary knowledge. While GenAI-driven Multi-Agent Systems (MASs) offer personalized solutions, their direct impact on students’ design thinking remains unclear. Furthermore, educators worry about students' overreliance on GenAI. Engaging students in complex, authentic problem-solving might mitigate this since GenAI lacks a physical understanding of the real world. However, its effectiveness in fostering critical knowledge is underexplored. To address these gaps, this study developed a personalized MAS to scaffold 38 secondary school students’ design thinking process during a semester-long STEM project. Pre- and post-test questionnaires were collected and paired-sample t-tests revealed significant improvements in both design thinking and GenAI critical knowledge. The findings indicate that a pedagogically structured MAS can effectively scaffold students’ complex engineering design processes. Crucially, by confronting GenAI's physical limitations during authentic tasks alongside teacher guidance, students developed epistemological vigilance rather than passive dependency, offering a scalable model for critical human-AI collaboration.

16:00
Teacher-Case-Driven RAG for Task Classification in Egocentric Agricultural Training Videos

ABSTRACT. Assessments in practical education vary across teachers and are difficult to encode as explicit rules. This paper presents an inference-time Retrieval-Augmented Generation (RAG) framework for adapting a Vision-Language Model (VLM) to a specific teacher without fine-tuning. The framework retrieves similar past assessment cases from the teacher’s assessment history and provides them as contextual examples during inference. We construct a dataset of teacher-annotated egocentric farming videos. The proposed framework is evaluated on task category classification and Good/Bad quality assessment. Experimental results demonstrate that presenting assessment cases improves both task category classification and task quality assessment. These findings suggest that RAG is an effective approach for reflecting teacher-specific judgment in video-based skill assessment.

16:15
Where Autonomy Ends: Evaluating Agentic LLM Reliability in Educational Administrative Workflows

ABSTRACT. The boundaries of reliable autonomy in agentic large language models (LLMs) remain insufficiently characterized, particularly in complex, document-intensive institutional workflows that require sustained reasoning across multiple information sources. This study investigates those boundaries within the context of educational administration by examining an Individual Program of Study (IPS) planning pipeline as a real-world case study. Using a rubric-based evaluation framework, the study assesses model performance across a multi-stage document analysis workflow, measuring accuracy, consistency, and reasoning quality at each stage of the process. The analysis identifies points at which autonomous performance degrades and traces the underlying reasoning behaviors associated with these failures. Findings indicate that retrieval augmentation alone does not ensure reliable execution of complex educational administrative tasks, especially when workflows require policy interpretation, cross-document synthesis, and judgment under ambiguity. Based on these results, the paper proposes a framework for AI delegation in curriculum planning that distinguishes tasks suitable for autonomous execution from those requiring human-in-the-loop (HITL) oversight. The framework offers practical guidance for institutions seeking to integrate agentic AI systems into administrative decision-making while maintaining reliability, accountability, and procedural integrity.

16:30
Self-Regulated Learning Profiles in SQL-Tutor: Trace-Based Learner Modelling of Repair and Progress

ABSTRACT. SQL-Tutor records how students submit SQL queries, receive feedback, and revise after constraint violations. This short paper uses these traces to examine whether students with similar prior knowledge worked with the tutor in different ways. The analytic sample included 83 students who completed both pre- and post-tests and had complete behavioral log measures. Seven standardized features were derived from the logs: adjusted learning-curve slope, learning-curve consistency, adjusted constraint recovery, help-seeking intensity, progression efficiency, improving-problem rate, and average logged-in time per attempt. Gaussian mixture modeling identified three profiles: Slow/Low Progress Learners, Efficient Progress Learners, and Persistent Self-Repair Learners. The profiles did not significantly differ in pre-test scores, but they differed in error trajectories, repair behavior, pace, progress, and learning efficiency. Post-test scores and normalized learning gains followed the expected direction but were not significantly different after correction. The analysis therefore treats repair and progress traces as learner-modelling evidence in their own right, not only as predictors of test scores. The study shows how domain-model traces from a constraint-based tutor can support interpretable profiles for future adaptive scaffolding.

16:45
Impacts of a Generative AI-Supported Concept Mapping Approach on Students’ Python Programming and Self-Efficacy

ABSTRACT. Generative AI learning companions are increasingly used in programming education, but students may still struggle to transform AI feedback into conceptual understanding and confidence. This study presents TaskInsighter-CM, a generative AI-supported learning companion that combines conversational self-explanation scaffolding with concept map–embedded error diagnosis. The system prompts students to explain their reasoning, identify conceptual errors in Python-related concept maps, receive rubric-based feedback, and revise their explanations. A quasi-experimental pretest–posttest study was conducted in a graduate Python course in Taiwan with 52 students. The experimental group (N = 35) used TaskInsighter-CM, while the control group (N = 17) followed the regular instructional arrangement. After controlling for pretest scores, the experimental group showed higher Python achievement and overall programming self-efficacy than the control group, with significant differences in Cooperation, Control, and Debug self-efficacy. Students also reported positive perceived usefulness and ease of use. The findings provide preliminary evidence that combining generative AI dialogue with concept map–based error diagnosis may support programming learning and process-oriented programming self-efficacy.

14:30-17:10 Session 36B: C1 Session P
Location: Savoy 2
14:30
Mirroring and Surprise to Understand You: Dynamical System Approach to Trace Learner’s States

ABSTRACT. Knowledge Tracing (KT) research in Intelligent Tutoring Systems (ITS) has long focused on estimating learners' knowledge states from response data, with recent advances in deep learning and Cognitive Diagnostic Models improving both accuracy and interpretability. However, a fundamental challenge remains: conventional KT approaches estimate knowledge states only at the moment of each response, and largely ignore the information contained in blank intervals—the time spans between consecutive responses. Moreover, existing models do not adequately capture the non-autonomous, nonlinear nature of learner dynamics, wherein the system's properties change over time as learning progresses. This paper proposes a dynamical systems framework for tracking learners' internal understanding states, motivated by how human teachers estimate learner understanding in practice. We argue that human teachers implicitly maintain an internal model of the learner and continuously refine it through prediction and error-based correction. From this perspective, we introduce a framework consisting of two components: a Mirror System Model and a Confidence Model. In the proposed framework, the learner is modeled as a non-autonomous nonlinear dynamical system whose internal state evolves over time. The learner's understanding is characterized by three global states—Not-Understanding (NU), Plateau (P), and Understanding (U)—each corresponding to a distinct attractor in phase space. The ITS maintains a Mirror agent that shares the same state variables as the genuine learner and generates predicted behaviors to compare against observed responses. The discrepancy between predicted and observed behavior, termed Surprise, is used to update the Mirror model dynamically. The Confidence Model manages the reliability of the current state estimate, capturing the degree to which the Mirror model approximates the true learner dynamics. To evaluate the proposed framework, computational simulations were conducted using three toy learner scenarios representing different learning trajectories. Results demonstrate that the Mirror agent with Surprise successfully tracks the learner's understanding state from observed response data. The framework offers high interpretability without requiring pre-training and highlights the importance of interaction-based refinement over static high-accuracy estimation as a design principle for ITS.

14:45
How Does AI-Empowered Socratic Questioning Influence Students’ Trust and Reliance?

ABSTRACT. Many students’ blind trust in generative AI (GenAI) tools has raised concerns about over-reliance and cognitive offloading in AI-assisted learning. This study distinguishes between two forms of trust: destructive trust and critical trust. Destructive trust refers to trusting AI mainly because it is convenient, efficient, or appears to work well, which may lead to over-reliance and negative learning consequences. Critical trust, in contrast, refers to trusting AI through careful evaluation, awareness of its limitations, and a balance between confidence and questioning. In the present study, we developed a Socratic Questioning AI Agent that defers direct answer generation and engages students through six types of Socratic questions, with the aim of exploring whether this approach can promote critical trust in AI. We examined how 40 postgraduate students used and responded to the designed agent. Self-reported questionnaire results showed an increasing trend in trust, a significant decrease in perceived ease of use, and no significant change in reliance. Further analysis of students’ reflective writing revealed active cognitive engagement during their use of the agent, suggesting the development of critical rather than destructive trust. These findings suggest that pedagogical design can reshape the typical relationship between trust and reliance in AI-assisted learning.

15:00
Development and Validation of an AI Research Literacy Scale for Research-Active Individuals

ABSTRACT. As artificial intelligence becomes increasingly integrated into research practice, research-active individuals increasingly rely on AI tools in activities such as research question generation, data analysis, and academic writing. However, validated instruments for assessing researchers' AI-related competencies remain scarce. This study introduces the concept of AI Research Literacy and develops a quantitative measurement scale. Drawing on AI literacy theory and the research activity process framework, we propose a five-dimensional model encompassing understanding of AI applicability in research, effective use, human–AI collaboration practices, evaluation of AI-generated results, and ethical awareness. Item generation was based on literature review and theoretical deduction, followed by expert evaluation and pilot testing, resulting in a 24-item scale. Exploratory and confirmatory factor analyses were conducted with a sample of 810 research-active individuals. The results support a five-factor structure with good model fit. The scale demonstrates satisfactory reliability as well as convergent and discriminant validity, and it performs significantly better than a single-factor model. The instrument provides a reliable tool for assessing AI-related competencies in research settings and offers empirical support for research training and governance practices in the context of AI integration.

15:25
When More Dialogue Does Not Help: Two Classroom Studies of Socratic Reflection in LLM-Supported Learning

ABSTRACT. Mathematics learners often develop procedural competence without corresponding conceptual understanding, in part due to limited opportunities for structured reflection in classroom-oriented digital environments. Prior research shows that reflection and self-explanation can support learning, but their design must balance effectiveness with instructional efficiency in real-world settings. We investigate the impact of Large Language Model (LLM)–supported reflection prompts embedded in a digital mathematics learning game for 5th- and 6th-grade students across two classroom studies. In Study 1 (N = 191), we compared three reflection formats following problem solving: (1) multi-turn, LLM-personalized Socratic dialogue, (2) open-ended self-explanation, and (3) multiple-choice self-explanation. All conditions produced significant learning gains, with no differences between formats; however, the more generative formats required substantially more time (1 and 2), indicating increased interaction burden without added benefit. To isolate whether these null effects were due to extended interaction or personalization, Study 2 (N = 33) compared a single LLM-personalized Socratic prompt with a standard open-ended self-explanation. Results again showed comparable learning gains, with no differences in outcomes or duration. Together, these findings identify a boundary condition for LLM-supported reflection: increasing dialogic interaction or personalization does not inherently enhance learning for 5th- and 6th-grade students and may introduce unnecessary interaction costs. For technology-enhanced learning systems, these results provide actionable design guidance -- prioritizing concise, targeted reflection over extended dialogue may better align with young students in the classroom. We discuss implications for designing efficient, classroom-ready reflective supports and consider when dialogic approaches may be more appropriate, such as for older learners.

15:40
LLM-Based Agent-Assessment for Supporting Quiz Correction in Learning by Quiz Creation

ABSTRACT. This study focuses on learning by quiz-creation as a learning strategy for fostering information literacy competence in Japanese education. It examines the extent to which LLM-based agent-assessment is reflected in learners’ quiz corrections. We propose a learning support system that enables learners to correct their own quizzes based on LLM-generated assessment results. An experimental study was conducted with 21 high school students in a quiz-creation activity on the Monte Carlo method. The results showed that 52.1% of valid assessment feedback items were reflected in learners’ corrections, suggesting that agent-assessment can support quiz correction. However, the reflection rates were polarized, indicating that the feedback was not equally effective for all learners. Feedback on typographical errors and terminology consistency was more likely to be corrected, whereas omission errors and feedback on the clarity of quiz statements were less likely to be reflected. These findings suggest that agent-assessment is effective for supporting formal improvements, but feedback presentation should be improved to reduce learners’ cognitive load during correction.

14:30-17:10 Session 36C: C7 Session I
Location: Savoy 3
14:30
Co-Creating Artifact-Mediated Participation Infrastructure for Language and Context-Responsive Science Teaching

ABSTRACT. This work examines how school-based micro-Communities of Practice (CoP) co-create, adapt, and circulate teaching artifacts to support diversity-responsive middle-school science teaching. Drawing on classroom observations, semi-structured interviews with middle school science teachers, and artifact data related to linguistically diverse Indian classrooms, an analysis is presented on how bilingual flashcards, diagrams, mnemonics, checklists, and digital routines are developed and used across classroom, school, home, and community contexts. The findings show that micro-CoPs involving teachers, counsellors, school leaders, parents, and external experts transform recurring participation barriers into reusable artifacts and routines. These artifacts function not merely as teaching aids but as participation infrastructure. The study contributes an artifact ecology account of language- and context-responsive teaching by showing how material, visual, linguistic, procedural, and digital artifacts mediate learner participation while carrying shared pedagogical repertoire across actors and contexts.

14:55
Empowering Learner Autonomy: The Impact of Selectable Generative AI Models in Math Scaffolded Question Generation

ABSTRACT. While generative AI offers highly adaptive scaffolding in education, its pervasive assistance risks undermining learner autonomy. Existing research predominantly treats LLMs models as the option of technical accuracy rather than analyzing how distinct model characteristics impact educational effects. To address this gap, this study proposes a paradigm shift toward active, self-regulated GenAI utilization. We integrated three distinct generative-AI models into an AI-powered math scaffolding tool, having students actively evaluate and select the model that best fits their learning needs. To evaluate this approach, we conducted an empirical study utilizing an adaptive question-generation tool and an e-book platform. We employed a Difference-in-Differences framework and an ordinal logistic regression model to rigorously analyze causality and the underlying psychological mechanisms. The results revealed distinct pedagogical trade-offs among the models regarding descriptive volume, conceptual density, and generation latency. Crucially, the sheer act of autonomously selecting a model specifically enhanced students’ subjective understanding. This selection did not artificially inflate motivation or relevance, rather than that these metrics remained grounded in students' prior academic abilities and system latency. Ultimately, this study demonstrates that only one LLM model would not be selected in technology design. Embedding explicit opportunities for autonomous model selection can empower human agency.

15:20
Can Generative AI Approximate Selected Aspects of Human Instructional Coaching Feedback? An Exploratory Comparative Analysis

ABSTRACT. Instructional coaching can support teachers to notice, interpret, and improve classroom practice, but access to expert coaching remains uneven because coaching is time-intensive and costly to scale. Generative artificial intelligence (GenAI) may offer a complementary mechanism for providing timely, rubric-aligned feedback. This exploratory study examines whether a GenAI-based instructional coach can approximate selected aspects of expert human coaching judgment when all coaches analyze the same classroom evidence under the same pedagogical framework. Two expert human coaches and one GenAI coach independently reviewed audio recordings from three classroom lessons in a K-12 school in Mumbai, India. All feedback was structured using Danielson's Framework for Teaching, Domain 3: Instruction. We compared feedback volume, issue-level convergence, scoring patterns, uniqueness of feedback, and time and cost requirements. Results show that the AI coach generated a comparable number of feedback points to one human coach, that 65% of clustered instructional issues were common across human and AI feedback, and that AI average scores were within the human-score range in two of three lessons. In the third lesson, the AI score was slightly above the human range. AI-generated feedback was substantially faster and lower cost than human coaching. These findings do not establish coaching effectiveness or substitutability. Rather, they suggest that, under constrained audio-only and rubric-aligned conditions, GenAI can participate in instructional noticing and first-pass feedback generation while requiring human oversight for contextual interpretation, prioritization, ethical use, and teacher uptake.

15:45
Effectiveness of Cluster Classroom Design on Foundation Students’ Non-Cognitive Learning Outcomes

ABSTRACT. This paper investigates the effectiveness of cluster classroom design on non-cognitive learning outcomes among foundation-level students in a Malaysian higher education context. Using a quasi-experimental pretest–posttest design, the study examines three key dimensions of a cluster classroom, namely, physical environment, teacher support, and social climate. Results from paired samples t-tests indicate significant improvements in teacher support and social climate within the experimental group. However, a one-way ANCOVA revealed no statistically significant difference in overall non-cognitive outcomes between the cluster and traditional classrooms. These findings indicate that changing physical layout alone does not cause significant improvement in students’ non-cognitive outcomes. Instead, the value of a cluster classroom lies in its role as a structural catalyst that successfully enables more responsive teacher support and collaborative social scaffolding.

14:30-17:10 Session 36D: C3 Session M
Location: Savoy 4
14:30
Designing Robot Roles for Effective Reflection on Shogi Matches

ABSTRACT. Reflection is essential for improving one’s skills in competitive games. However, for beginners, the perceived difference in skill between themselves and their opponents leads to a bias in their attention and in how they interpret their opponents' feedback, preventing them from engaging in effective reflection. Therefore, the opponent's skill level should be selected based on the learning stage and objectives. In this paper, we assign the roles of ‘professional’, ‘amateur’, and ‘beginner’ to a social robot as an opponent, and intentionally elicit the focus of reflection and the attitude towards receiving feedback from learners, thereby supporting their reflection effectively. The results of a case study show that learners playing Shogi matches to exchange opinions with the role-assigned robot can perceive the difference in skill as intended and take the intended attention and attitude towards receiving feedback from the robot.

14:45
A Web-based TPV Assessment System Using Personalized Gaze Heatmaps as Stimulated Recall Cues: Design and Evaluation

ABSTRACT. Teacher Professional Vision (TPV), the ability to notice instructionally relevant classroom events and reason about their implications for teaching and learning, is an important component of teacher expertise. Existing approaches for assessing TPV often rely on laboratory-based eye-tracking equipment, controlled environments, or resource-intensive assessment procedures, limiting their scalability in teacher education contexts. This paper presents the design and evaluation of a web-based TPV assessment system that combines webcam-based gaze tracking with stimulated recall. Using Webgazer.js, the system captures participants’ gaze behavior while viewing a classroom image on their own devices and generates personalized gaze heatmaps that are later used as recall cues to elicit reflective reasoning. 21 pre-service and in-service teachers completed a three-phase assessment protocol consisting of timed viewing, self-paced viewing, and heatmap-mediated stimulated recall. Gaze data were analyzed across predefined classroom Areas of Interest, while verbal responses were coded using a TPV reasoning rubric encompassing Description, Explanation, and Prediction dimensions. Results indicate that the protocol can be deployed outside laboratory settings and produces verbal data that can be reliably coded. Gaze patterns showed directional trends consistent with prior TPV research, while stimulated recall elicited richer action-oriented reasoning than unprimed observation alone. The findings demonstrate the feasibility of integrating low-cost webcam eye-tracking and personalized visual feedback into TPV assessment, offering a scalable approach for studying teacher noticing and reasoning in authentic settings.

15:00
Design and Evaluation of Compound Emotional Expressions in Social Robot

ABSTRACT. In face-to-face communication, individuals infer others' emotions not only from verbal content but also from nonverbal cues such as facial expressions, voice, and gestures. However, existing training methods for emotion recognition focus on single, clearly expressed emotions, despite real-world interactions often being complex and partially concealed. In this paper, we define emotional expressions in which positive emotions conceal negative ones as Compound Emotional Expressions. We also design robot-based emotional expressions to represent simple and compound emotions in a human-like manner, in which a social robot combines nonverbal behaviors including facial expressions, voice, and gestures to generate both single and compound emotional expressions. We conducted a case study to evaluate whether designed single and compound emotional expressions were recognized as intended. The results showed high recognition accuracy for single emotional expressions. Compound emotional expressions were also recognized to some extent and perceived contextually plausible. These findings indicate the possibility of representing compound emotions with a social robot through combined multiple nonverbal behaviors.

15:15
Efficiency and Classification Accuracy of Standard Error and Confidence Interval Stopping Rules in Computerized Adaptive Licensure Testing

ABSTRACT. This study compared standard error (SE) and confidence interval (CI) stopping rules in licensure computerized adaptive testing (CAT) using a simulation design that varied item difficulty, item discrimination, and examinee ability distribution. The SE rule produced shorter tests when item discrimination was at or above the medium level, whereas the CI rule yielded higher classification accuracy and lower error rates across conditions, despite requiring more items near the cut score. These findings suggest that the SE rule may be useful when test length must be constrained, while the CI rule is more appropriate when pass/fail classification validity is prioritized.

15:30
ScrumBoard: An Open-Source Project Management Tool as a Research Platform for Software Engineering Education

ABSTRACT. Team-based capstone projects in software engineering degrees often use project management tools, which can aid students in managing their projects and provide them with skills transferrable to tools they will encounter in the industry. However, commercial project management tools sometimes present issues in education contexts, such as useful features being paywalled or missing, and customization for the education context being limited because they are closed-source. Further, research dependent on data extracted from project management tools can also be limited by data commercial tools limiting what data can be exported. A solution for these issues is a purpose-built, open-source project management tool. In this paper, we present ScrumBoard, an open-source project management tool developed for both the education context and to support research. We discuss design decisions in ScrumBoard that are beneficial for conducting research with recommendations for researchers developing their own tools, and provide case studies of research projects ScrumBoard has facilitated to show practical applications of ScrumBoard’s research-oriented design. Through these case studies, we demonstrate that a purpose-built, open-source tool such as ScrumBoard, which is equipped with extensive logging and feature flags, and is customizable to our needs, enables research that would otherwise be difficult or impossible to undertake with commercial project management tools.

15:45
Who Does the Reasoning? An Activity–Community of Inquiry Analysis of Students’ AI-Mediated Thinking Across Two Pedagogical GAI Agents

ABSTRACT. Generative AI (GAI) agents are increasingly used in educational tasks, but less is known about how they redistribute students’ reasoning during learning. This qualitative multiple-case study analyzes interaction logs from two pedagogical agents used in computational-thinking lesson-design tasks. We define AI-mediated thinking as the interactionally visible ways learners’ reasoning, judgment, regulation, and agency are supported, delegated, or displaced during student–AI interaction. Using an AI-adapted Activity–Community of Inquiry (A–CoI) framework, episodes were analyzed through activity-theory components and the distribution of cognitive, teaching/guiding, and social presence across students, agents, and generated artifacts. Four modes were identified: delegated completion thinking, co-constructive object formation, regulated object grounding, and elicited reasoning synthesis. The Pedagogical Guided Agent tended to support artifact completion or co-construction, depending on student contribution, whereas the Socratic Questioning Agent delayed production, gated progress, and elicited students’ scenario details, variables, criteria, and decision rules before synthesis. Findings show that the educational meaning of GAI use depends on how agents configure the object of activity, participation rules, division of labor, and presences. Educational GAI agents should sequence generation after problem grounding, student-first elicitation, and pressure testing so that AI output synthesizes rather than replaces student reasoning.

16:10
Automated Assessment of Constructed Computational Artefacts

ABSTRACT. Automated assessment in computing courses often focuses on final answers, selected options, short text responses, or conventional program output. However, many computing tasks require students to construct structured computational artefacts whose correctness depends on internal representation, local well-formedness, and domain-specific behaviour. This paper proposes an assessment design pattern for such tasks, developed and illustrated in the context of a third-year artificial intelligence course. In the proposed pattern, code is used as the submission medium, but the assessed object is the artefact produced by that code. The marking code can then inspect the artefact structurally, check local validity conditions, and evaluate its behaviour according to the semantics of the relevant topic. The pattern is illustrated through three AI cases: belief-network modelling, constraint-satisfaction problems, and alpha-beta pruning in games. These cases show how automated tests can distinguish between different failure modes, including representational errors, modelling errors, parameterisation errors, and errors in reasoning traces. The approach was implemented in Python and deployed through Moodle and CodeRunner using server-side execution. Deployment across formative and summative assessments suggests that the pattern is practical within existing assessment infrastructure and can support more diagnostic feedback than final-answer grading alone.

16:35
A Five-Layer Architecture for Managing Uncertainty in Multimodal Learning Analytics Data

ABSTRACT. MMLA systems estimate educational events from sensor data, but current LRS designs record uncertain estimates as definitive facts, risking distortion of instructional decisions. This study proposes a five-layer hierarchical architecture (Raw Data, Feature, Interpretation, Event, Context) that preserves inference uncertainty and controls its propagation to educational decision-making. We evaluate the architecture using a teaching behavior recognition system and confirm that the Event layer gate mechanism effectively suppresses pedagogically harmful misclassifications while minimizing loss of correct events. We discuss implications for interoperability and cross-institutional data exchange.

14:30-17:10 Session 36E: C3 Session N
Location: Windsor
14:30
Privacy-Preserving LLM-Assisted Learning Analytics with Natural Language Interfaces for Educational Data

ABSTRACT. Educational institutions increasingly collect multi-source student data, including attendance records, academic performance, daily mood self-reports using a weather metaphor, and health room usage. These data can support learning analytics and evidence-informed educational decision-making, but their use is constrained by privacy risks, fragmented data systems, and limited technical expertise among practitioners. Recent large language models (LLMs) offer promising natural language interfaces for data exploration; however, LLM-assisted analytics may expose sensitive records, enable small-group re-identification, or support inappropriate profiling if connected directly to student-level databases. This paper presents a privacy-preserving LLM-assisted learning analytics platform designed for secure educational environments. The system integrates municipal school data through a four-layer architecture consisting of raw, clean, feature, and Model Context Protocol (MCP) layers. The raw and clean layers store original and normalized student-level data, but they are never exposed to the LLM or end users. Instead, the LLM accesses only aggregated, k-anonymity-protected feature tables through whitelisted MCP tools. The platform combines architectural separation, schema guards, read-only access, k-anonymity suppression, and audit logging to protect against unauthorized access, small-group disclosure, and unsafe automated operations. The prototype was implemented using PostgreSQL, Python-based ETL pipelines, MCP tools, and a locally deployed LLM environment. We demonstrate how the system enables natural language queries about attendance, well-being, and learning outcomes while enforcing privacy-by-design principles. This work contributes a reproducible design pattern for responsible LLM-assisted learning analytics in sensitive educational contexts.

14:45
Digital Twin Enhanced Simulation in Healthcare Education: Designing Adaptive Training Pathways for Clinical Reasoning and Decision-Making

ABSTRACT. Healthcare education has been resoundingly improved by the adoption of digital technologies and personalization, integration of data with the simulation-based learning. Still, though, much simulation environments are founded on predefined training routes which would be unable to cohere with the evolving clinical reasoning expertise diagnostic performance and cognitive-affective state of the learner. In this work, we introduce an adaptive simulation framework backed by a digital twin to deliver tailor-made clinical training in healthcare education. Within the framework, each learner is represented as a dynamic digital twin wherein learner-profile information (age, education level scales, and clinical simulation performance) determines its diagnostic accuracy, decision time, confidence, cognitive load, engagement (e.g., arousal vs. relaxation) and feedback type-level & adaptation according to previous studies. It builds a learner digital twin that is updated for every iteration of the simulation and then applied to guide adaptive training pathways, case complexity and feedback strategies. The method used was a quantitative research design with simulation-based activities to explore the data of how learners progressed in their learning. This study analyzed learner-session data to compare pre- and post-clinical reasoning scores, differences in diagnostic accuracy and decision time between simulation rounds, patterns of learner state indicators using clustering and used multivariate logistic regression to identify predictors of mastery status. It will introduces a structured approach for the design of intelligent healthcare education systems which may address a learner-based needs, improve clinical reasoning, assist in diagnostic decision making and improve professional skill development.

15:00
Multimodal Learner Engagement Recognition in Online Learning Through Facial Expression and Eye Gaze Fusion

ABSTRACT. In online learning systems, student engagement is a key factor that determines the effectiveness of learning outcomes. Existing automated approaches often rely on unimodal analysis, such as facial expression recognition or eye-gaze monitoring alone, which fails to capture the complex interaction between affective states and attentional dynamics. To address this limitation, this study proposes a hybrid multimodal deep learning framework that integrates facial expression cues with eye-gaze behavior for robust engagement detection. Convolutional Neural Networks (CNNs) are employed to extract spatial facial features, while Long Short-Term Memory (LSTM) networks model temporal gaze sequences to capture attention dynamics. A hybrid fusion strategy combining feature-level concatenation and decision-level weighting is applied to exploit complementary information across modalities. Experiments conducted on the DAiSEE benchmark dataset demonstrate that the proposed multimodal fusion model significantly outperforms unimodal baselines, achieving an accuracy of 87.1% and an F1-score of 0.85. These findings confirm the effectiveness of multimodal integration for resolving ambiguous engagement states and enabling scalable real-time engagement assessment in online learning environments.

15:15
Tools over Subjects: Tag-Based Exploration and Co-Browsing in a Learning Analytics Case-Sharing Portal

ABSTRACT. Adoption of learning analytics is hindered by a lack of locally relevant practice cases. To bridge this gap, we developed the LEAF Evidence Portal, a repository of case articles that describe the adoption of learning analytics in classrooms. The case articles are tagged by different tag categories such as school level and tool. In this paper, we analyze two months of portal access logs to identify what drives user exploration of case articles. The results show that learning analytics tool is the primary driver for user interest and co-browsing. Conversely, users tend to browse across different subjects. These findings suggest that recommender systems for case articles should prioritize tools over subject-specific filtering to better support learning analytics adoption.

15:30
Development and Evaluation of a Personality-Based Serendipity-Oriented Location Recommendation Chatbot

ABSTRACT. This study presents the development and evaluation of a personality-based, serendipity-oriented location recommendation chatbot. While serendipity has been widely discussed in recommender system research, few systems have operationalized personality traits as a mechanism to intentionally induce serendipitous experiences. To address this gap, we designed a chatbot that incorporates users’ Big Five personality traits (TIPI-J) into the recommendation process. Specifically, one personality dimension was inverted to generate recommendations that partially deviate from users’ typical preferences while maintaining contextual relevance. The system was implemented using a large language model (LLM) to dynamically generate recommendations based on personality traits and users’ prior location selections. An experimental study was conducted with 38 university students, comparing a serendipity-oriented condition (n = 20) with a non-serendipity control condition (n = 18). Participants evaluated recommended locations in terms of interest, unexpectedness, perceived serendipity, and willingness to visit. Results showed that although differences in interest and perceived serendipity did not reach statistical significance, medium effect sizes were observed for unexpectedness and serendipity. Importantly, willingness to visit was significantly higher in the serendipity-oriented condition (Mann–Whitney U = 116, p = .016), with a moderate effect size. Ordinal logistic regression further indicated that group membership significantly predicted visit willingness even after controlling for personality traits, suggesting that the observed effect was attributable to the system design rather than individual personality differences. These findings demonstrate the feasibility of inducing serendipitous experiences through personality-based design strategies in chatbot-driven recommendation systems. The study contributes to the development of LLM-based recommender systems that move beyond preference matching toward intentional experience design.

15:45
Communicating Urban Environmental Data: A Scoping Review of Citizen-Facing Environmental Dashboards

ABSTRACT. The increasing availability of urban environmental data creates new opportunities for residents to access information about air quality and environmental noise. However, access to data does not necessarily translate into the ability to interpret and use it in everyday decision-making. This study examines citizen-facing environmental dashboards as a communicative and potentially educational layer between complex environmental monitoring systems and urban residents. A scoping review methodology was employed. The review included 52 open-access publications and was designed as a descriptive and interpretive evidence-mapping study. The analysis identified seven functional types of digital environmental interfaces. Based on the synthesis, a seven-level model of environmental communication is proposed. The findings indicate a gradual shift from dashboards focused primarily on presenting environmental indicators towards systems that support spatial and temporal interpretation, forecasting, risk assessment, scenario analysis, recommendations, and natural-language interaction. The findings may support urban authorities, environmental agencies, public health organizations, and civic technology designers in developing more pedagogically informed digital interventions that facilitate residents’ understanding of environmental data and support action to reduce pollution.

16:00
Identity Flagging from Cursor Traces via Disentangled Identity Embeddings and Clustering Methods

ABSTRACT. The computer mouse is a ubiquitous human-computer interaction device that requires delicate and fine-grained motor control. Explicit training of mouse motor skills could improve the efficiency and effectiveness of tasks performed with a computer. To empower the next generation of mouse tutoring systems, we need reliable computational models that can help a system distinguish between user behavior, task features, and noise. We re-implemented a state-of-the-art model, the DisMouse disentanglement framework, on the Balabit Mouse Dynamics Challenge Dataset, and evaluated whether disentangled representations can support test-taker identity signatures and integrity-risk flagging from interaction logs. We observed that the model’s identity space forms clear user-specific clusters, while the model’s dynamics space remains largely user-agnostic. We then aggregated window-level representations into session-level signatures, and applied clustering methods to flag behaviorally aberrant sessions with 71-73% accuracy. Our results suggest that disentangled identity embeddings from mouse movement are interpretable and can serve as a practical signal for behavior detection.

16:15
Semantics-Aware Augmentation of Educational Robotics Programs

ABSTRACT. Collecting large and diverse datasets of student robotics programs is inherently difficult because activities depend on physical hardware, classroom settings, and teacher-designed tasks, limiting the data available for AI-supported learning analytics and automated program analysis. This paper presents a semantics-aware data augmentation framework for block-based educational robotics programs. We represent LEGO SPIKE Prime projects as mutable block graphs and define a set of controlled, semantics-preserving transformations over visible, pedagogically meaningful constructs including ports, units, actuator commands, variables, expressions, conditionals, and loops. Semantic equivalence is grounded in an abstract robot model, enabling transformations to preserve program behavior independently of concrete hardware configuration. An empirical evaluation on approximately 40,000 student programs shows that augmentation strength is easily controllable, and that generated variants increase both structural and parameter-level diversity. In a downstream task-classification experiment, augmented training data improved accuracy for each evaluated classifier family. These results demonstrate that semantics-preserving transformation-based augmentation is a practical strategy for expanding educational robotics program datasets and supporting AI-based program analysis in a domain where large-scale data collection remains a persistent challenge.

16:30
Real-Time Prediction of Student Failure from Java Quiz Interaction Logs

ABSTRACT. Real-time prediction of student failure may enable timely intervention during programming activities. Using 29,059 Java quiz attempts from 877 students, we investigate whether eventual failure can be predicted from partially observed interaction traces. We analyze 121 debugging and code-writing questions using a prefix-based framework and compare a Hidden Markov Model (HMM), a Long Short-Term Memory network (LSTM), and XGBoost. At a representative threshold of 0.70, HMM and XGBoost achieve higher recall and longer median lead times than LSTM, while LSTM produces substantially lower false-positive rates. Median lead times range from approximately 5–10 minutes across models and task types, indicating that partial interaction traces can provide an actionable window for intervention. These findings demonstrate the feasibility of early-risk detection from programming interaction logs while revealing important tradeoffs among detection coverage, alert timeliness, and false-alarm rates.

16:45
A Syntactic Probabilistic Representation of Programming Assignments for Controllable Similarity Analysis

ABSTRACT. Programming assignments are important learning materials in introductory programming education, and computational representations of assignments can support tasks such as curriculum analysis, assignment organization, and similarity-based retrieval. However, programming assignment similarity is not uniquely defined: two assignments may be considered similar with respect to their overall solution structure, loop usage, conditional branching, required data structures, or other aspects. Existing code-based approaches often represent assignments from a fixed perspective, making it difficult to control which structural aspect is emphasized in comparison. This study proposes an AST-based probabilistic representation of programming assignments that enables controllable similarity analysis from selected syntactic perspectives. We evaluate the proposed representation using 47 assignments and learner solutions from an introductory Python programming course. The results show that changing the focal syntactic element produces different and interpretable similarity relationships among assignments. An expert evaluation by the course instructor further suggests that the proposed representation can identify assignment relationships that are meaningful within the target course context. These findings indicate that the proposed representation provides an interpretable and controllable way to analyze programming assignments through solution-code structure.

14:30-17:10 Session 36F: C5 Session C
Location: Clarendon
14:30
Gamified Gaze Interaction in Immersive Virtual Reality for Attention Training: Effects on Learner Performance
PRESENTER: Dai-Yi Wang

ABSTRACT. Sustaining attention is a critical challenge in multimedia-rich digital learning environments, where learners are frequently exposed to multiple sources of information. While gamification and immersive technologies have been widely used to enhance engagement, limited research has examined how gaze interaction can be integrated as an active mechanism for attention training. This study proposes a gamified attention training system implemented in an immersive virtual reality (VR) environment with gaze-based interaction. The system incorporates goal-oriented visual search tasks, immediate feedback, and progressive challenges, while transforming gaze behavior into a core interaction modality that directly drives task performance. A quasi-experimental design was employed to compare training effects between VR and computer-based (PC) environments, as well as across learners with different attention tendencies. The results indicate that learners demonstrated significant improvement in visual search performance, particularly in task completion time, after training. The VR condition showed stronger improvement compared to the PC condition, suggesting that immersive gamified environments can enhance attentional engagement. In addition, learners with attention-deficit tendencies exhibited more pronounced improvement in the VR condition. This study contributes to the field of educational gamification and game-based learning by demonstrating how gaze interaction can be integrated as a gameplay mechanic within immersive environments, supporting both attention training and attention-aware learning design.

14:45
Flow Experience in a Scaffolded 2D-to-3D Game Development Transition
PRESENTER: Jeeyoung Hwang

ABSTRACT. This study compared primary school students' flow experiences between 2D game development and its 3D codebase extension. Twenty-nine fourth graders built a 2D maze game in MakeCode Arcade and, one week later, extended the same code into a 3D maze game through a Raycasting extension. Flow was measured after each condition using a nine-dimensional Learning Flow Scale. Overall flow scores were significantly higher in the 3D condition than in the 2D condition, particularly in three of the nine sub-dimensions. The findings suggest that scaffolded 2D-to-3D transitions may offer an effective pathway for sustaining motivation and enhancing engagement in game-based programming education.

15:00
A Gamified Vocabulary Learning System Integrating Self-Regulated Learning for Elementary Students: Effects on Achievement, Self-Efficacy, and Satisfaction

ABSTRACT. With the digital transformation of education, stimulating schoolchildren's motivation to learn English and cultivating their digital self-regulated learning capabilities have become critical challenges in contemporary teaching. This study developed a digital vocabulary learning system that integrates game-based learning elements with Self-Regulated Learning. Tailored to the characteristics of English vocabulary, the system constructs four core modules: image-based learning, immediate feedback mechanisms, process management and tagging, and external motivation conversion extension. A six-week instructional experiment was conducted with 53 fourth-grade elementary school students. The results of ANCOVA and t-tests indicate that the experimental group utilizing this system significantly outperformed the paper-based control group in both English vocabulary retention (p = .006, ηp² = .142) and learning self-efficacy (p<.001, ηp²=.412), while also demonstrating an exceptionally high level of learning satisfaction (d = 1.41). This study confirms that the synergistic effect between digital gamified tools and self-regulated learning strategies can effectively assist students in self-monitoring, thereby deepening the self-regulated learning cycle from both efficacy and motivational dimensions.

15:15
Gamifying the Learning of Data Structures and Algorithms in Python

ABSTRACT. Gamification in education is increasingly prevalent. However, its application to complex and abstract learning domains remains less explored. One such domain would be the learning of Data Structures and Algorithms (DSA). DSA presents a challenge as learners are required to strengthen the relationship between four different parts: abstract concepts, visual representations, algorithmic logic and code implementation for any DSA concept to be considered completely learnt. Many tools that already exist support learning only part of the learning process, while the learners need help in going through the full process, from concept to code. A key missing link is the connection between conceptual understanding, algorithm logic, and code implementation. This paper aims to solve this problem by connecting the dots, with a stronger linkage between concepts, visualisation and code. To best represent each concept and its intricacies with games, this project explores the different game genres: puzzle, board, card and action games, can represent different DSA topics. This will enable players to have an improved learning experience with better conceptual understanding of the DSA concepts as they engage and interact more with the games as compared to traditional learning methods of DSA.

15:30
Urban Choices: A Serious Game for Urban Sustainability Education
PRESENTER: Samarth Kumar

ABSTRACT. Serious games are widely advocated for environmental education, yet few are designed rigorously or measured in ways that preserve engagement. This paper presents Urban Choices, a browser-based narrative game for India's undergraduate environmental studies curriculum, designed around the Design-Play-Experience (DPE) framework to make measurement an integral part of play rather than a post-hoc add-on. The game operationalises UNESCO's eight sustainability competencies through eight scenes spanning a character's first three months in a shared apartment. A cascade architecture embeds early consumption decisions as hidden consequences that surface later, operationalising systems thinking without explicit instruction. The theoretical framework-DPE, ICAP explanation probes, and competency-aligned learning design divides labour among disciplines (game design, learning science, narrative) by giving each a lane while making trade-offs explicit. A pilot playtest with 43 undergraduates found that (1) in-flow telemetry captured every research-critical event with 97.7% completeness and 100% scene-to-capstone retention, demonstrating that unobtrusive measurement and engagement can coexist; (2) players' pre-play framing persisted into later unrelated decisions at well-above-chance rates, supporting the measurement of prior knowledge activation; and (3) explanation depth rose across three ordered probes from active to constructive reasoning, consistent with knowledge construction observable inside gameplay. The paper contributes a design pattern hidden consequences, measurement folded into fiction, discipline-bridged architecture applicable beyond this curriculum, and releases the artifact and telemetry schema openly for re-authoring.

15:45
Uncertainty Design and Learner Curiosity in a Japanese Onomatopoeia Learning Game A Prior Knowledge Perspective

ABSTRACT. Uncertainty design is considered an important element in promoting engagement and curiosity in digital game-based learning (DGBL). However, prior research has often treated uncertainty as a binary design factor, focusing mainly on whether uncertainty is present or absent, while paying limited attention to how different uncertainty mechanisms are structured. In particular, little is known about how such mechanisms shape different forms of curiosity or how learners’ prior knowledge may condition these effects. To address this gap, this study examined two design-level uncertainty conditions that differed in information transparency and event predictability: predictable uncertainty (PU) and unpredictable uncertainty (UU). An experimental study was conducted with 55 undergraduate students using “Dream Island”, a selfdeveloped game for learning Japanese onomatopoeia. Participants were assigned to either the PU or UU condition, and subgroup comparisons were further conducted based on prior knowledge level. The results showed no significant differences between the PU and UU conditions on overall curiosity measures. However, within the UU condition, learners with low prior knowledge reported higher deprivation-type curiosity than those with high prior knowledge, whereas this pattern was not observed in the PU condition. Interest-type curiosity showed no clear between-group differences. In addition, participants in both conditions showed significant and comparable improvements from pretest to posttest. These findings suggest that the effects of uncertainty design are not uniform, but vary according to curiosity type and learners’ prior knowledge. The study offers evidence for a more mechanism-oriented and personalized approach to DGBL design.

17:10-18:00 Session 37: Closing Ceremony

Closing Ceremony of ICCE 2026

Location: Savoy Ballroom