Title: The Great Convergence: Recommendation, Search, and Conversation in Real-World AI Systems
Abstract: For most of their history, recommendation, search, and conversational interfaces have been built as separate systems, with distinct architectures, evaluation methods, and engineering teams. This separation is rapidly disappearing. Embeddings, foundation-model backbones, and large language models are pushing these three paradigms toward a single underlying stack, in which a query, a click, a scroll, and a natural-language request are increasingly different entry points into the same retrieval-and-ranking machinery. This talk examines what such convergence means in practice for the applications we build, and the new challenges it brings, both theoretical and practical. On the theoretical side, unifying interaction-based and semantic signals raises hard questions about evaluation, cold-start behavior, bias and alignment. On the practical side, deploying these systems in the real world forces an uncomfortable trade-off: they must be not only accurate, but also fast and economically viable at the scale of hundreds of millions of users — a constraint that reshapes architectural decisions far more than benchmark leaderboards suggest. Drawing on experience from both academic research and the production systems behind Recombee, the talk explores how to scale converged recommendation–search–conversation systems while keeping latency and cost under control, and how to ensure user satisfaction and AI alignment across very different domains, from news and media, through e-commerce, to education. The central argument is that the next generation of information systems will be defined less by raw model capability and more by how thoughtfully we engineer the interaction between accuracy, affordability, and alignment.
GenAI and IS/Computing Professionals: A Framework for the Evolution of Practice
ABSTRACT. Generative AI is rapidly reshaping Information Systems and Computing practice, yet its operationalization within organizations remains insufficiently understood. The study draws on focus groups with experts from a large multinational consultancy to examine how GenAI is currently integrated into business operations by IS and Computing professionals. Findings indicate that most organizations operate at early stages of AI maturity where GenAI functions primarily as a tool or assistant with some emerging uses of supervised agents. Building on these insights, the paper proposes a multidimensional GenAI Operations Progression Framework that outlines five maturity stages progressing from GenAI as a tool to GenAI workforce model. The framework integrates Work Practices and AI-Human collaborations, operational models of AI maturity, organizational enablement from education to institutionalization and governance from basic GenAI acceptance to AI embedded processes and operations.
Understanding Gamification Management: Practitioner Perspectives on Project Organization and Challenges
ABSTRACT. Gamification is widely used to increase motivation and engagement, yet implementations often remain ad hoc and poorly structured, leading to inconsistent effects and ethical concerns.
While prior research has primarily focused on the design of gamification, less attention has been paid to how gamification is managed in practice.
To address this gap, we introduce the concept of \textit{Gamification Management} and investigate how practitioners understand, organize, and experience challenges in gamification projects. We conducted semi‑structured expert interviews with nine practitioners and applied qualitative content analysis.
Our results reveal that practitioners largely share a common conceptual understanding of gamification but organize projects in highly heterogeneous ways, with limited use of structured frameworks, testing strategies, and formalized decision‑making processes.
Keychallenges include a lack of resources, insufficient knowledge of game design and gamification theory, as well as, process‑related difficulties.
We argue that the recurring quality issues in gamification are not solely a design problem but also a management problem.
Our findings highlight the need for more systematic approaches to managing gamification and contribute to a better understanding of gamification as an organizational practice within IS development and operations.
ABSTRACT. While Artificial Intelligence (AI) is a transformative technology, its organizational adoption remains a major socio-technical challenge.
To structure the fragmented discourse relevant to Information Systems (IS) Engineering, this paper conducts a bibliometric analysis of 1,575 Web of Science publications (2019–2026).
The analysis identifies four foundational strands: acceptance theories, strategic firm-level adoption, human-AI trust, and innovation diffusion.
Thematic mapping reveals technical infrastructure, organizational readiness, Human-AI Interaction and Governance, and contextual concerns as central dimensions.
A comparison with socio-technical frameworks and recent reviews suggests that current research remains fragmented across these dimensions.
Interpreted through Requirements Engineering literature, these findings indicate that adoption concerns such as trust, explainability, and ethics can inform non-functional requirements (NFRs) and socio-technical design considerations for more adoption-aware IS Engineering.
ITLingo-Chatbot: A Conversational Assistant for Knowledge-Grounded Requirements Engineering
ABSTRACT. Requirements engineering remains a critical yet labor-intensive phase of software development, in which unrestricted natural-language specifications often introduce ambiguity, inconsistency, and incompleteness. Recent advances in large language models (LLMs) create new opportunities to support requirements documentation, but directly applying these models can reproduce the limitations of informal natural language. This paper presents the ITLingo-Chatbot, a knowledge-grounded conversational assistant designed to support the creation and refinement of structured requirement specifications. The system integrates LLMs with retrieval-augmented generation (RAG) using a curated knowledge base that includes grammar definitions, validation rules, specification examples, and theoretical documentation for multiple specification languages and controlled natural languages. A prototype implementation demonstrates how users can generate, refine, and validate specification fragments through an interactive workflow. Generated artifacts may subsequently be manually validated using existing ITLingo tools.
A Knowledge Graph Treatment to the IO-BPG Extension of BPMN
ABSTRACT. This paper reports a deployment with knowledge graph capabilities for a BPMN extension previously presented in the recent IS literature – namely the IO-BPG extension for interorganizational business process governance. According to the rationale of that extension, BPMN is insufficiently expressive to describe collaborative networks – as they require business processes to be distributed across multiple organizations, raising governance concerns related to data security, roles, risk-related operations and regulatory compliance. The IO-BPG extension aiming to fill the gap was originally designed through a Design Science Research process, however its deployment was focused on the notation and conceptual levels, leaving the tool-level operationalization out of scope. The work reported in this paper applies the Agile Modelling Method Engineering (AMME) methodology to develop a modeling tool that provides the visual syntax of the original IO-BPG extension augmented with a Knowledge Graph treatment, to enable semantic queries and traceability over any diagrammatic content created with IO-BPG – a necessary enabler for the reporting capabilities envisioned by the original proposition. Feasibility is demonstrated in two scenarios, one from the original IO-BPG paper and one newly introduced here.
Adaptive Model-Driven Inference in Intelligent Information Systems Using the DIPAS Framework for River Monitoring
ABSTRACT. Environmental monitoring systems operate under uncertainty due to incomplete observations, measurement noise, and limitations of static models. This paper proposes a model-driven framework for environmental information systems built upon the authors’ original concept of the Dynamic Intelligent Process Automation System (DIPAS). The approach integrates structured domain specifications, process-based models, and adaptive inference within a unified architecture. Domain models are treated as formal representations that are continuously transformed into executable inference mechanisms, forming digital twins that are continuously updated using streaming sensor data. DIPAS employs a Predictive Data-Adaptive Learning Mechanism (PDALM) that improves model-based estimation through adaptive, innovation-driven adjustments. The framework demonstrates how conceptual models can be operationalized within AI pipelines and used in real-time analytics. A large language model (LLM) interacts with structured model outputs, supporting model-based reasoning and decision-making. The proposed framework contributes to the development of information systems by integrating requirements, models, and AI-based implementation.
HEXACO Personality Dimensions and HEXAD User Typology: A Correlational Analysis
ABSTRACT. This study examines the relationship between HEXACO personality traits and HEXAD gamification user types in higher education. Gamification is used to enhance engagement and motivation, yet its effectiveness varies between individuals. To better understand these differences, this study adopts an exploratory approach combining personality traits and user typologies. The data were gathered from 26 university students using only validated HEXACO and HEXAD questionnaires. The Pearson correlations analysis revealed that Extraversion was positively associated with multiple HEXAD user types, including the Philanthropist, the Socialiser, the Free Spirit and the Achiever. In contrast, the Disruptor type showed negative correlations with Honesty–Humility and Emotionality dimensions. The findings provide preliminary evidence that personality traits may be associated with gamified system engagement and highlight the importance of personalized gamification design. Despite the small sample size, this study provides preliminary insights into individual differences in gamified environments and offers directions for future research.
Development of Engineer 5.0 Competencies: A Case Study of Team Research Projects
ABSTRACT. Industry 5.0 emphasizes the importance of technical and human aspects that are crucial for addressing economic and social challenges. The new requirements for the competencies needed to meet the demands of the contemporary world refer directly to engineers who are responsible for the development and implementation of new technologies. The article presents models of Engineer 5.0 competencies in the context of Industry 5.0. The main goal of the paper is to present an educational project established at Gdańsk University of Technology, Team Research Projects, as an example of the Project-Based Learning (PBL) method, designed to support engineering students’ development of interpersonal and analytical competencies required for Engineer 5.0. The concept of the course is presented, examples of TRP topics are provided, and the process of building project teams is described. Students’ opinions regarding the educational outcomes of the course are also discussed.
A Comparative Analysis of Lecturer and Student Attitudes Toward Generative AI Policies in Higher Education
ABSTRACT. This article presents a study of the attitudes of students and lecturers at a single Polish technical university toward institutional policies governing the use of generative artificial intelligence (Gen AI). The study was conducted to understand what kind of policy the academic community expects: one that protects academic integrity while not hindering innovation and that is perceived as fair. The survey involved 144 respondents from various study fields and used an online questionnaire covering eight topic areas. Three findings stand out: three-quarters of respondents do not know whether their institution has a Gen AI policy; students and lecturers diverge sharply in how they judge the impact of Gen AI on academic integrity and independent thinking; and both groups reject outright bans while favouring clear, selective rules. Engineering students show greater acceptance of AI generation than students in the humanities and social sciences. Because the data come from one institution, the results are exploratory; we situate them against comparable international studies and derive a three-tier hybrid policy model.
Empowering and Impactful IS Education for the Young Generation
ABSTRACT. Youth remain an underexplored stakeholder group in Information Systems (IS) research. Educating them for their digital future should be a pivotal concern of ours. This study, inspired by the Participatory Design (PD) tradition, explores ways by which to offer empowerment and impact-oriented IS education. Literature on IS education lacks PD inspired approaches. We conducted ten participatory game design workshops with youth, who designed and evaluated a literacy game for youth. Their contributions shaped multiple aspects of the game: character design and visual elements were implemented directly, while other ideas evolved into new features or inspired them. Their input also influenced game mechanics and narrative flow, and some ideas travelled beyond the initial project, informing virtual reality environments. We demonstrate an example of PD inspired, empowerment and impact-oriented education, which enabled youth to impact not only immediate design outcomes but also future development trajectories.
Elaborating Explainable AI for Software Fault Prediction: Interpretability Techniques and Performance Insights
ABSTRACT. Software Fault Prediction (SFP) is an important task of software engineering, which attempts to discover the faulty modules in source code proactively to improve quality and reducing maintenance costs. The machine learning (ML) models has significantly enhanced fault prediction capability, yet these methods are black box, and limits in interpretability and trust. Explainable AI (XAI) approaches, such SHapley Additive exPlanations (SHAP), are useful to get insights, and quantify the impact of specific characteristics to the model decisions. The study proposed a robust and interpretable framework for SFP which integrates class imbalance handling (SMOTE), optimisation (using hyperparameter tuning (SMOHY)), SHAP-based explainability and Feature Sensitivity (FS) analysis to produce accurate and transparent models using 41 ML classifiers against 41 opensource java projects. SMOHY improved the efficiency and predictability of the SFP model by 22.38% compared to ORGD and 6.5% compared to SMOTE, attaining highest median AUC value of 0.82. ExtraTrees and RandomForest algorithms achieve the highest median accuracies of 87% and 86%. The SHAP values revealed RFC as the most significant feature, offering a visual depiction of its influence on model predictions. FS analysis demonstrated that RFC metrics are the most important metric affecting SFP model performance, validating SHAP’s applicability in SFP.
From Technology to Trust: Advancing Human-Centric AI and Digital Sovereignty
ABSTRACT. As 5G-enabled smart cities to become increasingly data-driven and interconnected, the rapid use of AI raises pressing concerns about privacy, ethics, and individual autonomy. Without Human-Centered AI (HCAI) frameworks and protections for digital sovereignty, citizen trust in urban technologies is at risk, threatening long-term societal acceptance. This paper examines how HCAI can be operationalized to build trustworthy smart-city ecosystems. It examines the intersection of digital sovereignty, privacy, and governance and proposes strategies to enhance citizen trust in 5G-enabled environments. A comparative analysis is conducted to identify principles, challenges, and frameworks for embedding trust in the design and deployment of HCAI. We propose a three-tier HCAI trust model comprising three interrelated layers: ethical, functional, and institutional trust. The model is validated statistically using a survey. Findings indicate a growing consensus on the need for AI systems that reflect human values and ensure transparent decision-making. However, significant gaps remain in aligning AI deployment with digital sovereignty and public trust, particularly regarding data control, accountability, user consent, and governance within 5G infrastructure. HCAI must be embedded into policy, design, and technology to build trustworthy smart cities.
A Structured Examination of Artificial Intelligence (AI) Adoption in Public Procurement: Determinants Across Technological, Organisational, and Environmental dimensions
ABSTRACT. Public procurement plays a pivotal role in public service delivery and strategic governance. Growing demands for its transparency, efficiency, and responsiveness have encouraged the exploration of AI in the public procurement process. AI offers potential benefits, including improved efficiency, enhanced effectiveness, greater automation in transactions, and streamlined supplier management. Despite these benefits, its adoption in public procurement is in its early stages. Existing research has largely focused on the discussion around the potential of AI in procurement processes, whereas limited attention has been paid to the determinants of AI adoption in public procurement. This paper conducts a Systematic Literature Review (SLR) to identify the determinants influencing AI adoption in public procurement. The findings are analysed using the Technology-Organisation-Environment (TOE) framework. The results reveal that organisational determinants vary across the procurement process, while the TOE dimensions differ across stages of the adoption process. Public value also emerges as an important dimension.
Artificial Intelligence in Distributed Energy Management: A Model for Optimizing Modern Power Systems
ABSTRACT. The transformation of power systems, driven by decarbonization and the rapid growth of distributed energy resources (DER), requires new methods of coordination, flexibility, and real-time control. The aim of this study is to develop and analyze an integrated distributed energy management model that combines hierarchical dynamic optimization with artificial intelligence methods to enhance system efficiency, stability, and adaptability. The paper presents a systematic literature review (2009–2025), supported by NLP and text-mining techniques, demonstrating the evolution of concepts from communication-oriented solutions to advanced DERMS platforms. Based on these findings, a research model is proposed, built on a network graph and power balance equations, extended with machine-learning predictors for renewable energy generation and demand, AI-supported state estimation, and reinforcement learning for adaptive control under uncertainty. The proposed approach enables the development of digital twin representations while preserving physical constraints and ensuring operational system security.
Architecting the Digital Product Passport in HVAC&R
ABSTRACT. The digital product passport (DPP) is a key pillar of the European Union’s circular economy action plan, enhancing access to sustainability information and promoting more responsible consumption. However, the challenges for inter-organizational data sharing are significant. This paper proposes a sector-wide information system architecture for Heating, Ventilation, Air Conditioning, and Refrigeration (HVAC&R), which is part of the electronics supply chain, a priority for DPP adoption. The results emerge from a systematic literature review and a field study made in collaboration with a leading HVAC&R business association. The findings show the (1) ecosystem interactions, (2) technologies, and (3) end-to-end interoperability requirements for HVAC&R DPPs. Moreover, the proposed architecture aims to design a DPP data lake for HVAC&R, thereby increasing the value of data-driven product improvement initiatives. Our work confirms that DPPs should be designed at the ecosystem level to ensure efficient data flows from the early stages of production.
Multi-Criteria AHP Model for Assessing ML Viability in Business Processes
ABSTRACT. This paper examines when the deployment of machine learning (ML) in a business process is economically justified and when alternative automation approaches may be more appropriate. Although ML adoption is widely discussed in technical and implementation-oriented literature, less attention has been paid to compact decision models for assessing ML suitability in specific business-process contexts. To address this gap, the paper proposes a multi-criteria assessment framework that combines fifteen literature-derived decision criteria with the Analytic Hierarchy Process (AHP). The criteria were weighted by three domain experts and applied to three real-world business processes: tutoring settlement, matching bank-statement transactions to customers, and conducting classes for students. The results show that the framework differentiates effectively among processes with different levels of ML suitability. The transaction-matching process achieved the highest score, the tutoring-settlement process was conditionally justified, and the process of conducting classes for students was not justified for ML deployment. These findings indicate the practical usefulness of a systematic, criteria-driven approach to ML deployment decisions.
Improving session-based recommender systems using popularity-based data augmentation
ABSTRACT. Recommender systems have become an effective approach to address the difficulties users face when interacting with online services, due to the vast amounts of data available. These systems provide customised content tailored to the individual preferences of each user.
While the literature introduces various new models of recommenders, there is a growing area of research focused on additional factors that can improve recommendation quality. One such factor is data augmentation, a pre-processing technique that can influence overall performance.
The aim of this paper is to present data augmentation techniques based on data popularity measures, and to compare them with baseline and advanced solutions. Furthermore, it analyses the influence of input data size on effectiveness and utility of the augmentation process.
Non-uniform Hamming metric learning and classification of categorical data
ABSTRACT. The non-uniform Hamming metric (dissimilarity function) is an extension of the Hamming metric where distances depend not only on the number of different features, but also on the concrete pair of values that are different in each feature. In this article a new method for non-uniform Hamming metric learning is developed. The method is based on the assumption that the data are divided into several classes that are formed according to the metric and minimizing the total inner-class squared distance. Numerical experiments confirm that the method recovers approximately the metric from the data. Moreover, using of the recovered metric can improve knn classification of categorical data with the Hamming metric and give competitive performance against other classical classifiers.
Automated Neural Structure Adaptation for Time Series Classification
ABSTRACT. This paper presents a novel neural time series classifier that operates on symbolic data representation. The proposed approach extends an embedding-based model called SAFE (Simple And Fast segmented word Embedding) by introducing a dynamic variant in which the neural network architecture changes during training. The design of the method utilizes a learning scheme called GrowingNN, which adapts neural architectures through Monte Carlo Tree Search (MCTS). In the proposed hybrid model, SAFE converts time series into symbolic words and maps them to dense embeddings. Then, GrowingNN classifies the embedding matrices, starting from a small network and dynamically growing or shrinking it during training using MCTS-guided structural modifications. On 14 benchmark datasets, the combined approach achieves results comparable to ResNet and standalone SAFE. Empirical experiments show that the new method constructs very compact models tailored to the dataset, where the smallest network has only 1545 trainable parameters it on average three times smaller than ROCKET and 72 times smaller than that of ResNet.
Hybrid Outlier Detection via Majority-Vote Pseudo-Labeling
ABSTRACT. Data quality remains a critical challenge in data-centric information systems, particularly in the presence of high-dimensional data, class imbalance, and limited labeled data. Undetected outliers may propagate through analytical pipelines and reduce the reliability of downstream decision-support services. This paper proposes a hybrid outlier detection framework that combines heterogeneous unsupervised detectors through decision-level majority voting to generate pseudo-labels, followed by supervised refinement in a reduced-dimensional feature space. The framework investigates how detector complementarity affects pseudo-label reliability and downstream model performance. Experiments on ADBench datasets show that detector diversity and partial error decorrelation improve detection robustness. The proposed approach outperforms standalone detectors and identifies conditions under which supervised refinement improves or degrades detection quality. By framing outlier detection as a modular validation layer, the framework supports the development of more reliable and deployable data-centric information systems.
Context-Aware Rule-Based Outlier Detection in Non-Stationary Time Series
ABSTRACT. Outlier detection in non-stationary time series is
challenging due to regime changes and strong contextual
dependence of system behavior. Observations numerically
typical in a global sense may violate structural
relationships that hold only within specific operating
regimes.
This paper proposes a context-aware rule-based framework
for outlier detection in complex time series. System
behavior is partitioned into deterministic contexts
representing volatility and trend regimes; for each
context, a set of interpretable IF--THEN rules describes
regular structural relationships. Outliers are identified
as observations that violate rules of the active context
or exhibit inter-context inconsistency---strong
disagreement between active and alternative context
rule sets.
The method is fully deterministic, requires no iterative
training or probabilistic density estimation, and
preserves interpretability at both rule and context
levels. Rule ranges are calibrated via non-parametric
percentile estimation with linear O(N) complexity.
Evaluation on three daily currency pairs (EUR/USD,
GBP/USD, USD/JPY, 2010--2022) under leave-one-year-out
cross-validation and on a synthetic nonlinear system
confirms that context conditioning increases sensitivity
to structural regime transitions while maintaining
stability compared to global or fully local approaches.
Efficient Compression of Deep Neural Networks via Structural Sparsification
ABSTRACT. Deep neural networks are widely used in data processing and analysis due to their high performance. However, a major drawback of such models is their large size, which imposes substantial memory requirements on storage, transmission, and computational systems. These demands can be significantly reduced through structural sparsification, including the approximation of convolutional kernels and the use of sparse neural layers. In this paper, we propose a unified approach that combines these techniques to compress a two-stage deep neural model composed of convolutional and fully connected networks, as commonly used in image classification tasks. The aim is to experimentally evaluate the extent to which this combination enables model compression while maintaining acceptable classification performance. Extensive experiments demonstrate that the proposed method can reduce the size of the model by up to 14 times, while improving the classification accuracy by approximately 1.2%, likely due to its regularization effect.
Novel spatial meta-feature weighted consensus clustering of geospatial data
ABSTRACT. Accurate identification of geospatial data has a fundamental impact on research and management of wildlife and animal behaviour. Consensus clustering literature neglects the fact that different algorithms are better suited to different dataset geometries, and this can be assessed prior to clustering. In this paper we propose the Spatial Meta-Feature Weighted (SMFW) Consensus Clustering algorithm, which improves the classic EAC for spatial animal data by replacing equal weighting with multi-algorithm adaptive weighting based on spatial meta-features, and by determining weights predictively without a training phase. In addition, this approach made it possible to remove the assumption regarding the number of clusters required by the methods used. Experiments on real and synthetic datasets for the primary ARI metric showed an average improvement of 22.5% compared to traditional approaches. The study's findings demonstrate that the proposed method is an effective tool for geospatial data clustering.
The Impact of Cloud Computing on Selected Dimensions of Enterprise Operations
ABSTRACT. This article analyzes the impact of cloud computing on small and medium-sized enterprises in Poland and identifies key effects. Empirical research, based on a CAWI survey of 410 enterprises, correlation analysis, and exploratory factor analysis (EFA), shows that cloud computing supports the strategic, operational, and information-decision dimensions of company operations. The greatest impact occurs in knowledge management, accelerating decision-making, improving information flow, and supporting innovation. Factor analysis grouped the 15 effects into three main dimensions: strategic, knowledge management, and operational, making CC's impact easier to interpret. The results emphasize that cloud implementation can increase competitiveness, operational efficiency, and adaptability, making it a crucial element of digital transformation.
Integration of Augmentation and Generative Models for Enhanced Gait-Based Authentication
ABSTRACT. This study investigates biometric gait identification using accelerometer and gyroscope signals.
The analysis is based on experiments conducted on the BUT gait database comprising 100 subjects
and a publicly available Signet dataset with recordings from 29 individuals. The selection
of data sets was based on the fact that the same group of participants was recorded on two separate
data collection days, which enabled cross-day validation. We focused on creating a two-step
method for generating artificial gait samples. For the pipeline involving data augmentation, we
decided to add a generative component. The incorporation of the generative model component
resulted in a substantial improvement in performance. For the BUT dataset, the F1-score increased
from 0.802 (augmentation only) to 0.891, whereas for the Signet dataset it improved
from 0.774 to 0.839.
Evaluating blockchain technology adoption opportunities for enhanced traceability and transparency in supply chain management systems engineering
ABSTRACT. Global supply chains are increasingly exposed to disruptions resulting from pandemics, geopolitical conflicts, and climate change, reinforcing the need for resilient and transparent digital supply chain solutions. The article explores the role of blockchain technology in supply chain management (SCM) systems engineering, focusing on its potential to enhance transparency and traceability in SCM. A two-stage qualitative research design was used, involving structured in-depth interviews with SCM and blockchain experts from several countries. The results show that while blockchain is perceived as having moderate and increasing relevance for addressing two key digital supply chain (DSC) challenges, it is not yet a primary design choice in SCM systems engineering due to limited maturity, integration complexity, and lack of dedicated solutions. The study highlights a strong need for engineered, interoperable blockchain-based architectures to support traceability of products and components, as well as transparency in DSC, especially in mature SCM systems. Microsoft Blockchain and VeChain sub-technologies and solutions are recognized by experts as having the highest potential for this. Moreover, there exists a strong mutual interpenetration of tracing and transparency aspects in SCM systems, where their separation is undesirable and uncommon.
Engaging with AI: A Delphi study to prioritise an IS research agenda
1. Overview of the Roadmap project 2. Review of the SLR 3. Introduction to the Delphi study 4. Review of the Delphi survey 4.1. Structure 4.2. Content ( issues, topics) 4.3. What’s missing 4.4. Methodology (iterations, saturation, venues, analysis, etc) 5. Summary and conclusions
Interpretable Relational kNN for Gene Expression Classification: Compact Global and Local Gene-Pair Explanations
ABSTRACT. High-dimensional gene expression data are difficult to analyze with value-based models because absolute expression levels are sensitive to preprocessing and technical variation. Pair-based methods such as TSP and kTSP improve interpretability by using within-sample gene orderings, but they rely on fixed global pair panels.
We investigate whether relational $k$-nearest neighbors (kNN) can provide a useful balance between neighborhood-based flexibility and gene-pair interpretability. Restricted relational representations are constructed from globally informative gene pairs, while individual predictions are explained through compact sets of locally active relations.
The approach is evaluated on seven public gene expression datasets. An ablation study compares Kendall-style and Footrule-based relational variants and is used to select a compact operating point. The selected restricted model is then compared with TSP, kTSP, and full Euclidean and RRM-kNN baselines. The results show the expected trade-off: restricting the relational representation reduces performance relative to full RRM-kNN, but yields compact, sample-specific explanations while remaining competitive with classical pair-based classifiers.
Overall, restricted relational kNN occupies an intermediate position between fixed pair-rule classifiers and full neighborhood-based models, linking global gene-pair selection with local sample-level explanations.
Vision-Assisted Multi-Agent Decision Support for Production Disruptions and Machine-Failure Risk in Manufacturing
ABSTRACT. This paper presents the architecture of a decision support system for a production environment that combines image analysis, adverse-event risk prediction, and multi-agent coordination. The study was conducted on data from a company producing precision aluminum valve bodies and components for the automotive sector. Images from station cameras, process signals, maintenance logs and planning information from MES/ERP were integrated. The proposed system uses a vision module to detect symptoms of degradation, a fusion model to estimate the risk of an adverse maintenance event within an 8-hour horizon, and an agent layer responsible for recommending maintenance activities and schedule changes. On the held-out test set, the best variant achieved F1 = 0.91 and AUROC = 0.96 and, in replay-based evaluation on historical production weeks, was associated with a 40.9% reduction in average weekly downtime relative to the reactive baseline. The results indicate that, in replay-based evaluation, the combination of visual perception and agent-based coordination is associated with higher predictive performance and better operational decision support than the compared baselines. Operational effects were estimated in a replay-based evaluation of historical production weeks rather than in a live online deployment.
A multi-objective parameter-efficient fine-tuning method for medical Polish language automatic speech transcription systems
ABSTRACT. Efficient fine-tuning of an automatic speech recognition system is a key challenge in machine learning research. A major obstacle is accurately recognizing jargon-heavy, domain-specific speech in expert discussions, such as those among lawyers, doctors, or engineers. We propose an automated machine-learning approach that fine-tunes the Whisper small model to enhance medical speech recognition. This fine-tuning uses three selected PEFT methods (LoRA, DoRA, and MoRA) and employs a multi-objective optimization framework, enabling a balance between the accuracy of medical speech recognition (evaluated on the ADMEDVOICE test subset) and general Polish speech recognition (evaluated on the Common Voice test subset). The set of Pareto-optimal solutions was found to be statistically significantly better than those obtained from randomly chosen configurations and to provide better general language performance than most solutions generated by the single-objective optimization process. Solutions on this front are competitive with state-of-the-art systems for speech recognition in the medical domain.
Robust Low-Resource Dysarthric Speech Recognition via Low-Rank Adaptation
ABSTRACT. This study investigates transformer fine-tuning strategies for low-resource dysarthric speech recognition across four languages (English, Polish, Italian, and Dutch). We compare full-parameter fine-tuning, Low-Rank Adaptation (LoRA), and a QK-MLP ablation that updates the same layers as LoRA but uses full-rank parameter updates, alongside regularisation techniques such as dropout and SpecAugment. LoRA outperforms full-parameter fine-tuning on five of six test sets, achieving new state-of-the-art word error rates (WER) on COPAS (23.86 vs. 29.0), EasyCall (17.15 vs. 39.2), and PLDD (69.59 vs. 71.91), and providing the first large-scale ASR evaluation on GidoLab. A direct comparison with the QK-MLP variant confirms that LoRA's advantage arises from its low-rank constraint acting as implicit regularisation, rather than simply from the selection of updated layers. Severity-stratified evaluation further shows that LoRA reduces cross-severity WER spread, indicating improved robustness to dysarthria severity variation and suggesting that the low-rank constraint prevents severity-specific overfitting.
Interactive Geospatial Mapping of Photovoltaic Energy Yield Using a Learned Weather-to-Production Model
ABSTRACT. This paper presents an interactive pipeline for geospatial mapping of photovoltaic (PV) energy yield, transforming a trained weather-to-production model into a spatial decision tool. The approach follows two stages. First, a predictive model is trained on reference PV data using selected meteorological features and consistent preprocessing (imputation and scaling). Second, the model is deployed in a mapping module that retrieves daily weather data for a chosen location and surrounding grid from the NASA POWER API, applies the saved preprocessing, and estimates PV yield for each point. Results are visualized as a heatmap, with point-level predictions and basic economic indicators based on user-defined system capacity and energy price. A~case study shows that the “train once, infer across space” approach enables fast and reproducible assessment of PV yield variability driven by weather conditions, while noting limitations such as site-specific effects and the need for multi-site calibration.
Neural Additive Model based framework for interpretable glaucoma screening
ABSTRACT. Glaucoma is a chronic, progressive eye disease projected to affect 111 million people by 2040, which has motivated numerous automatic detection methods, most of them relying on black-box models with limited interpretability. We introduce an inherently interpretable framework for glaucoma screening that extracts 20 clinically relevant concepts: cup-to-disc ratios, ISNT-sector areas, shape descriptors, and regional color statistics, from color fundus images and feeds them to a Neural Additive Model in which each concept is handled by a dedicated subnetwork. Despite containing only 2,587 trainable parameters, the model rivals far larger deep neural networks: on the Rim-One DL benchmark, it reaches an accuracy of 0.901, an F1 score of 0.855, and an AUROC of 0.945 on the by-Random split, and remains competitive under domain shift on the by-Hospital split (accuracy 0.851, F1 0.774, AUROC 0.894), while outperforming logistic-regression and gradient-boosting baselines trained on the same concepts. Crucially, every prediction is accompanied by intrinsic, per-concept local explanations that remain clinically meaningful, obtained without any post hoc explanation method, thereby supporting the transparency requirements placed on medical AI and contributing to greater trust in AI systems in medicine.