Sample-Efficient Reinforcement Learning
The core of my thesis: making RL converge faster by exploiting structure that is already available — teacher signals, prior knowledge, planner guidance, or inductive bias. Teacher-assisted exploration and pseudo-label-driven actor-critic acceleration cut convergence time substantially without changing the underlying algorithm class.
TA-Explore (AAMAS 2023)
Accelerated Actor-Critic (Inf. Sciences 2024)
Human-Inspired RL (J. Supercomputing 2025)
NARS vs. Q-learning (AGI 2023)
BLADE-TD (under review)Rule-Based Grid World (under review)
Federated & Distributed Optimization
Convergence guarantees where the convenient assumptions fail — without data similarity, under compressed and clipped gradients, under model mismatch between clients, and under adversarial participants. Theory paired with reference implementations, extending to distributed inference and resource-aware serving in edge–cloud settings.
FL without Data Similarity (IEEE TBD 2024)
Distributed SGDM (IEEE TCNS 2025)
CompFedRL (ECML-PKDD 2024)
FedTD(0) (ECML-PKDD 2025)
Trust-Aware Decentralized FL (preprint 2026)
Collaborative Distributed Inference (preprint 2026)
Generative Markov Model (preprint 2026)
Reinforcement Learning for Control & Cyber-Physical Systems
Learning-based control where safety, partial observability and physical constraints matter: multi-agent deep RL for smart power converters with Hitachi Energy and KTH, model-predictive-control-guided RL for power electronics, and adaptive traffic signal control under noisy, nonstationary conditions.
PriPG-RL (preprint 2026)
Active Inference for Traffic Control (IEEE WF-IoT 2026)
Energy-Efficient Cloud Data Centres (Trans. Machine Intelligence 2020)
MPC-Guided Fast RL (under review)
Agentic & LLM-Based Systems
How to make systems built on large language models trustworthy enough to deploy: auditing fault containment in LLM-driven agents, grounding generation in retrieved evidence for high-stakes scientific prediction, sample-efficient test-time scaling, and the socio-technical questions agentic AI raises. This is also the research strand closest to my industry work.
Socio-Technical Agentic AI (preprint 2025)
GICA (TMLR 2026)Fault Containment in LLM–RL Agents (under review)Evidence-Grounded AMP Prediction (under review)
Learning from Imbalanced & Incomplete Data
Systematic study of resampling methods — SMOTE, ADASYN and alternatives — against gradient-boosted and ensemble classifiers across imbalance regimes, alongside cost-sensitive Transformers for industrial prognostics and clustering that tolerates up to 50% missing data. Developed in part with Scania CV AB and Linköping University.
Cost-Sensitive Transformer (Cluster Computing 2026)
RF & XGBoost under Imbalance (Technologies 2025)
Churn Prediction Review (MLKE 2025)
SMOTE & ADASYN in Telecom Churn (ICWR 2024)
Computer Vision & Multimodal Learning
Plant and leaf identification with deep CNNs and Vision Transformers, automatic multimodal fusion, and semantic segmentation for autonomous driving — several strands developed together with the Master's students I supervise, and several leading to joint publications.
SWP-LeafNET (ESWA 2022)
Multimodal Plant ID (Frontiers Plant Sci. 2025)
Human Activity Recognition (ICSPIS 2019)
Leaf Transfer Learning (ICSPIS 2018)
Autonomous Driving Segmentation (CSICC 2021)
Hierarchical Kannada-MNIST (CSICC 2021)
ML for PEDOT Conductivity (Phys. Rev. Materials 2025)