Learning Reusable Options by Decomposing Neural Policies
Advances in Neural Information Processing Systems, 2026
Didec, a neural option-discovery method based on policy masking.
Hello, I'm Kiarash
I’m a PhD student in Computing Science at the University of Alberta, advised by Levi Lelis. I work in reinforcement learning, on problems where structure helps an agent generalize: discovering options for hierarchical RL, representing policies as programs, and using LLMs as world models.
Before my PhD I was a support researcher at Huawei in Edmonton for two years, working on image-quality assessment, ISP and compiler optimization. I did my MSc at UAlberta with Martin Müller, where I studied Monte Carlo Tree Search with imperfect models, and my BSc in Computer Engineering at Iran University of Science and Technology.
Advances in Neural Information Processing Systems, 2026
Didec, a neural option-discovery method based on policy masking.
International Conference on Machine Learning, 2026
Re-evaluates prior claims that programmatic policies generalize better than neural policies on out-of-distribution tasks.
AAAI Conference on Artificial Intelligence, 2024 · * equal contribution
UA-MCTS adapts Monte Carlo Tree Search to inaccurate transition models. Based on my MSc thesis.