Ongoing · Upadhyay Lab, Dept. of CSIS, BITS Pilani Hyderabad

Reinforcement Learning for Adaptive Mutational Pathway Modelling

Danny Muzata • Dr. Prajna D. Upadhyay

Reinforcement Learning Evolutionary Dynamics Mutational Pathways Machine Learning

Overview

Adaptive evolution under selection can be framed as a sequential decision process: at each step, a population "chooses" among available mutations, and the choice that maximises long-term fitness is not always the one that looks best immediately — a landscape can be rugged, with fitness valleys that must be crossed to reach a higher peak.

This ongoing project, carried out in collaboration with the Upadhyay Lab (Dept. of CSIS), explores reinforcement learning formulations of this problem — treating genotypes as states, mutations as actions, and fitness as reward — to model and predict adaptive mutational trajectories, complementing the fitness-landscape work on PfDHFR and DHPS carried out in the Chowdhury Lab.

Key Directions

  • Formulating protein sequence space as a Markov decision process over mutations
  • Learning policies that reproduce empirically observed adaptive walks
  • Connecting RL-predicted trajectories to epistatic fitness-landscape models
  • Presented in part at the Graduate Forum, IndoML 2025

Researcher

Portrait of Danny Muzata

Led by , a biomedical data scientist at BITS Pilani, Hyderabad Campus, in collaboration with Dr. Prajna D. Upadhyay's group in the Dept. of CSIS.

Google Scholar  •  ORCID