About Me
I am a Postdoctoral Researcher at MIT CSAIL, working with Prof. Omar Khattab. My current research focuses on developing the foundations of Pedagogical Reinforcement Learning for AI agents—learning paradigms that enable agents to improve through interaction, feedback, and long-horizon experience rather than relying solely on scalar terminal rewards. My broader vision is to build continually learning AI agents that can autonomously acquire new capabilities while remaining robust, efficient, and aligned with human objectives.
I received my Ph.D. in Computer Science from the University of Maryland, College Park, where I was fortunate to be advised by Prof. Furong Huang and Prof. Dinesh Manocha. During my Ph.D., I had the opportunity to collaborate closely with Prof. Mengdi Wang (Princeton University), Prof. Amrit Singh Bedi (University of Central Florida), Dr.Mohammad Ghavamzadeh (Qualcomm AI Research) and Prof. Aldo Pacchiano (Boston University), working on reinforcement learning, AI alignment, optimization, and decision-making under uncertainty.
I have also spent time at Google Research, where I worked with Alekh Agarwal, Rahul Kidambi, and collaborators on reinforcement learning algorithms for long-horizon sequential decision making, with a particular emphasis on exploration, credit assignment, and efficient learning. Prior to my Ph.D., I worked as a Research Scientist at Walmart Global Tech, where I developed machine learning methods for large-scale recommendation and optimization systems.
Recent News
- July 2026: Joined MIT CSAIL as a Postdoctoral Researcher working on Foresight RL for Agents
- May 2026: Released our recent work on Pedagogical RL [Blog], [Detailed Thread]
- May 2026: Completed my Ph.D. in Computer Science at the University of Maryland, College Park with Prof. Furong Huang and Prof. Dinesh Manocha [Details]
- May 2026: Received the UMD Computer Science Certificate of Outstanding Achievement [Details]
- 2025: Our work COLLAB studied Test-time Scaling with Mixture of Agents appeared at ICLR 2025 [Paper]
- 2025: Our work IMMUNE studied robustness against jailbreak attacks and appeared at CVPR 2025 [Paper]
- 2024: Our work on inference-time alignment, Transfer Q*, appeared at NeurIPS 2024 [Paper]
- 2024: Our work on pluralistic alignment, MaxMin-RLHF, appeared at ICML 2024 [Paper]
- 2024: Our work PARL on distribution shift in RLHF appeared at ICLR 2024 [Paper]
Selected Publications
IMMUNE: Provable Safety Alignment Against Jailbreaks for Multi-modal LLMs
Souradip Chakraborty, et al.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025).
[Paper] [Detailed Thread]COLLAB: Test-time Scaling with Mixture of Agents
Souradip Chakraborty, et al.
International Conference on Learning Representations (ICLR 2025).
[Paper] [Detailed Thread]Transfer Q*: Provably Efficient Test-time Scaling and Personalization for LLM Alignment
Souradip Chakraborty, et al.
Advances in Neural Information Processing Systems (NeurIPS 2024).
[Paper] [Detailed Thread]MaxMin-RLHF: Toward Equitable Alignment of Large Language Models with Diverse Human Preferences
Souradip Chakraborty, et al.
International Conference on Machine Learning (ICML 2024).
[Paper] [Detailed Thread]PARL: Principled Online Reinforcement Learning from Human Feedback
Souradip Chakraborty, et al.
International Conference on Learning Representations (ICLR 2024).
[Paper] [Extended Version] [Tweet]STEERING: Stein Information Directed Exploration for Model-Based Reinforcement Learning
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Mengdi Wang, Furong Huang, Dinesh Manocha.
International Conference on Machine Learning (ICML 2023).
[Paper]Posterior Coreset Construction with Kernelized Stein Discrepancy for Model-Based Reinforcement Learning
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Brian M. Sadler, Furong Huang, Pratap Tokekar, Dinesh Manocha.
AAAI Conference on Artificial Intelligence (AAAI 2023, Oral).
[Paper]Dealing with Sparse Rewards in Continuous Control Robotics via Heavy-Tailed Policy Optimization
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Pratap Tokekar, Dinesh Manocha.
IEEE International Conference on Robotics and Automation (ICRA 2023).
[Paper]
For a complete publication list, see my Google Scholar.
Selected Patents
- Gregory Dixon, Souradip Chakraborty, Ojaswini Chhabra, Mallikharjuna Mv: Reverse Engineering Food Ingredient Share estimation using Constrained Optimization, US Patent, Walmart Ref. 6031US01.
- Souradip Chakraborty, Abhishek Mishra, Somedip Karmakar: Systems and methods for Unsupervised image processing, US Patent, Ref. US11688049B2.
- Pranay Dugar, Souradip Chakraborty: Automated planogram anomaly detection with Computer Vision, US Patent, Ref. US11669843B2.
- Souradip Chakraborty, Mani Garlapati: Systems and methods for identifying negotiable items, US Patent, Walmart Ref. 5928US01.
- Souradip Chakraborty, Rajesh Shreedhar Bhat, Mani Garlapati System and Method For Automated Electronic Catalogue Management and Image Quality Assessment, US Patent, Walmart Ref. 5118US01.
Broader Impact & Open-Source Contributions
- Served as one of the student organizers of summer AI camps at UMD Fall’23 as a part of the Iribe Initiative for Inclusion and Diversity with the AI4ALL and Maryland Center for Women in Computing
- Outstanding Reviewer Award (Travel award) at Neurips 2022 and Neurips 2023 (consecutive 2years in a row)
- Outstanding Reviewer Award (Travel award) at AISTATS 2023
- Recognized by Google as a Google Developer Expert in Machine Learning for my open-source contributions & mentorships in the field of AI and ML.
- AI vs COVID-19 BioMedical Research Initiaive with Google Research : Developed BioMedBERT for biomedical researchers, doctors, and virologists, to augment their ability to sift through biomedical knowledge and existing research to extract novel insights and help them make new drug discoveries, COLING’2019
Selected Blogs & Articles
- Souradip Chakraborty, Amlan Das, Sai Yashwanth: Risks and Caution on Applying PCA for Supervised Learning Problems, published on Towards Data Science, Medium, 2019.
- Souradip Chakraborty, Rajesh Shreedhar Bhat: Why Not Mean Squared Error (MSE) as a Loss Function for Logistic Regression?, published on Towards Data Science, Medium, 2019. Trending in Machine Learning Category. (> 53K views)
- Souradip Chakraborty: Dimensionality Reduction in Supervised Framework and Partial Least Square Regression
- Souradip Chakraborty An Attempt - Detection of COVID-19 Presence from Chest X-ray Scans Using CNN & Class Activation Maps
- Souradip Chakraborty Tutorial on Bayesian Machine Learning
