Hello, I'm

Mehdi Jafarnia

Research Engineer · Google DeepMind

Ph.D. in Electrical Engineering · Reinforcement Learning & Foundation Models

About Me

Mehdi Jafarnia

I am a Research Engineer at Google DeepMind in Mountain View, CA. My work centers on reinforcement learning, post-training, and foundation models. I investigate novel uncertainty estimation and exploration paradigms for Large Language Models (LLMs), designing methods that improve model reasoning, data efficiency, and reliability in complex environments.

Prior to Google DeepMind, I spent a year at Uber Technologies, where I worked on dynamic pricing and cost estimation models.

I received my Ph.D. in Electrical Engineering from the University of Southern California (USC), alongside concurrent Master's degrees in Computer Science and Applied Mathematics. Prior to USC, I completed dual B.S. degrees in Electrical Engineering and Computer Science at Sharif University of Technology.

Experience

RL, Exploration & Post-Training
Jul 2022 — Present Mountain View, CA

Google DeepMind

Research Engineer

My work focuses on advancing the core capabilities of large language models through reinforcement learning and post-training. I develop novel exploration algorithms to dramatically improve data efficiency, applying these online RL techniques across complex downstream domains such as code generation and reasoning.

Dynamic Pricing
Oct 2021 — Jul 2022 San Francisco, CA

Uber Technologies Inc.

Software Engineer

Worked on the Dynamic Pricing team, developing real-time trip cost prediction and marketplace pricing algorithms across high-throughput global traffic.

Selected Publications

2026
Efficient Exploration at Scale
M. Asghari, C. Chute, V. Dwaracherla, X. Lu, M. Jafarnia, Z. Wen, B. Van Roy
2024
A Bayesian Learning Algorithm for Unknown Zero-Sum Stochastic Games with an Arbitrary Opponent
M. Jafarnia, R. Jain, A. Nayyar
2023
Posterior Sampling-Based Online Learning for the Stochastic Shortest Path Model
M. Jafarnia, L. Chen, R. Jain, H. Luo
2022
Online Learning for Unknown Partially Observable MDPs
M. Jafarnia, R. Jain, A. Nayyar
2021
Implicit Finite-Horizon Approximation and Efficient Optimal Algorithms for Stochastic Shortest Path
L. Chen, M. Jafarnia, R. Jain, H. Luo
2020
Model-Free Reinforcement Learning in Infinite-Horizon Average-Reward Markov Decision Processes
C. Wei, M. Jafarnia, H. Luo, H. Sharma, R. Jain

Get in Touch