Member of Technical Staff, Gradium
Pretraining and post-training large-scale streaming speech LLMs across architecture, data quality, evaluation, and post-training methodology.
Paris, France
Research Scientist | Multimodal & Speech Foundation Models
I build large-scale multimodal and streaming speech LLMs, from tokenization and pretraining through reinforcement-learning post-training and evaluation. I am currently a Member of Technical Staff at Gradium and was previously a Staff Research Scientist at Meta.
Pretraining and post-training large-scale streaming speech LLMs across architecture, data quality, evaluation, and post-training methodology.
Contributed to speech pretraining for Llama 3; co-led speech tokenization and pretraining design for Llama 4 Speech, and reinforcement-learning post-training for full-duplex speech models. Led post-training for speech-to-speech translation models deployed in Instagram AI dubbing.
Developed large-scale multilingual end-to-end speech recognition models.
Developed multimodal emotion and sentiment recognition systems, including low-latency models using deep reinforcement learning.
Multimodal models and streaming speech LLMs, including tokenization, pretraining, and post-training.
Reinforcement learning, instruction following, model evaluation, naturalness, and expressivity.
Large-scale distributed GPU training, data quality, experimentation, and evaluation pipelines.