Paris, France

Egor Lakomkin

Research Scientist | Multimodal & Speech Foundation Models

I build large-scale multimodal and streaming speech LLMs, from tokenization and pretraining through reinforcement-learning post-training and evaluation. I am currently a Member of Technical Staff at Gradium and was previously a Staff Research Scientist at Meta.

Egor Lakomkin

Experience

Paris

Member of Technical Staff, Gradium

Pretraining and post-training large-scale streaming speech LLMs across architecture, data quality, evaluation, and post-training methodology.

Aachen

Staff Research Scientist, Meta Superintelligence

Contributed to speech pretraining for Llama 3; co-led speech tokenization and pretraining design for Llama 4 Speech, and reinforcement-learning post-training for full-duplex speech models. Led post-training for speech-to-speech translation models deployed in Instagram AI dubbing.

Aachen

Applied Scientist, Amazon Alexa

Developed large-scale multilingual end-to-end speech recognition models.

Hamburg

Doctoral Researcher, University of Hamburg

Developed multimodal emotion and sentiment recognition systems, including low-latency models using deep reinforcement learning.

Selected publications

Google Scholar ↗
  1. Efficient Streaming LLM for Speech Recognition ICASSP, 2025
  2. The Llama 3 Herd of Models arXiv, 2024
  3. AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs NAACL, 2024 · Long Papers
  4. End-to-End Speech Recognition Contextualization with Large Language Models ICASSP, 2024

Research focus

Foundation models

Multimodal models and streaming speech LLMs, including tokenization, pretraining, and post-training.

Learning and evaluation

Reinforcement learning, instruction following, model evaluation, naturalness, and expressivity.

Systems

Large-scale distributed GPU training, data quality, experimentation, and evaluation pipelines.