I work on making LLM training and inference more efficient. To that end, I develop algorithms that improve model quality at low precision and run efficiently on the latest hardware.
More broadly, I study which limits of low-precision learning are fundamental and which can be overcome with better algorithms. Ultimately, the goal is better models for a given hardware and time budget: faster computation lets us train larger models, or the same model on more tokens.
QUASAR treats quantization-aware training as a weight reconstruction problem, continuously fitting a loss-aware quantized representation as the latent full-precision weights evolve. Its analysis connects weight reconstruction quality to the optimization dynamics and final quantized-model loss.
LOOT predicts each token from the rest of the sequence, allowing masked diffusion language models to learn from every token and self-correct during generation.