Stable and Low-Precision Training for Large-Scale Vision-Language Models
First page

Paper summary
Methods for accelerating and stabilizing large VLM training.
Ask this paper
01
Mixed-precision techniques: Introduces stable training strategies for bfloat16/float16 mixed precision of large VLMs.
02
Training speedup: Significantly accelerates VLM training while avoiding common instabilities (loss spikes, NaN).
03
Scale-friendly: Scales to the largest open-source VLMs, enabling more research at serious scale.
04
Infrastructure contribution: Practical infrastructure advances that benefited the entire VLM research community.