If you ve been training Vision-Language-Action (VLA) or world-action models, you ve probably noticed that the training infrastructure feels a bit unoptimized compared to the mature stack we have for LLMs. Most researchers end up relying on each model's official reference repository, which is great for reproducibility, but throughput was rarely the main priority.
LoongForge is a modular, scalable, high-performance training framework for LLMs, VLMs, diffusion, and embodied models. Up to ~5× training speedup, native NVIDIA GPU and Kunlun XPU support.