Light Origins has released a preview of Light-O1, a six-billion-parameter reasoning text-to-action model intended to generate whole-body humanoid motion from natural-language instructions.
The preview includes weights and inference materials
The Hugging Face model card lists BF16 weights, the tokenizer, processor configuration, a 22-joint shared action representation and an action decoder. The preview is licensed under Apache 2.0, with inference code linked from the public Light-O1 repository.
Human-video pretraining is the core research claim
Light Origins says it trained a series of four-billion-parameter base models across six budgets from 3.75 billion to 120 billion multimodal tokens, with the largest corresponding to 100,000 hours of recovered human action. It reports power-law reductions in prediction loss and pose error after adaptation across human and robot datasets.
Benchmarks remain developer-evaluated
The company reports a 79.3% macro success rate on 24 simulated RoboCasa GR-1 tasks and stronger motion-generation ratings than cited baselines. RuntimeWire independently confirms the release and public artifacts, but the benchmark results, scaling-law fits and real-world demonstrations have not been independently reproduced.
