Video carries culture, ideas, and connection around the world. Video is a universal language. Unlike text, making video is a physical act, and editing high-dimensional video data is genuinely hard. So the few who hold the means of production create, and the rest of us consume. They tell the stories and we all listen.
Video can be the most effective medium of visual storytelling. Video — if it were as universally manipulable as text — could be the most effective medium of thought.
To get there we have been training Vela, a family of generative models that produce and manipulate video. Adoption of Vela 1 outran everything we projected, and in directions we did not predict. Which tells us we should go faster.
The Orrery API
We are launching the Orrery API in beta today with the newest family of Vela 1.6 models. It covers high-quality text-to-video, image-to-video, video extension, loop creation, and our camera-control capabilities.
Through our research and engineering we aim to bring about an age of visual literacy in which anyone can make the thing they can picture, and the distance between an idea and an artifact keeps shrinking.