The competition is going up a notch for generative AI models applied to video. On February 15, OpenAI released Sora, a new text-to-video model that takes a step forward in rendering quality. Has OpenAI won the war? Absolutely not, according to the CEO of the start-up Runway, Cristóbal Valenzuela.
The latter, in the process, wrote on X (ex-Twitter) “game on” in response to the various video demonstrations published by OpenAI with Sora. Runway is developing its own video generation model, the second version of which (called Gen-2) can process instructions up to 320 characters long and costs $0.05 per second of generated video.
If you can imagine it, you can generate it.
Gen-2 is now available on web and mobile: pic.twitter.com/tbo1t7JGeJ
— Runway (@runwayml) June 7, 2023
Runway claims its tool is already used by companies like Google, Publicis Group, Microsoft, New Balance, Nvidiaor even Vox.
Meta and Google, pioneers in video generation
OpenAI and Runway are obviously not the only ones in this market. Adobe, Google, Meta and Nvidia are also in the running. It was Meta who first launched its “Make-A-Video” tool at the end of September 2022. It is capable of transforming a few words or lines of text into a short video. The system can also create videos from images or take existing videos and create new, similar ones.
A few days later, it was Google's turn to present its Imagen Video solution. This announcement followed the presentation of Imagen (a solution for transforming text into images). Imagen Video can produce videos with a resolution of 1280 x 768 pixels at 24 frames per second.
Adobe and Nvidia are also in the race
In April 2023, other actors enter the dance. The Toronto AI laboratory of graphics card giant Nvidia unveils VideoLDM, capable of generating temporally coherent videos lasting a few seconds at a resolution of 1280 x 2048 pixels. The firm takes particular advantage of a high-resolution text-to-video synthesis model based on the open source Stable Diffusion model from Stability AI.
In parallel, Nvidia researchers also train prediction models to enable the generation of temporally consistent long videos lasting several minutes. Their resolution is, however, much less impressive, reduced to 512 x 1024 pixels. For its part, Adobe launched Firefly, a family of generative artificial intelligence models. Developed in partnership with Nvidia, it allows you to quickly create and modify images using natural language instructions. Premiere Pro, its editing software, should soon benefit from AI functions, including the integration of text video editing.
A feature purported to allow users to trim and rearrange video based on automatically detected transcriptions of lyrics extracted from video clips. “With just one button, we can generate 1,000 versions of the same video that are localized,” told Reuters Ivo Manolov, Adobe vice president for digital audio and video enterprise offerings. The advertising industry is targeted as a priority.
An accelerating pace of innovation
At the end of January, Google made headlines again with Lightwhich the firm describes as “a spatio-temporal diffusion model for video generation”. Based on a single reference image, Lumiere can generate videos in the target style using refined text-image model weights, it reads. It stands out for its Space-Time U-Net architecture which generates the entire temporal duration of the video at once, via a single pass through the model. This approach makes it possible to generate 80 images at 16 frames per second, specify the researchers behind the model.
Stability AI also took the plunge and presented last week SVD 1.1a delivery model for AI videos “more coherent”. And if we follow Cristóbal Valenzuela's reasoning, advances in artificial intelligence-generated videos could accelerate. “A year's progress is now happening in months. Months of progress will soon start happening in days. Days of progress will soon start happening in hours.”
Selected for you