The Yandex team has introduced a beta version of the YandexART (Vi) neural network for creating five-second videos. According to the press service, the model has learned to recreate smooth movements of objects in the frame, such as a dog running, a leaf falling from a tree, or a firework explosion.

The neural network can be used by regular users to create unique animated phone screensavers, as well as by bloggers, animators, and other professionals. YandexART (Vi) is already available in the "Shedevroom" app.

The company released a previous version of the model for generating videos based on text descriptions in August of last year. The previous solution allowed for the creation of animations where it appeared that the camera, rather than the object, was moving. Additionally, objects changed significantly from frame to frame during generation.

YandexART (Vi) has learned to recreate realistic movements and consider the relationship between frames, resulting in more cohesive and smoother videos. To enable the neural network to perform this task, it was trained on videos featuring moving objects, such as a car driving or a cat sneaking.

The neural network generates a sequence of frames that seamlessly transition into each other, forming a smooth video. The model takes a text description from the user about what should be in the frame (for example: "A rhinoceros dancing hip-hop in a gloomy forest") and creates an image from which the animation begins. The model then gradually transforms digital noise into a sequence of frames, based on this image and text prompt.