ModelScope Text-to-Video Technical Report
Abstract
ModelScopeT2V synthesizes videos from text by integrating spatio-temporal blocks and combining VQGAN, a text encoder, and a denoising UNet to achieve state-of-the-art performance.
This paper introduces ModelScopeT2V, a text-to-video synthesis model that evolves from a text-to-image synthesis model (i.e., Stable Diffusion). ModelScopeT2V incorporates spatio-temporal blocks to ensure consistent frame generation and smooth movement transitions. The model could adapt to varying frame numbers during training and inference, rendering it suitable for both image-text and video-text datasets. ModelScopeT2V brings together three components (i.e., VQGAN, a text encoder, and a denoising UNet), totally comprising 1.7 billion parameters, in which 0.5 billion parameters are dedicated to temporal capabilities. The model demonstrates superior performance over state-of-the-art methods across three evaluation metrics. The code and an online demo are available at https://modelscope.cn/models/damo/text-to-video-synthesis/summary.
Community
মানুষটি উঠে দারাবে
Cinematic product commercial, dark and intense atmosphere. Start in a nearly black environment with dramatic low-key lighting. A premium dark brown wooden table is barely visible through the shadows. The camera begins below the table level and slowly, smoothly rises upward, revealing the surface of the wooden table in an extremely cinematic way.
As the camera reaches the tabletop, three premium wireless earbuds are revealed, carefully positioned with a clean and luxurious composition: Anker Soundcore R50i, Soundcore T13 ANC, and Soundcore T13 ANC 2. Keep the products realistic, accurate to their original designs, with no changes to their shapes, colors, logos, or physical details.
At the exact moment the products are fully revealed, water suddenly begins falling from above, creating dramatic splashes and droplets across the earbuds and wooden table. Use high-detail slow-motion water physics, realistic reflections, cinematic highlights, and visible water droplets in the air.
The camera slowly and smoothly pushes forward toward the products. During the close-up movement, the camera motion must remain slow, controlled, smooth, and premium, gradually increasing the feeling of tension without becoming fast or shaky. Use macro close-ups of water droplets hitting the earbuds, realistic reflections on their surfaces, and dramatic shadows.
The soundtrack should be dark, cinematic, suspenseful, and adrenaline-building. Start with a subtle low bass and atmospheric tension, then gradually build intensity with deep pulses, rising cinematic impacts, and a strong sense of danger and anticipation. No vocals. The music should create the feeling that something powerful and dangerous is about to happen.
Style: ultra-realistic cinematic product advertisement, luxury commercial, dark atmosphere, high contrast, dramatic lighting, volumetric shadows, realistic water simulation, shallow depth of field, slow camera movement, intense suspense, premium technology advertisement, 4K, highly detailed, cinematic color grading.
Important: The camera must rise slowly from the wooden table, reveal all three earbuds, then move slowly closer to them. Keep the product designs completely accurate and recognizable. The scene must feel dark, dangerous, luxurious, intense, and adrenaline-filled.
Get this paper in your agent:
hf papers read 2308.06571 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 17
ali-vilab/i2vgen-xl
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 561
Collections including this paper 0
No Collection including this paper
