Document-to-video
Also known as: doc-to-video, slides to video, PDF to video
Document-to-video is an emerging generative capability where an AI model accepts a file — a PDF report, a PowerPoint deck, a spreadsheet — and outputs a video that interprets and animates the content. It is a meaningful step beyond standard text-to-video: instead of asking a human to translate their document into a written prompt, the model reads the source material directly and decides how to render it visually.
The capability entered builder conversations in August 2026 when Alibaba's Wan3.0 launched in public beta with support for documents including PDFs, slides, spreadsheets, and web pages alongside traditional text, image, and audio inputs. Alibaba calls the feature Omni-Reference. The practical use cases most builders are exploring are marketing content (turning a product brief or investor deck into a short video), internal communications (animating a report), and content at scale (generating many short videos from a structured data source without writing individual prompts for each).
As of mid-2026, document-to-video is early. Output quality varies depending on how dense the source document is, and no model has published benchmarks specifically for this input type. The more important structural shift is what it signals: multimodal input is moving from 'image plus text' toward 'any structured file type,' which changes how builders think about the source material for generation workflows.