【AI前沿】xAI upgrades Imagine Video 1.5: New Image and Voice Reference Features, and Native 1080p Video Generation

2026-08-03

AI NEWSLatest AI NewsArticlexAI upgrades Imagine Video 1.5: New Image and Voice Reference Features, and Native 1080p Video GenerationPublished in Latest AI NewsTime :Aug 3, 2026Read :4minutexAI has launched a major upgrade for its video generation model Imagine Video 1.5, adding image reference, voice reference, text-to-video generation, and native 1080p output capabilities, further enhancing role consistency, scene control, and video quality in AI video creation. Currently, the text-to-video and 1080p video generation features are available on the Grok Imagine web version, iOS, and Android apps. The image and voice reference features have been initially opened to SuperGrok Heavy and SuperGrok Plus users in the United States, with plans to expand to more users later.This update provides users with a more flexible video generation process. Creators can directly generate videos through text descriptions, or upload reference images to control characters, products, or scenes. They can also combine character images with voice samples, allowing the AI to maintain consistent appearance and voice across different shots. xAI stated that Imagine Video 1.5 supports up to seven visual references per generation, used to fix facial features of characters, product characteristics, or environmental elements, thus achieving more precise content control.In the multi-reference mode, users can retain character identity while replacing the background, keep the scene unchanged while changing the character, or adjust only the action while maintaining the overall visual setup. The voice reference feature further addresses the issue of character continuity in AI video production, ensuring stable facial features and voice performance for the same character across multiple segments.For developers, the xAI API now supports calling image references, text-to-video generation, and native 1080p video generation capabilities through thegrok-imagine-video-1.5model. Developers can generate videos by passing parameters such as prompts, reference image URLs, video duration, aspect ratio, and resolution through the API. The voice reference feature requires integration via request methods.Imagine Video 1.5 was released last month, and xAI stated that this version optimized motion performance, physical simulation, and audio generation. This upgrade further enhances the model’s control over identity, scenes, and product details, making AI video generation gradually transition from simple image transitions to a multimodal creation tool closer to professional film production processes.