qwencloud-video-generation
Create videos asynchronously using QwenCloud's Wan models across multiple modes: text-to-video, image-to-video, first-and-last-frame transitions, role-play character animation, and video editing. Submit your request and poll for completion—no synchronous waiting required.
qwencloud-video-generation creates videos from text, images, or frame pairs using Wan models with async processing.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-07-22
qwencloud-video-generation creates videos from text, images, or frame pairs using Wan models with async processing. Create videos asynchronously using QwenCloud's Wan models across multiple modes: text-to-video, image-to-video, first-and-last-frame transitions, role-play character animation, and video editing. Submit your request and poll for completion—no synchronous waiting required.
Use it when
- Yes, qwencloud-video-generation supports image-to-video conversion.
- qwencloud-video-generation provides five primary modes: text-to-video (generate from prompts), image-to-video (animate still images).
Verify before relying
Read SKILL.md below before installing (15 files). Open directory: indexed for reading, not audited.
Install
QwenCloud/qwencloud-ai/qwencloud-video-generation · repository language: Python
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
How does qwencloud-video-generation create video from text description?
qwencloud-video-generation uses Wan models to convert text prompts into videos asynchronously. You submit a text description, and the system processes it in the background. Poll the API to check completion status rather than waiting synchronously for results. This approach handles complex prompts efficiently without blocking your application.
Can I make animation from image using this skill?
Yes, qwencloud-video-generation supports image-to-video conversion. You can animate a still image into a video, or use first-and-last-frame mode to create smooth transitions between two frames. The async architecture means you submit your image and check back for the generated video once processing completes.
What video generation modes does qwencloud-video-generation offer?
qwencloud-video-generation provides five primary modes: text-to-video (generate from prompts), image-to-video (animate still images), first-and-last-frame transitions, character role-play animation, and video editing with advanced functions. Each mode leverages Wan models and operates asynchronously for reliable batch processing.
Is qwencloud-video-generation available as an API?
Yes, qwencloud-video-generation is available as a video generation API under Apache-2.0 license. You can create videos programmatically by submitting requests and polling for async completion. This enables integration into applications needing multi-shot video synthesis or automated content creation workflows.
How does async video generation work in qwencloud-video-generation?
qwencloud-video-generation processes all video requests asynchronously. Submit your generation request (text prompt, image, frames, or editing parameters), receive a job identifier, then poll the API periodically to check status. Once complete, retrieve your generated video—no synchronous waiting blocks your application.
Can qwencloud-video-generation create reference-based animations?
Yes, qwencloud-video-generation supports reference-based video creation and character role-play animation modes. These specialized features let you generate videos with specific character animations or style references, processed through the same async pipeline as other generation modes.
SKILL.md
Rendered from the published skill. Quoted content, verbatim.
> Agent setup: If your agent doesn't auto-load skills (e.g. Claude Code), > see agent-compatibility.md once per session.
Qwen Video Generation
Generate videos using Wan models. All tasks are asynchronous — submit, then poll until completion. This skill is part of qwencloud/qwencloud-ai.
> ⚠️ Critical Parameter Differences by Mode: > - kf2v (First+Last Frame): Duration is fixed at 5 seconds — other values will fail.
(truncated - see the full file via the links below)
File tree — 15 files
skills/video/qwencloud-video-generation/SKILL.md
skills/video/qwencloud-video-generation/references/agent-compatibility.md
skills/video/qwencloud-video-generation/references/api-guide.md
skills/video/qwencloud-video-generation/references/examples.md
skills/video/qwencloud-video-generation/references/execution-guide.md
skills/video/qwencloud-video-generation/references/merge-media.md
skills/video/qwencloud-video-generation/references/polling-guide.md
skills/video/qwencloud-video-generation/references/prompt-guide.md
skills/video/qwencloud-video-generation/references/request-fields.md
skills/video/qwencloud-video-generation/references/sources.md
skills/video/qwencloud-video-generation/references/workflows.md
skills/video/qwencloud-video-generation/scripts/gossamer.py
skills/video/qwencloud-video-generation/scripts/qwencloud_lib.py
skills/video/qwencloud-video-generation/scripts/video.py
skills/video/qwencloud-video-generation/scripts/video_lib.py
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Generate a video from a text description or prompt”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Generate videos through multiple input modes—text descriptions, single images, frame transitions, character role-play, or video editing—powered by Qianwen's Wan models. All operations run asynchronously; submit your request and poll for completion. The skill auto-detects your task and routes to the right model, with detailed reference guides for prompt engineering, polling patterns, and media workflows.
Wan Video lets you create videos programmatically through AceDataCloud's API, supporting text-to-video, image-to-video, and reference video transfer workflows. Choose from multiple models optimized for different generation types, with output resolutions from 480P to 1080P and optional audio generation.
This skill wraps Aliyun's Wan video generation models through the DashScope SDK, enabling both text-to-video and image-to-video workflows. It standardizes video.generate requests with support for prompt control, duration, frame rate, resolution, seed, and motion parameters across multiple Wan model variants.
Access Qwen's language models for text generation, multi-turn conversations, code writing, and function calling through an OpenAI-compatible interface. The skill supports multiple Qwen variants optimized for different tasks—from general-purpose models to specialized code and reasoning versions—with flexible model selection and streaming output.
Qwen Vision lets you understand images and videos through Qwen's specialized VL and QVQ models. Extract text via OCR, analyze charts and tables, perform visual reasoning, and compare multiple images—all with built-in support for thinking mode and high-resolution processing.
Turn written text into high-quality spoken audio using QwenCloud's TTS models. Choose between fast standard synthesis (qwen3-tts-flash) or instruction-guided style control (qwen3-tts-instruct-flash), or opt for premium quality via CosyVoice. Select from multiple voices and languages to match your content needs.
More skills qwencloud-image-generation (Apache-2.0) · Happyhorse (unlicensed)