qianwen-vision
qianwen-vision lets you process images and videos through Qwen's vision models to extract text, understand visual content, and reason about complex scenes. Use it for OCR, chart/table analysis, multi-image comparison, and video comprehension across multiple model options tuned for speed, precision, or reasoning depth.
qianwen-vision analyzes images and videos using Qwen models to extract text, describe content, and perform visual reasoning.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-06-17
qianwen-vision analyzes images and videos using Qwen models to extract text, describe content, and perform visual reasoning. qianwen-vision lets you process images and videos through Qwen's vision models to extract text, understand visual content, and reason about complex scenes. Use it for OCR, chart/table analysis, multi-image comparison, and video comprehension across multiple model options tuned for speed, precision, or reasoning depth.
Use it when
- Yes.
- qianwen-vision can understand video content by analyzing video frames to extract key information and provide comprehension.
Verify before relying
Read SKILL.md below before installing (15 files). Open directory: indexed for reading, not audited.
Install
QianWen-AI/qianwen-ai/qianwen-vision · repository language: Python
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What can qianwen-vision do with images and videos?
qianwen-vision processes images and videos through Qwen's vision models to extract text, understand visual content, and reason about complex scenes. You can use it for OCR on receipts and documents, analyze charts and tables, compare multiple images, and extract key information from video frames.
Can I extract text from screenshot using OCR with qianwen-vision?
Yes. qianwen-vision supports OCR capabilities to extract text from screenshots, documents, receipts, invoices, and tables. It can read and digitize text from image files, making it useful for document scanning and data extraction workflows.
Does qianwen-vision support video analysis?
qianwen-vision can understand video content by analyzing video frames to extract key information and provide comprehension. This enables video frame analysis and summarization, helping you derive insights from video material.
How does qianwen-vision handle visual reasoning tasks?
qianwen-vision performs visual reasoning by solving problems step-by-step from images. It can analyze charts, work through math problems shown in images, and answer visual questions about complex scenes, supporting reasoning-focused workflows.
Can qianwen-vision compare multiple images?
Yes. qianwen-vision can compare two or more images side by side and analyze visual differences between them, making it useful for multi-image comparison tasks and identifying changes across related images.
What license does qianwen-vision use?
qianwen-vision is released under the Apache-2.0 license, allowing broad use and modification while maintaining appropriate attribution requirements.
SKILL.md
Rendered from the published skill. Quoted content, verbatim.
> Agent setup: If your agent doesn't auto-load skills (e.g. Claude Code), > see agent-compatibility.md once per session.
Qwen Vision (Image & Video Understanding)
Analyze images and videos using Qwen VL and QVQ models. This skill is part of QianWen-AI/qianwen-ai.
Skill directory
Use this skill's internal files to execute and learn. Load reference files on demand when the default path fails or you need
(truncated - see the full file via the links below)
File tree — 15 files
skills/vision/qianwen-vision/SKILL.md
skills/vision/qianwen-vision/references/agent-compatibility.md
skills/vision/qianwen-vision/references/api-guide.md
skills/vision/qianwen-vision/references/curl-examples.md
skills/vision/qianwen-vision/references/execution-guide.md
skills/vision/qianwen-vision/references/ocr.md
skills/vision/qianwen-vision/references/prompt-guide.md
skills/vision/qianwen-vision/references/sources.md
skills/vision/qianwen-vision/references/visual-reasoning.md
skills/vision/qianwen-vision/scripts/analyze.py
skills/vision/qianwen-vision/scripts/gossamer.py
skills/vision/qianwen-vision/scripts/ocr.py
skills/vision/qianwen-vision/scripts/qianwen_lib.py
skills/vision/qianwen-vision/scripts/reason.py
skills/vision/qianwen-vision/scripts/vision_lib.py
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Analyze and describe images or videos to extract information and insights”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
This skill walks you through QianWen API credential setup and validation. It handles both standard and Token Plan key types, manages environment variables and .env files, and provides verification steps to confirm your authentication is working correctly.
qianwen-text lets you interact with Qwen language models for text generation, conversation, and code writing through an OpenAI-compatible interface. Choose from multiple Qwen models including qwen3.6-plus (recommended default), specialized code models, and reasoning variants. The skill handles API authentication, provides execution guides, and includes prompt engineering references.
Qwen Vision lets you understand images and videos through Qwen's specialized VL and QVQ models. Extract text via OCR, analyze charts and tables, perform visual reasoning, and compare multiple images—all with built-in support for thinking mode and high-resolution processing.
Qwen Audio TTS turns written text into natural-sounding speech using Qwen's TTS engine. Choose from multiple voices and models—including the fast qwen3-tts-flash for standard tasks or instruction-guided variants for tone control—then output audio directly to file.
Generate videos through multiple input modes—text descriptions, single images, frame transitions, character role-play, or video editing—powered by Qianwen's Wan models. All operations run asynchronously; submit your request and poll for completion. The skill auto-detects your task and routes to the right model, with detailed reference guides for prompt engineering, polling patterns, and media workflows.
Create images from text descriptions or edit existing ones using Wan and Qwen Image models. Supports style transfer, subject consistency across multiple reference images, text rendering in images, and interleaved text-image output for tutorials and guides. Choose from multiple models optimized for different tasks—from fast drafts to high-resolution 4K generation.
More skills qwencloud-text (Apache-2.0)