$npx skillfedfor your agent

qwencloud-vision

Qwen Vision lets you understand images and videos through Qwen's specialized VL and QVQ models. Extract text via OCR, analyze charts and tables, perform visual reasoning, and compare multiple images—all with built-in support for thinking mode and high-resolution processing.

Qwen Vision analyzes images and videos using Qwen's VL and QVQ models for understanding, reasoning, and text extraction.

AI-generated summary based on this skill's SKILL.md

★ 34  0 Apache-2.0updated by QwenCloud

Decision gist · record as of 2026-07-22

Qwen Vision analyzes images and videos using Qwen's VL and QVQ models for understanding, reasoning, and text extraction. Qwen Vision lets you understand images and videos through Qwen's specialized VL and QVQ models. Extract text via OCR, analyze charts and tables, perform visual reasoning, and compare multiple images—all with built-in support for thinking mode and high-resolution processing.

manual: git clone https://github.com/QwenCloud/qwencloud-ai → cp -r qwencloud-ai/skills/vision/qwencloud-vision ~/.claude/skills/qwencloud-vision
skills/vision/qwencloud-vision/SKILL.md · version 3592f0cb

Use it when

  • qwencloud-vision extracts text and structured data from images through OCR capabilities.
  • Yes, qwencloud-vision performs visual reasoning on charts, diagrams, and math problems.

Verify before relying

Read SKILL.md below before installing (15 files). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

QwenCloud/qwencloud-ai/qwencloud-vision · repository language: Python

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What can qwencloud-vision analyze or describe in image and video content?

qwencloud-vision uses Qwen's specialized VL and QVQ models to analyze and describe images and videos. The skill can identify objects, scenes, text, and visual elements within your media. It supports high-resolution processing and thinking mode for deeper analysis, making it suitable for detailed visual understanding tasks across photos, screenshots, and video frames.

How do I extract text from screenshot OCR using qwencloud-vision?

qwencloud-vision extracts text and structured data from images through OCR capabilities. Simply provide a screenshot or image containing text, and the skill will recognize and extract the content. This works for documents, signs, handwriting, and printed text, converting image-based text into machine-readable format for further processing or analysis.

Can qwencloud-vision perform visual reasoning on math problems?

Yes, qwencloud-vision performs visual reasoning on charts, diagrams, and math problems. The skill analyzes mathematical expressions, geometric figures, and problem layouts within images. Using Qwen's reasoning capabilities and thinking mode, it can interpret complex visual information and provide solutions or explanations for mathematical content.

What does this chart show—can qwencloud-vision read charts?

qwencloud-vision reads and interprets charts, graphs, and diagrams. It understands chart types, axes, legends, and data relationships to explain what information is being presented. The skill can also extract numerical values and trends from visual representations, making it useful for data analysis and report comprehension.

Does qwencloud-vision support comparing multiple images?

qwencloud-vision compares multiple images and analyzes video frames. You can provide two or more images to identify similarities, differences, and relationships between them. This capability extends to frame-by-frame video analysis, allowing you to track changes, movements, or content evolution across sequential visual inputs.

Can qwencloud-vision extract table data from document screenshots?

qwencloud-vision understands document layout and parses tables from screenshots. It recognizes table structures, cell boundaries, and content organization, then extracts the data in structured format. This makes it effective for digitizing tabular information from PDFs, scanned documents, or spreadsheet screenshots.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

> Agent setup: If your agent doesn't auto-load skills (e.g. Claude Code), > see agent-compatibility.md once per session.

Qwen Vision (Image & Video Understanding)

Analyze images and videos using Qwen VL and QVQ models. This skill is part of qwencloud/qwencloud-ai.

Skill directory

Use this skill's internal files to execute and learn. Load reference files on demand when the default path fails or you need

(truncated - see the full file via the links below)

File tree — 15 files
skills/vision/qwencloud-vision/SKILL.md
skills/vision/qwencloud-vision/references/agent-compatibility.md
skills/vision/qwencloud-vision/references/api-guide.md
skills/vision/qwencloud-vision/references/curl-examples.md
skills/vision/qwencloud-vision/references/execution-guide.md
skills/vision/qwencloud-vision/references/ocr.md
skills/vision/qwencloud-vision/references/prompt-guide.md
skills/vision/qwencloud-vision/references/sources.md
skills/vision/qwencloud-vision/references/visual-reasoning.md
skills/vision/qwencloud-vision/scripts/analyze.py
skills/vision/qwencloud-vision/scripts/gossamer.py
skills/vision/qwencloud-vision/scripts/ocr.py
skills/vision/qwencloud-vision/scripts/qwencloud_lib.py
skills/vision/qwencloud-vision/scripts/reason.py
skills/vision/qwencloud-vision/scripts/vision_lib.py

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Analyze or describe image and video content using Qwen vision models”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

qianwen-vision
by QianWen-AI · QianWen-AI/qianwen-ai

qianwen-vision lets you process images and videos through Qwen's vision models to extract text, understand visual content, and reason about complex scenes. Use it for OCR, chart/table analysis, multi-image comparison, and video comprehension across multiple model options tuned for speed, precision, or reasoning depth.

Apache-2.0updated Jun 2026
★ 38repo stars
qwencloud-text
by QwenCloud · QwenCloud/qwencloud-ai

Access Qwen's language models for text generation, multi-turn conversations, code writing, and function calling through an OpenAI-compatible interface. The skill supports multiple Qwen variants optimized for different tasks—from general-purpose models to specialized code and reasoning versions—with flexible model selection and streaming output.

Apache-2.0updated Jul 2026
★ 34repo stars
qwencloud-video-generation
by QwenCloud · QwenCloud/qwencloud-ai

Create videos asynchronously using QwenCloud's Wan models across multiple modes: text-to-video, image-to-video, first-and-last-frame transitions, role-play character animation, and video editing. Submit your request and poll for completion—no synchronous waiting required.

Apache-2.0updated Jul 2026
★ 34repo stars
qwencloud-audio-tts
by QwenCloud · QwenCloud/qwencloud-ai

Turn written text into high-quality spoken audio using QwenCloud's TTS models. Choose between fast standard synthesis (qwen3-tts-flash) or instruction-guided style control (qwen3-tts-instruct-flash), or opt for premium quality via CosyVoice. Select from multiple voices and languages to match your content needs.

Apache-2.0updated Jul 2026
★ 34repo stars
qwencloud-image-generation
by QwenCloud · QwenCloud/qwencloud-ai

Create images from text descriptions or edit existing ones with Wan and Qwen Image models. This skill handles text-to-image generation, style transfer, subject consistency across reference images, and interleaved text-image output for tutorials and guides.

Apache-2.0updated Jul 2026
★ 34repo stars
qianwen-text
by QianWen-AI · QianWen-AI/qianwen-ai

qianwen-text lets you interact with Qwen language models for text generation, conversation, and code writing through an OpenAI-compatible interface. Choose from multiple Qwen models including qwen3.6-plus (recommended default), specialized code models, and reasoning variants. The skill handles API authentication, provides execution guides, and includes prompt engineering references.

Apache-2.0updated Jun 2026
★ 38repo stars

More skills qianwen-ops-auth (Apache-2.0) · qianwen-audio-tts (Apache-2.0) · qianwen-image-generation (Apache-2.0)

Tags
multimodal-aidocument-extractionvisual-qavideo-analysistext-recognitionchart-parsingimage-captioningreasoning-enginestructured-data-extraction