skillfed

qianwen-vision

qianwen-vision lets you process images and videos through Qwen's vision models to extract text, understand visual content, and reason about complex scenes. Use it for OCR, chart/table analysis, multi-image comparison, and video comprehension across multiple model options tuned for speed, precision, or reasoning depth.

qianwen-vision analyzes images and videos using Qwen models to extract text, describe content, and perform visual reasoning.

AI-generated summary based on this skill's SKILL.md

38 0 Apache-2.0 updated by QianWen-AI

Install

QianWen-AI/qianwen-ai/qianwen-vision · repository language: Python

git clone https://github.com/QianWen-AI/qianwen-ai
cp -r qianwen-ai/skills/vision/qianwen-vision ~/.claude/skills/qianwen-vision
npx skillfed install QianWen-AI/qianwen-ai/qianwen-vision

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What can qianwen-vision do with images and videos?

qianwen-vision processes images and videos through Qwen's vision models to extract text, understand visual content, and reason about complex scenes. You can use it for OCR on receipts and documents, analyze charts and tables, compare multiple images, and extract key information from video frames.

Can I extract text from screenshot using OCR with qianwen-vision?

Yes. qianwen-vision supports OCR capabilities to extract text from screenshots, documents, receipts, invoices, and tables. It can read and digitize text from image files, making it useful for document scanning and data extraction workflows.

Does qianwen-vision support video analysis?

qianwen-vision can understand video content by analyzing video frames to extract key information and provide comprehension. This enables video frame analysis and summarization, helping you derive insights from video material.

How does qianwen-vision handle visual reasoning tasks?

qianwen-vision performs visual reasoning by solving problems step-by-step from images. It can analyze charts, work through math problems shown in images, and answer visual questions about complex scenes, supporting reasoning-focused workflows.

Can qianwen-vision compare multiple images?

Yes. qianwen-vision can compare two or more images side by side and analyze visual differences between them, making it useful for multi-image comparison tasks and identifying changes across related images.

What license does qianwen-vision use?

qianwen-vision is released under the Apache-2.0 license, allowing broad use and modification while maintaining appropriate attribution requirements.

SKILL.md

rendered from the published skill — quoted content, verbatim

> Agent setup: If your agent doesn't auto-load skills (e.g. Claude Code), > see agent-compatibility.md once per session.

Qwen Vision (Image & Video Understanding)

Analyze images and videos using Qwen VL and QVQ models. This skill is part of QianWen-AI/qianwen-ai.

Skill directory

Use this skill's internal files to execute and learn. Load reference files on demand when the default path fails or you need

(truncated - see the full file via the links below)

Read as markdown · JSON record · Browse the source repository

File tree — 15 files
skills/vision/qianwen-vision/SKILL.md
skills/vision/qianwen-vision/references/agent-compatibility.md
skills/vision/qianwen-vision/references/api-guide.md
skills/vision/qianwen-vision/references/curl-examples.md
skills/vision/qianwen-vision/references/execution-guide.md
skills/vision/qianwen-vision/references/ocr.md
skills/vision/qianwen-vision/references/prompt-guide.md
skills/vision/qianwen-vision/references/sources.md
skills/vision/qianwen-vision/references/visual-reasoning.md
skills/vision/qianwen-vision/scripts/analyze.py
skills/vision/qianwen-vision/scripts/gossamer.py
skills/vision/qianwen-vision/scripts/ocr.py
skills/vision/qianwen-vision/scripts/qianwen_lib.py
skills/vision/qianwen-vision/scripts/reason.py
skills/vision/qianwen-vision/scripts/vision_lib.py

Related skills

Tags

multimodal-ai document-processing text-recognition visual-qa video-analysis chart-reading object-detection reasoning-engine content-extraction