{"enrichment":{"faq":[{"a":"qwencloud-vision uses Qwen's specialized VL and QVQ models to analyze and describe images and videos. The skill can identify objects, scenes, text, and visual elements within your media. It supports high-resolution processing and thinking mode for deeper analysis, making it suitable for detailed visual understanding tasks across photos, screenshots, and video frames.","q":"What can qwencloud-vision analyze or describe in image and video content?"},{"a":"qwencloud-vision extracts text and structured data from images through OCR capabilities. Simply provide a screenshot or image containing text, and the skill will recognize and extract the content. This works for documents, signs, handwriting, and printed text, converting image-based text into machine-readable format for further processing or analysis.","q":"How do I extract text from screenshot OCR using qwencloud-vision?"},{"a":"Yes, qwencloud-vision performs visual reasoning on charts, diagrams, and math problems. The skill analyzes mathematical expressions, geometric figures, and problem layouts within images. Using Qwen's reasoning capabilities and thinking mode, it can interpret complex visual information and provide solutions or explanations for mathematical content.","q":"Can qwencloud-vision perform visual reasoning on math problems?"},{"a":"qwencloud-vision reads and interprets charts, graphs, and diagrams. It understands chart types, axes, legends, and data relationships to explain what information is being presented. The skill can also extract numerical values and trends from visual representations, making it useful for data analysis and report comprehension.","q":"What does this chart show\u2014can qwencloud-vision read charts?"},{"a":"qwencloud-vision compares multiple images and analyzes video frames. You can provide two or more images to identify similarities, differences, and relationships between them. This capability extends to frame-by-frame video analysis, allowing you to track changes, movements, or content evolution across sequential visual inputs.","q":"Does qwencloud-vision support comparing multiple images?"},{"a":"qwencloud-vision understands document layout and parses tables from screenshots. It recognizes table structures, cell boundaries, and content organization, then extracts the data in structured format. This makes it effective for digitizing tabular information from PDFs, scanned documents, or spreadsheet screenshots.","q":"Can qwencloud-vision extract table data from document screenshots?"}],"shadow_tags":["multimodal-ai","document-extraction","visual-qa","video-analysis","text-recognition","chart-parsing","image-captioning","reasoning-engine","structured-data-extraction"],"summary_rewrite":"Qwen Vision lets you understand images and videos through Qwen's specialized VL and QVQ models. Extract text via OCR, analyze charts and tables, perform visual reasoning, and compare multiple images\u2014all with built-in support for thinking mode and high-resolution processing."},"files":[{"bytes":19870,"path":"skills/vision/qwencloud-vision/SKILL.md","sha256":"c208bd49d5af80f0c0f99357dec66d6ffa480d5776f66c95765027e53be656e1","url":"https://skillfed.io/files/QwenCloud/qwencloud-ai/qwencloud-vision/3592f0cb/SKILL.md"}],"id":"QwenCloud/qwencloud-ai/qwencloud-vision","links":{"html":"https://skillfed.io/QwenCloud/qwencloud-ai/qwencloud-vision","md":"https://skillfed.io/QwenCloud/qwencloud-ai/qwencloud-vision.md","repo":"https://github.com/QwenCloud/qwencloud-ai"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":0,"language":"Python","last_updated":"2026-07-22","license":"Apache-2.0","name":"qwencloud-vision","publisher":"QwenCloud","stars":34},"relations":{"categories":["vision-ai","document-processing","multimodal-analysis"],"similar":[{"id":"QianWen-AI/qianwen-ai/qianwen-vision"},{"id":"QwenCloud/qwencloud-ai/qwencloud-text"},{"id":"QwenCloud/qwencloud-ai/qwencloud-video-generation"},{"id":"QwenCloud/qwencloud-ai/qwencloud-audio-tts"},{"id":"QwenCloud/qwencloud-ai/qwencloud-image-generation"},{"id":"QwenCloud/qwencloud-ai/qwencloud-ops-auth"},{"id":"QianWen-AI/qianwen-ai/qianwen-text"},{"id":"QwenCloud/qwencloud-ai/qwencloud-model-selector"},{"id":"QianWen-AI/qianwen-ai/qianwen-audio-tts"},{"id":"QianWen-AI/qianwen-ai/qianwen-image-generation"}]},"slug":{"owner":"QwenCloud","repo":"qwencloud-ai","skill":"qwencloud-vision"},"version":"3592f0cb"}
