Dashscope
Connect Dashscope to harness Alibaba Cloud's Qwen models for multimodal content creation. Generate images via qwen-image-2.0-pro, synthesize speech with qwen3-tts-flash, or transcribe audio with word-level timing via qwen3-asr-flash-filetrans—all through native DashScope endpoints.
Dashscope integrates Alibaba Cloud's Qwen models for image generation, text-to-speech, and speech-to-text with word-level timestamps.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-07-24
Dashscope integrates Alibaba Cloud's Qwen models for image generation, text-to-speech, and speech-to-text with word-level timestamps. Connect Dashscope to harness Alibaba Cloud's Qwen models for multimodal content creation. Generate images via qwen-image-2.0-pro, synthesize speech with qwen3-tts-flash, or transcribe audio with word-level timing via qwen3-asr-flash-filetrans—all through native DashScope endpoints.
Use it when
- Dashscope integrates as a connector into your agent or workflow, allowing you to call Alibaba Cloud's Qwen models directly.
- Dashscope provides three primary functions: generate images using qwen-image-2.0-pro, synthesize speech with qwen3-tts-flash.
Install
calesthio/OpenMontage/dashscope · repository language: Python
generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What is Dashscope and what can it do?
Dashscope connects you to Alibaba Cloud's Qwen models for multimodal content creation. It enables image generation via qwen-image-2.0-pro, speech synthesis with qwen3-tts-flash, and audio transcription with word-level timing via qwen3-asr-flash-filetrans through native DashScope endpoints.
How do I integrate Dashscope with my agent or workflow?
Dashscope integrates as a connector into your agent or workflow, allowing you to call Alibaba Cloud's Qwen models directly. Configure it with your DashScope API credentials, then invoke image generation, text-to-speech, or speech-to-text capabilities as steps in your automation pipeline.
How to use Dashscope for image generation and audio tasks?
Dashscope provides three primary functions: generate images using qwen-image-2.0-pro, synthesize speech with qwen3-tts-flash, and transcribe audio with word-level timing via qwen3-asr-flash-filetrans. Each integrates into your workflow as a discrete tool you can chain with other steps.
What is the dashscope configuration process?
Configure Dashscope by providing your DashScope API credentials and selecting which Qwen model endpoint you need—qwen-image-2.0-pro for images, qwen3-tts-flash for speech synthesis, or qwen3-asr-flash-filetrans for audio transcription. Once set up, the skill is ready to use in your agent.
Is Dashscope open source and what license does it use?
Dashscope is released under the AGPL-3.0 license, making it open source. This means you can use, modify, and distribute it freely under the terms of the AGPL-3.0 agreement.
Can Dashscope handle multimodal content creation?
Yes, Dashscope is built for multimodal content creation. It supports image generation, text-to-speech synthesis, and speech-to-text transcription all within a single skill, letting you combine these capabilities in complex workflows.
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Integrate Dashscope with my agent or workflow”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Convert text to natural-sounding speech or transcribe audio files using Alibaba Cloud's DashScope API. This skill handles both speech synthesis with customizable voice styles and speech recognition for audio up to 12 hours long, supporting 30+ languages and multiple audio formats.
Doubao TTS produces natural Mandarin and multilingual audio narration via Volcengine's Speech 2.0 API. It returns word-level timing metadata for precise subtitle synchronization, making it ideal for video projects requiring accurate caption alignment. Configure your voice preference and speech rate, then generate samples before committing to full narrations.
Z-Image Turbo enables fast text-to-image generation through Alibaba Cloud's DashScope multimodal API. Control output dimensions, randomization seed, and optional prompt enhancement while managing costs based on your feature selection.
This skill harnesses Alibaba Cloud DashScope to create, modify, and interpret images through multiple AI models. It handles text-to-image generation, image editing with local or remote files, and visual content analysis—routing each task to the appropriate model and script automatically.
This skill provides standardized image generation using Alibaba's Qwen models through the DashScope SDK. It normalizes requests across qwen-image variants with support for prompts, negative prompts, dimensions, seeds, and reference images, while handling authentication via environment variables or credential files.
Turn written text into high-quality spoken audio using QwenCloud's TTS models. Choose between fast standard synthesis (qwen3-tts-flash) or instruction-guided style control (qwen3-tts-instruct-flash), or opt for premium quality via CosyVoice. Select from multiple voices and languages to match your content needs.
More skills qianwen-audio-tts (Apache-2.0) · aliyun-wan-i2v (MIT)