comfy-aimdo
AI Model Dynamic Offloader for ComfyUI
Decision gist · record as of 2026-08-14
Yes, if you are building or using ComfyUI-like workflows with multiple large models on limited VRAM and are willing to accept the constraints: Nvidia GPU only, specific PyTorch/CUDA versions, and active management of model priority and allocator flushing. The license status is unclear, so verify terms first. Not suitable for general PyTorch projects or non-Nvidia hardware.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Nvidia GPU, PyTorch 2.8+, CUDA 12.8+, and Python 3.9+.
- Windows 11+ or Linux only.
- Low friction installation with no runtime dependencies.
License · maintenance · safety
(unclear) — License status is unclear—no SPDX identifier or raw license text is available. Verify the actual license terms in the repository before using in commercial or proprietary projects.
last release 2026-08-04 (10 days) · last repo commit 2026-08-04 · 55 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,835,202 downloads/mo, #3,504 on PyPI
Alternatives
Verify before relying
pip install comfy-aimdo
import comfy_aimdo
# Create a VBAR, allocate tensors, call fault() to load weights on demand
# See examples/example.py in the repository for full usage pattern- Performance overhead of the offloading mechanism compared to standard PyTorch allocation
- Compatibility with specific PyTorch versions beyond the stated 2.8+ minimum
- Real-world stability and fragmentation behavior under sustained production workloads
What it is and what it does
comfy-aimdo is a custom PyTorch VRAM allocator designed to handle GPU memory pressure by dynamically offloading model weights to system memory. Instead of the standard PyTorch allocator, it uses CUDA's virtual address reservation APIs to create Virtual Base Address Registers (VBARs) for models—reserving address space without consuming VRAM upfront. Tensors are allocated within these VBARs and only faulted into actual GPU memory when needed via an explicit `fault()` call. If VRAM is insufficient, the allocator falls back to temporary GPU tensors that are garbage-collected after use, effectively spilling to system memory.
The allocator implements a priority system where more recently created VBARs take precedence, and within a VBAR, lower addresses have higher priority. When a weight is evicted due to memory pressure, a watermark is set to prevent repeatedly faulting in already-offloaded weights. Applications can also call `prioritize()` to promote an existing model to top priority. The design assumes regular weight access patterns and recommends flushing the PyTorch caching allocator between model runs to avoid fragmentation. This is specialized infrastructure for scenarios like ComfyUI where multiple large models must coexist with limited VRAM.
Use it for
- Load and run multiple large language or diffusion models sequentially on a single GPU without OOM errors
- Implement dynamic model swapping in inference pipelines where model priority changes based on workflow order
- Optimize VRAM utilization in multi-model workflows by offloading lower-priority weights to system memory
- Reduce memory fragmentation in long-running applications that repeatedly load and unload different models
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building or using ComfyUI-like workflows with multiple large models on limited VRAM and are willing to accept the constraints: Nvidia GPU only, specific PyTorch/CUDA versions, and active management of model priority and allocator flushing.
The license status is unclear, so verify terms first. Not suitable for general PyTorch projects or non-Nvidia hardware.
Install
comfy-aimdo on PyPI
Before you install
Low friction installation with no runtime dependencies. Active maintenance with recent releases; however, the package is young (first release January 2026) and relatively niche, so production stability remains unproven.
Requires Nvidia GPU, PyTorch 2.8+, CUDA 12.8+, and Python 3.9+. Windows 11+ or Linux only.
License in practice
License status is unclear—no SPDX identifier or raw license text is available. Verify the actual license terms in the repository before using in commercial or proprietary projects.
Quickstart
pip install comfy-aimdo
import comfy_aimdo
# Create a VBAR, allocate tensors, call fault() to load weights on demand
# See examples/example.py in the repository for full usage pattern
Verify before relying
- Performance overhead of the offloading mechanism compared to standard PyTorch allocation
- Compatibility with specific PyTorch versions beyond the stated 2.8+ minimum
- Real-world stability and fragmentation behavior under sustained production workloads
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 10 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,835,202 / month, #3,504 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: comfy_aimdo-0.4.13-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pytorch vram allocator”
- comfy-aimdoA PyTorch VRAM allocator that dynamically offloads model weights to…
- mmgpOptimizes GPU memory usage for running large generative models on…
- unslothUnsloth accelerates training and fine-tuning of large language models…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also mmgp · nvidia-resiliency-ext · instanttensor · torch · flashoptim · cuequivariance-ops-torch-cu12 · tensorizer · binpacking · omnimalloc · cufile-python