onnxslim
OnnxSlim: A Toolkit to Help Optimize Onnx Model
Decision gist · record as of 2026-08-14
Yes. OnnxSlim is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and installs with low friction. It is widely adopted in production ML frameworks (NVIDIA, HuggingFace, ultralytics) and has strong community adoption. Install if you work with ONNX models and need to optimize for speed or size.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.8 or later.
- Input must be a valid ONNX model file.
- Low friction installation with a pure Python wheel.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution in both open-source and commercial projects with minimal restrictions.
last release 2026-08-01 (13 days) · last repo commit 2026-08-01 · 513 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,281,234 downloads/mo, #4,117 on PyPI
Alternatives
Verify before relying
pip install onnxslim
import onnx
import onnxslim
model = onnx.load("model.onnx")
slimmed_model = onnxslim.slim(model)
if slimmed_model:
onnx.save(slimmed_model, "slimmed_model.onnx")- Specific optimization techniques used (pruning, quantization, graph simplification, or others)
- Whether accuracy preservation is measured quantitatively or qualitatively
- Inference speed improvements are claimed but not quantified in the excerpt
What it is and what it does
OnnxSlim is a toolkit for optimizing ONNX neural network models by reducing operator count and model size while maintaining inference accuracy. It provides both a command-line interface and a Python API for model optimization. The package depends on onnx for model handling, sympy for symbolic computation, ml-dtypes for numeric types, packaging for version management, and colorama for terminal output formatting.
The tool is designed for developers deploying neural networks who need faster inference without sacrificing model accuracy. It integrates into major ML frameworks and deployment pipelines, having been adopted by NVIDIA TensorRT-Model-Optimizer, HuggingFace optimum and transformers.js, ultralytics, and others. Installation is straightforward via pip with no compiled dependencies.
Use it for
- Optimize ONNX models for edge deployment where model size and inference latency are critical constraints.
- Reduce inference latency in production serving pipelines while maintaining model accuracy.
- Prepare pre-trained models for mobile or embedded inference with lower computational overhead.
- Integrate model optimization into automated ML pipelines as a preprocessing step before deployment.
- Benchmark and compare inference performance before and after optimization on target hardware.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
OnnxSlim is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and installs with low friction. It is widely adopted in production ML frameworks (NVIDIA, HuggingFace, ultralytics) and has strong community adoption. Install if you work with ONNX models and need to optimize for speed or size.
Install
onnxslim on PyPI
Before you install
Low friction installation with a pure Python wheel. Actively maintained with a release in the last two weeks and a recent commit history. Five runtime dependencies are all stable, widely-used packages.
Requires Python 3.8 or later. Input must be a valid ONNX model file.
License in practice
MIT license permits unrestricted use, modification, and distribution in both open-source and commercial projects with minimal restrictions.
Quickstart
pip install onnxslim
import onnx
import onnxslim
model = onnx.load("model.onnx")
slimmed_model = onnxslim.slim(model)
if slimmed_model:
onnx.save(slimmed_model, "slimmed_model.onnx")
Verify before relying
- Specific optimization techniques used (pruning, quantization, graph simplification, or others)
- Whether accuracy preservation is measured quantitatively or qualitatively
- Inference speed improvements are claimed but not quantified in the excerpt
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagescoloramaml-dtypesonnxpackagingsympy |
| Maintenance | Actively maintained 13 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,281,234 / month, #4,117 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Topic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: onnxslim-0.1.95-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “reduce onnx model size”
- onnxslimOnnxSlim reduces the size and operator count of ONNX models while…
- sne4onnxExtracts subgraphs from ONNX model files by specifying input and…
- scs4onnxCompresses ONNX model files by deduplicating constant tensor values…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also nudenet · nvidia-modelopt · optimum-onnx · optimum · onnxruntime_extensions · onnxoptimizer · mnn · ssi4onnx · onnxruntime-gpu · sit4onnx