$npx skillfedfor your agent

ai-edge-quantizer

A quantizer for advanced developers to quantize converted AI Edge models.

With conditionsPyPI Software DevelopmentReleased Jul 2026169.7K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — ai_edge_quantizer-0.8.0-py3-none-any.whl
v0.8.0 · released 2026-07-13 · Python >=3.10 · 7 runtime deps: absl-py, immutabledict, numpy, scipy, ml_dtypes, ai-edge-litert, litert-lm-builder

Yes, if you are quantizing LiteRT models for edge deployment. The package is actively maintained, has low install friction, carries a permissive license, and provides a clear API with multiple quantization strategies. No known vulnerabilities. Best suited for advanced developers; requires understanding of quantization trade-offs and model format requirements. Start with dynamic quantization recipes if you lack calibration data.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • Input model must be an unquantized FP32 LiteRT model in FlatBuffer format with .tflite extension.
  • TensorFlow (tf-nightly) is listed as a dependency in the documentation.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows use in commercial and proprietary projects with minimal restrictions.

last release 2026-07-13 (32 days) · last repo commit 2026-08-14 · 189 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 169,674 downloads/mo, #10,410 on PyPI

Verify before relying

pip install ai-edge-quantizer

from ai_edge_quantizer import quantizer, recipe

qt = quantizer.Quantizer("path/to/input.tflite")
qt.load_quantization_recipe(recipe.dynamic_wi8_afp32())
qt.quantize().export_model("/path/to/output.tflite")
  • Whether tf-nightly is an actual runtime dependency or only a build/development requirement
  • Hardware compatibility details for each quantization strategy beyond the CPU/GPU vs NPU recommendation
  • Performance benchmarks or typical latency/size improvements across different model types
Same gist for agents: .md · .json

What it is and what it does

AI Edge Quantizer is a tool for converting unquantized LiteRT models into quantized versions optimized for edge device deployment. It targets advanced developers working with resource-constrained environments, particularly for GenAI and large language models. The package provides three quantization strategies: dynamic quantization (weights quantized, activations remain float, no calibration needed), weight-only quantization (reduced model size with float computation), and static quantization (both weights and activations quantized, requires calibration data). Users define quantization behavior through recipes that specify which operators to quantize, bit-widths, symmetry, and granularity settings.

The workflow is straightforward: instantiate a Quantizer with an input .tflite file, load a quantization recipe (either from built-in templates or custom-defined), then quantize and export. The package depends on numpy, scipy, absl-py for core functionality, plus Google's ai-edge-litert and litert-lm-builder for LiteRT model handling. It supports Python 3.10–3.13 on Linux and macOS, with active maintenance and nightly releases.

Use it for

  • Reduce model size and memory footprint for deployment on mobile and embedded devices without calibration data using dynamic quantization.
  • Optimize inference latency on NPU hardware by applying static quantization with calibration on representative data.
  • Selectively quantize specific operators or layers while keeping others in FP32 to balance quality and performance.
  • Convert GenAI and large language models to 4-bit or 8-bit weight quantization for on-device inference.
  • Experiment with mixed-precision quantization strategies to find the optimal trade-off between model size and accuracy.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are quantizing LiteRT models for edge deployment.

The package is actively maintained, has low install friction, carries a permissive license, and provides a clear API with multiple quantization strategies. No known vulnerabilities. Best suited for advanced developers; requires understanding of quantization trade-offs and model format requirements. Start with dynamic quantization recipes if you lack calibration data.

Install

ai-edge-quantizer on PyPI

Before you install

Low friction install with a pure-Python wheel. Active maintenance with recent releases and passing unit tests. Depends on established packages (numpy, scipy, absl-py) plus Google's ai-edge-litert and litert-lm-builder, which may require additional setup.

Requires Python 3.10 or later. Input model must be an unquantized FP32 LiteRT model in FlatBuffer format with .tflite extension. TensorFlow (tf-nightly) is listed as a dependency in the documentation.

License in practice

Apache-2.0 permissive license allows use in commercial and proprietary projects with minimal restrictions.

Quickstart

pip install ai-edge-quantizer

from ai_edge_quantizer import quantizer, recipe

qt = quantizer.Quantizer("path/to/input.tflite")
qt.load_quantization_recipe(recipe.dynamic_wi8_afp32())
qt.quantize().export_model("/path/to/output.tflite")

Verify before relying

  • Whether tf-nightly is an actual runtime dependency or only a build/development requirement
  • Hardware compatibility details for each quantization strategy beyond the CPU/GPU vs NPU recommendation
  • Performance benchmarks or typical latency/size improvements across different model types

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
7 packages
absl-pyimmutabledictnumpyscipyml_dtypesai-edge-litertlitert-lm-builder
MaintenanceActively maintained 32 days since the last release
Last repo commit
First released
Downloads169,674 / month, #10,410 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: MathematicsTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules

Evidence: ai_edge_quantizer-0.8.0-py3-none-any.whl

Tags

Capabilities
model quantization litertedge device model optimizationneural network weight quantizationon-device ml model compressiontflite model quantizergenai model optimizationdynamic quantization framework
Topics
model-compressionedge-mlquantization
PyPI keywords
On-Device MLAIGoogleTFLiteQuantizationLLMsGenAI

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “model quantization litert”

  • ai-edge-quantizerQuantizes LiteRT models to reduce size and improve inference…
  • litert-torchConverts PyTorch models to .tflite format for on-device deployment on…
  • qwixQwix is a JAX quantization library that applies Quantization-Aware…

Give your agent the search over MCP, or paste the wish link into any chat.

More Software Development packages

typing-extensions Worth it
PyPI · Software Development · released Jul 2026

Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.

PSF-2.0pure Python · 3.9+
1.9Bdownloads / mo
numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
fastapi Worth it
PyPI · Software Development · released Jul 2026

FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.

MITpure Python · 3.10+
568.6Mdownloads / mo
annotated-doc With conditions
PyPI · Software Development · released Jul 2026

Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.

MITpure Python · 3.9+
456.2Mdownloads / mo
typer Worth it
PyPI · Software Development · released Aug 2026

Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.

Install it if you are building CLIs in Python.

MITpure Python · 3.10+
369.3Mdownloads / mo
distlib With conditions
PyPI · Software Development · released Jun 2026

Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.

permissive licensepure Python
323.3Mdownloads / mo

See also ai-edge-litert · optimum-quanto · ai-edge-litert-nightly · litert-converter · tosa-adapter-model-explorer · litert-torch · diffq · ai-edge-model-explorer · auto-gptq · qwix

Further reading