sagemaker-inference
Open source toolkit for helping create serving containers to run on Amazon SageMaker.
Decision gist · record as of 2026-08-14
Yes, but with caution. Install if you are actively deploying models to Amazon SageMaker and need a structured framework for containerized inference. The permissive Apache 2.0 license poses no barrier. However, the abandoned repository status (last commit 2023-11-20) means you should verify compatibility with your target SageMaker version and Python runtime before committing to production use. If you are starting a new project, check whether SageMaker's prebuilt containers or newer alternatives better suit your needs.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Must be installed within a Docker container build; requires multi-model-server as a peer dependency and integration with a serving entrypoint script.
- High install friction due to no runtime dependencies and requirement to be integrated into a Docker build process rather than installed standalone.
- Repository is archived and abandoned as of 2023-11-20, with no active maintenance.
License · maintenance · safety
Apache License 2.0 (permissive) — Licensed under Apache License 2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
last release 2023-10-25 (1024 days) · last repo commit 2023-11-20 · 413 stars · archived
0 known vulnerabilities (OSV.dev, 2026-08-14) · 315,403 downloads/mo, #7,687 on PyPI
Alternatives
Verify before relying
# In Dockerfile:
RUN pip3 install multi-model-server sagemaker-inference
# In handler implementation:
from sagemaker_inference import content_types, decoder, encoder
from sagemaker_inference.default_handler_service import DefaultHandlerService
class MyHandler(DefaultHandlerService):
pass- Whether the abandoned repository status affects long-term compatibility with current SageMaker versions or Python releases beyond 3.10.
- Whether Multi Model Server remains actively maintained and compatible with modern deployment environments.
- Specific performance characteristics or throughput limits for the serving stack.
What it is and what it does
SageMaker Inference Toolkit is a library that packages a model serving stack for deployment on Amazon SageMaker. It abstracts the complexity of setting up inference endpoints by providing handler interfaces for model loading, input preprocessing, prediction, and output serialization. The toolkit is designed to be embedded in Docker containers and works with Multi Model Server to handle incoming inference requests.
The library is intended for developers building custom inference containers for SageMaker. It provides base classes and utilities (decoder, encoder, content type handlers) that you extend to define how your specific model should be loaded and served. However, the repository has been archived and is no longer actively maintained as of late 2023, which means it may not receive updates for compatibility with newer Python versions or SageMaker features.
Use it for
- Build a custom Docker inference container for a PyTorch or TensorFlow model to deploy on SageMaker.
- Implement multi-model serving where a single container handles multiple model versions or types.
- Add standardized input/output handling (JSON, CSV, NPZ) to a model serving pipeline.
- Create a handler service that integrates with SageMaker's model server lifecycle (initialization and request handling).
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, but with caution.
Install if you are actively deploying models to Amazon SageMaker and need a structured framework for containerized inference. The permissive Apache 2.0 license poses no barrier. However, the abandoned repository status (last commit 2023-11-20) means you should verify compatibility with your target SageMaker version and Python runtime before committing to production use. If you are starting a new project, check whether SageMaker's prebuilt containers or newer alternatives better suit your needs.
Install
sagemaker-inference on PyPI
Before you install
High install friction due to no runtime dependencies and requirement to be integrated into a Docker build process rather than installed standalone. Repository is archived and abandoned as of 2023-11-20, with no active maintenance.
Must be installed within a Docker container build; requires multi-model-server as a peer dependency and integration with a serving entrypoint script.
License in practice
Licensed under Apache License 2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
Quickstart
# In Dockerfile:
RUN pip3 install multi-model-server sagemaker-inference
# In handler implementation:
from sagemaker_inference import content_types, decoder, encoder
from sagemaker_inference.default_handler_service import DefaultHandlerService
class MyHandler(DefaultHandlerService):
pass
Verify before relying
- Whether the abandoned repository status affects long-term compatibility with current SageMaker versions or Python releases beyond 3.10.
- Whether Multi Model Server remains actively maintained and compatible with modern deployment environments.
- Specific performance characteristics or throughput limits for the serving stack.
Package facts
| License | Apache License 2.0 permissive |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Abandoned 1,024 days since the last release |
| Last repo commit | repository archived |
| First released | |
| Downloads | 315,403 / month, #7,687 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseNatural Language :: EnglishProgramming Language :: PythonProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: sagemaker_inference-1.10.1.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ml inference docker container”
- sagemaker-inferenceProvides a model serving stack for deploying machine learning models…
- cogCog packages machine learning models into production-ready Docker…
- sagemaker-containersProvides tools to build Docker containers compatible with Amazon…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also kserve · sagemaker-schema-inference-artifacts · sagemaker-serve · model-hosting-container-standards · sagemaker-containers · sagemaker-training · bentoml · multi-model-server · model-archiver · sagemaker