sagemaker-serve
SageMaker Serve package for model serving and deployment
Decision gist · record as of 2026-08-14
Yes, if you are building ML workflows on AWS SageMaker and need a unified serving layer. The package is actively maintained, has low install friction, carries a permissive Apache 2.0 license, and integrates with the SageMaker ecosystem. However, verify that its API and deployment model match your specific serving requirements before committing, as the fact sheet does not detail its exact interface or whether it supports your model types.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; AWS credentials and SageMaker access for actual deployment; torch and onnxruntime may require system-level dependencies.
- Low install friction with a pure Python wheel.
- Active maintenance status with a release 3 days old.
License · maintenance · safety
permissive license (permissive) — Apache License 2.0 (permissive) allows commercial use, modification, and distribution with minimal restrictions—primarily requiring license and copyright notice preservation in derivative works.
last release 2026-08-11 (3 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,328,127 downloads/mo, #4,051 on PyPI
Alternatives
Verify before relying
pip install sagemaker-serve
from sagemaker_serve import ...
# Deploy and serve ML models on SageMaker- Specific API surface and main classes/functions available in the package
- Whether the package supports local testing or requires live SageMaker endpoints
- Integration patterns with sagemaker-core and sagemaker-train
- Performance characteristics and scalability limits for model serving
What it is and what it does
SageMaker Serve is a Python package for deploying and serving machine learning models on Amazon SageMaker. It abstracts the complexity of model hosting by providing a unified interface to SageMaker's inference infrastructure, working alongside sagemaker-core and sagemaker-train to complete the ML lifecycle from training through production serving.
The package depends on boto3 for AWS API interaction, torch and onnxruntime for model inference, and includes testing infrastructure (pytest, tqdm, psutil) and monitoring tools (mlflow, tritonclient). It targets developers building end-to-end ML pipelines on AWS who need to move trained models into production endpoints without managing the underlying SageMaker deployment details directly.
Use it for
- Deploy trained PyTorch or ONNX models to SageMaker endpoints for real-time inference
- Manage model serving infrastructure and endpoint lifecycle on AWS
- Integrate model serving with SageMaker training pipelines for automated MLOps workflows
- Monitor and track model performance in production using MLflow integration
- Test model serving configurations locally before deploying to SageMaker
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building ML workflows on AWS SageMaker and need a unified serving layer.
The package is actively maintained, has low install friction, carries a permissive Apache 2.0 license, and integrates with the SageMaker ecosystem. However, verify that its API and deployment model match your specific serving requirements before committing, as the fact sheet does not detail its exact interface or whether it supports your model types.
Install
sagemaker-serve on PyPI
Before you install
Low install friction with a pure Python wheel. Active maintenance status with a release 3 days old. Depends on 14 runtime packages including sagemaker-core, sagemaker-train, boto3, and ML frameworks (torch, onnxruntime), which may require additional system dependencies or AWS credentials.
Requires Python 3.10 or later; AWS credentials and SageMaker access for actual deployment; torch and onnxruntime may require system-level dependencies.
License in practice
Apache License 2.0 (permissive) allows commercial use, modification, and distribution with minimal restrictions—primarily requiring license and copyright notice preservation in derivative works.
Quickstart
pip install sagemaker-serve
from sagemaker_serve import ...
# Deploy and serve ML models on SageMaker
Verify before relying
- Specific API surface and main classes/functions available in the package
- Whether the package supports local testing or requires live SageMaker endpoints
- Integration patterns with sagemaker-core and sagemaker-train
- Performance characteristics and scalability limits for model serving
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 14 packagessagemaker-coresagemaker-trainboto3botocoredeepdiffmlflowsagemaker_schema_inference_artifactspytesttqdmpsutiltritonclientonnxonnxruntimetorch |
| Maintenance | Actively maintained 3 days since the last release |
| First released | |
| Downloads | 1,328,127 / month, #4,051 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12 |
Evidence: sagemaker_serve-1.19.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “sagemaker model serving”
- sagemaker-serveProvides model serving and deployment functionality for machine…
- sagemaker-inferenceProvides a model serving stack for deploying machine learning models…
- sagemaker-containersProvides tools to build Docker containers compatible with Amazon…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also cog · model-hosting-container-standards · sagemaker · sagemaker-data-insights · sagemaker-inference · tensorflow-serving-api · truss · sagemaker-mlops · sagemaker-train · sagemaker-core