$npx skillfedfor your agent

bentoml

BentoML: The easiest way to serve AI apps and models

Worth itPyPI LibrariesReleased May 2026274.7K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — bentoml-1.4.39-py3-none-any.whl
v1.4.39 · released 2026-05-07 · Python >=3.9 · 42 runtime deps: a2wsgi, aiohttp, aiohttp-asgi-connector, aiosqlite, attrs, cattrs, click-option-group, click

Yes. BentoML is actively maintained, permissively licensed, and widely adopted. It solves a real problem—reducing boilerplate for production model serving—and has no known vulnerabilities. The 42 dependencies are expected for a full-featured serving framework. Install if you need to deploy ML models as scalable APIs; skip if you only need lightweight HTTP routing or are committed to a different serving paradigm.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python ≥3.9.
  • Local testing requires the ML framework (e.g., PyTorch, TensorFlow) and any model dependencies to be installed separately in your environment.
  • Low install friction with a pure Python wheel.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for proprietary applications.

last release 2026-05-07 (99 days) · last repo commit 2026-08-03 · 8,788 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 274,689 downloads/mo, #8,188 on PyPI

Verify before relying

pip install bentoml

import bentoml

@bentoml.service()
class MyModel:
    @bentoml.api()
    def predict(self, data: str) -> str:
        return f"Result: {data}"

# Run: bentoml serve
  • Whether the 42 runtime dependencies are all mandatory or some are optional/conditional based on use case.
  • Performance characteristics under high concurrency or GPU load compared to other serving frameworks.
  • Stability and breaking-change frequency of the API across minor versions.
Same gist for agents: .md · .json

What it is and what it does

BentoML is a production-ready framework for turning ML model inference code into scalable REST APIs. You define a service class with decorated methods, and BentoML handles HTTP routing, request batching, containerization, and deployment orchestration. It abstracts away boilerplate for common serving patterns—dynamic batching, multi-model composition, GPU utilization, and observability—while remaining agnostic to the underlying ML framework (PyTorch, TensorFlow, scikit-learn, etc.).

The framework is designed for both local development and cloud deployment. Locally, you run `bentoml serve` to test your API. For production, you package your service into a standardized Bento artifact, then generate a Docker image with `bentoml containerize`. It also integrates with BentoCloud for managed hosting. The dependency footprint is substantial (42 runtime packages including aiohttp, pydantic, and opentelemetry), reflecting its full-featured nature as an application framework rather than a lightweight library.

Use it for

  • Deploy a fine-tuned LLM or vision model as a REST API without writing HTTP boilerplate or managing async concurrency yourself.
  • Build a multi-model inference pipeline where requests flow through sequential or parallel model stages with automatic batching.
  • Package a model service with all dependencies and environment config into a reproducible Docker image for on-premise or cloud deployment.
  • Add observability and monitoring to model inference with built-in OpenTelemetry and Prometheus instrumentation.
  • Optimize GPU utilization across multiple concurrent inference requests using adaptive batching and worker parallelization.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

BentoML is actively maintained, permissively licensed, and widely adopted. It solves a real problem—reducing boilerplate for production model serving—and has no known vulnerabilities. The 42 dependencies are expected for a full-featured serving framework. Install if you need to deploy ML models as scalable APIs; skip if you only need lightweight HTTP routing or are committed to a different serving paradigm.

Install

bentoml on PyPI

Before you install

Low install friction with a pure Python wheel. Active maintenance with recent commits and a large community. Requires Python ≥3.9 and brings 42 runtime dependencies including aiohttp, pydantic, and opentelemetry instrumentation—manageable for a framework of this scope.

Requires Python ≥3.9. Local testing requires the ML framework (e.g., PyTorch, TensorFlow) and any model dependencies to be installed separately in your environment.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for proprietary applications.

Quickstart

pip install bentoml

import bentoml

@bentoml.service()
class MyModel:
    @bentoml.api()
    def predict(self, data: str) -> str:
        return f"Result: {data}"

# Run: bentoml serve

Verify before relying

  • Whether the 42 runtime dependencies are all mandatory or some are optional/conditional based on use case.
  • Performance characteristics under high concurrency or GPU load compared to other serving frameworks.
  • Stability and breaking-change frequency of the API across minor versions.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
42 packages
a2wsgiaiohttpaiohttp-asgi-connectoraiosqliteattrscattrsclick-option-groupclickcloudpicklefsspechttpxhttpx-wsjinja2kantokunumpynvidia-ml-pyopentelemetry-apiopentelemetry-instrumentation-aiohttp-clientopentelemetry-instrumentation-asgiopentelemetry-instrumentationopentelemetry-sdkopentelemetry-semantic-conventionsopentelemetry-util-httppackagingpathspecpip-requirements-parserprometheus-clientpsutilpydanticpython-dateutil
MaintenanceActively maintained 99 days since the last release
Last repo commit
First released
Downloads274,689 / month, #8,188 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software Development :: Libraries

Evidence: bentoml-1.4.39-py3-none-any.whl

Tags

Capabilities
model serving frameworkml inference apiai model deploymentbatch servingdocker containerization ml
Topics
model-servingmlopsinference-api
PyPI keywords
BentoMLCompound AI SystemsLLMOpsMLOpsModel DeploymentModel InferenceModel Serving

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “model serving framework”

  • bentomlBentoML is a Python framework for building and deploying REST APIs…
  • sagemaker-inferenceProvides a model serving stack for deploying machine learning models…
  • sglangSGLang is a serving framework that runs large language models and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also tensorflow-serving-api · sagemaker-inference · sagemaker-serve · mlserver · truss · kserve · vllm · multi-model-server · litserve · zenml