inference-cli
With no prior knowledge of machine learning or device-specific deployment, you can deploy a computer vision model to a range of devices and environments using Roboflow Inference CLI.
Decision gist · record as of 2026-08-14
Yes, if you need a straightforward CLI for deploying Roboflow-compatible vision models locally or to edge devices. Low install friction, active maintenance, and no known vulnerabilities make it reliable. However, verify that the GPL-3.0 and AGPL-3.0 licenses on bundled models (YOLOv5, YOLOv8) align with your project's licensing requirements before committing to production use.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Docker must be installed and running on your machine; Python 3.10 or later required.
- Low install friction with a pure-Python wheel.
- Active maintenance as of 2026-08-14 with 2416 repository stars.
License · maintenance · safety
permissive license (permissive) — Distributed under Apache 2.0, permissive and commercial-friendly. Note that individual models bundled with inference (YOLOv5, YOLOv8) carry AGPL-3.0 or GPL-3.0 licenses; verify compatibility with your use case.
last release 2026-08-14 (0 days) · last repo commit 2026-08-14 · 2,416 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 341,369 downloads/mo, #7,404 on PyPI
Alternatives
Verify before relying
pip install inference-cli
inference server start --port 9001
inference infer ./image.jpg --project-id my-project --model-version 1 --api-key my-api-key- Whether Docker must be pre-installed or if the CLI handles installation automatically.
- Whether the tool supports GPU inference without additional NVIDIA driver setup beyond nvidia-ml-py.
- Performance characteristics and latency for typical inference workloads on different device types.
What it is and what it does
Inference CLI is a lightweight command-line interface for deploying and running computer vision models locally or via Roboflow's hosted API. It wraps the parent inference package to provide server management (start, stop, status) and single-image inference commands without requiring custom Docker image building. The tool automatically detects your device (x86 CPU, ARM64, or NVIDIA GPU) and pulls the appropriate Docker image, handling the containerization details transparently.
You use it to start a local inference server on a specified port, then send images to it for predictions in JSON format. It supports object detection, classification, instance segmentation, and foundation models (CLIP, SAM), and can route requests to Roboflow's serverless hosted API instead. The CLI depends on 21 runtime packages including docker, click, typer for command parsing, opencv-python and pillow for image handling, and system introspection tools (nvidia-ml-py, py-cpuinfo) to select the right inference backend.
Use it for
- Start a local inference server on your machine and run batch inference on images without writing Python code.
- Deploy YOLOv5 or YOLOv8 or custom Roboflow models to edge devices (ARM64, x86) using the CLI's automatic device detection.
- Run inference on a single image via command line and parse the JSON output for integration into shell scripts or CI/CD pipelines.
- Switch between local and hosted inference (Roboflow serverless) by changing the --host parameter without code changes.
- Monitor and manage inference server lifecycle (start, check status, stop) from the command line in production environments.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need a straightforward CLI for deploying Roboflow-compatible vision models locally or to edge devices.
Low install friction, active maintenance, and no known vulnerabilities make it reliable. However, verify that the GPL-3.0 and AGPL-3.0 licenses on bundled models (YOLOv5, YOLOv8) align with your project's licensing requirements before committing to production use.
Install
inference-cli on PyPI
Before you install
Low install friction with a pure-Python wheel. Active maintenance as of 2026-08-14 with 2416 repository stars. Requires Python 3.10 or later and Docker for local server operation.
Docker must be installed and running on your machine; Python 3.10 or later required.
License in practice
Distributed under Apache 2.0, permissive and commercial-friendly. Note that individual models bundled with inference (YOLOv5, YOLOv8) carry AGPL-3.0 or GPL-3.0 licenses; verify compatibility with your use case.
Quickstart
pip install inference-cli
inference server start --port 9001
inference infer ./image.jpg --project-id my-project --model-version 1 --api-key my-api-key
Verify before relying
- Whether Docker must be pre-installed or if the CLI handles installation automatically.
- Whether the tool supports GPU inference without additional NVIDIA driver setup beyond nvidia-ml-py.
- Performance characteristics and latency for typical inference workloads on different device types.
Package facts
| License | permissive license permissive |
| Python support | Capped below the current Python release <3.13,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 21 packagesrequestsdockerclicktyperrichPyYAMLsupervisionopencv-pythontqdmnvidia-ml-pypy-cpuinfoaiohttpbackoffpandaspybase64pydanticurllib3tldextractdataclasses-jsonpillownumpy |
| Maintenance | Actively maintained 0 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 341,369 / month, #7,404 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyTopic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Image RecognitionTopic :: Software DevelopmentTyping :: Typed |
Evidence: inference_cli-1.4.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “computer vision inference CLI”
- inference-cliA command-line tool for running computer vision inference locally via…
- voxel51-etaETA is an extensible computer vision and machine learning analytics…
- roboflowRoboflow is a Python client for the Roboflow computer vision…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also inference-models · inference-sdk · roboflow · yolov5 · sahi · supervision · openvino · ultralytics · matrice-inference · codeflash