xinference-client
Client for Xinference
Decision gist · record as of 2026-08-14
Yes, if you are running a Xinference server and need a Python client to access it. The package is actively maintained, has low install friction, carries no known vulnerabilities, and uses a permissive license. It is not useful standalone—it requires a running Xinference server to connect to.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a running Xinference server accessible at the specified URL (e.g., http://localhost:9997).
- Low friction installation with a pure Python wheel and four common dependencies (requests, aiohttp, typing-extensions, pydantic).
- Actively maintained with recent releases; last commit 2026-07-31.
License · maintenance · safety
Apache License 2.0 (permissive) — Apache License 2.0 is permissive, allowing commercial and private use with minimal restrictions—suitable for most projects without licensing concerns.
last release 2026-07-31 (14 days) · last repo commit 2026-07-31 · 9 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 204,341 downloads/mo, #9,610 on PyPI
Alternatives
Verify before relying
pip install xinference-client
from xinference_client import RESTfulClient as Client
client = Client("http://localhost:9997")
model_uid = client.launch_model(model_name="chatglm2")
model = client.get_model(model_uid)
model.chat("What is the largest animal?", chat_history=[], generate_config={"max_tokens": 1024})- Whether the package supports async/await patterns beyond aiohttp dependency presence.
- Supported Xinference server versions and compatibility guarantees.
- Whether model launch/retrieval operations have timeout or retry mechanisms.
What it is and what it does
xinference-client is a Python HTTP client for the Xinference model serving platform, allowing developers to programmatically launch, retrieve, and query language models running on a Xinference server. It wraps REST API calls into a Python-friendly interface, handling model lifecycle management (launch, get) and inference requests (chat completions with configurable generation parameters).
The package is built on standard HTTP libraries (requests, aiohttp) and integrates with pydantic for data validation. It's designed for scenarios where models are hosted on a separate Xinference server and accessed remotely, rather than for local model execution. The client abstracts away HTTP details and returns structured responses (chat completion objects with token counts, finish reasons, and assistant messages).
Use it for
- Launch and query language models on a remote Xinference server from a Python application without direct model management.
- Build chatbot or conversational AI applications that delegate inference to a centralized Xinference deployment.
- Integrate model inference into data pipelines or batch processing workflows using a simple client API.
- Prototype or test different language models by swapping model names without changing client code.
- Monitor and manage model lifecycle (launch, retrieve state) programmatically across multiple applications.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are running a Xinference server and need a Python client to access it.
The package is actively maintained, has low install friction, carries no known vulnerabilities, and uses a permissive license. It is not useful standalone—it requires a running Xinference server to connect to.
Install
xinference-client on PyPI
Before you install
Low friction installation with a pure Python wheel and four common dependencies (requests, aiohttp, typing-extensions, pydantic). Actively maintained with recent releases; last commit 2026-07-31.
Requires a running Xinference server accessible at the specified URL (e.g., http://localhost:9997).
License in practice
Apache License 2.0 is permissive, allowing commercial and private use with minimal restrictions—suitable for most projects without licensing concerns.
Quickstart
pip install xinference-client
from xinference_client import RESTfulClient as Client
client = Client("http://localhost:9997")
model_uid = client.launch_model(model_name="chatglm2")
model = client.get_model(model_uid)
model.chat("What is the largest animal?", chat_history=[], generate_config={"max_tokens": 1024})
Verify before relying
- Whether the package supports async/await patterns beyond aiohttp dependency presence.
- Supported Xinference server versions and compatibility guarantees.
- Whether model launch/retrieval operations have timeout or retry mechanisms.
Package facts
| License | Apache License 2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesrequestsaiohttptyping-extensionspydantic |
| Maintenance | Actively maintained 14 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 204,341 / month, #9,610 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Operating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonTopic :: Software Development :: Libraries |
Evidence: xinference_client-3.1.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “xinference client library”
- xinference-clientA Python client library for connecting to and managing Xinference…
- conjure-python-clientA lightweight HTTP client library built on requests that handles RPC…
- paypalhttpPayPalHttp is a generic HTTP client library designed to work with…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also ollama · twentyc.rpc · cvprac · anthropic · restfly · ai21 · python-http-client · reka-api · langchain-nvidia-ai-endpoints · ogx_open_client