$npx skillfedfor your agent

smg-grpc-servicer

SMG gRPC servicer implementations for LLM inference engines (vLLM, MLX, TokenSpeed, SGLang)

Worth itPyPI Artificial IntelligenceReleased Jul 20261.1M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — smg_grpc_servicer-0.8.0-py3-none-any.whl
v0.8.0 · released 2026-07-30 · Python >=3.10 · 4 runtime deps: smg-grpc-proto, grpcio, grpcio-reflection, grpcio-health-checking

Yes. The package solves a real problem—exposing LLM inference engines as gRPC services—with low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install it if you need to serve LLM inference over gRPC to remote clients or integrate multiple inference backends into a distributed system.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.10.
  • Each backend (vLLM, MLX, TokenSpeed, SGLang) must be installed separately; the servicer itself provides only the gRPC bridge.
  • Low friction: pure Python wheel with four runtime dependencies (smg-grpc-proto, grpcio, grpcio-reflection, grpcio-health-checking).

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions.

last release 2026-07-30 (15 days) · last repo commit 2026-08-14 · 461 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,118,607 downloads/mo, #4,344 on PyPI

Verify before relying

# For vLLM:
pip install smg-grpc-servicer[vllm]
vllm serve meta-llama/Llama-2-7b-hf --grpc

# For MLX:
pip install smg-grpc-servicer[mlx]
python -m smg_grpc_servicer.mlx --model meta-llama/Llama-2-7b-hf --host 0.0.0.0 --port 50051
  • Whether gRPC service discovery and reflection work out-of-the-box or require additional configuration.
  • Performance characteristics and latency overhead of the gRPC bridge relative to native backend APIs.
  • Load-balancing and concurrency limits when serving multiple concurrent inference requests.
Same gist for agents: .md · .json

What it is and what it does

smg-grpc-servicer is a bridge library that wraps four popular open-source LLM inference engines—vLLM, MLX, TokenSpeed, and SGLang—behind a unified gRPC interface. Instead of calling each engine's native Python API directly, you run one of the supported backends with gRPC enabled (or via the servicer's CLI modules), and clients connect over gRPC to send inference requests and receive responses.

The package isolates backend dependencies using optional extras ([vllm], [mlx], [sglang]) so you only install what you need; TokenSpeed integrates via an external runtime. It depends on grpcio, grpcio-reflection, and grpcio-health-checking for service discovery and health checks. The library is actively maintained, supports Python 3.10–3.13, and carries no known vulnerabilities.

Use it for

  • Serve a vLLM model over gRPC to multiple remote clients without exposing HTTP endpoints.
  • Run MLX inference on macOS and expose it as a gRPC service for distributed workloads.
  • Build a microservice architecture where LLM inference is decoupled from application logic via gRPC.
  • Load-balance inference requests across multiple backend instances using gRPC service discovery.
  • Integrate TokenSpeed or SGLang inference into a polyglot system using language-agnostic gRPC clients.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package solves a real problem—exposing LLM inference engines as gRPC services—with low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install it if you need to serve LLM inference over gRPC to remote clients or integrate multiple inference backends into a distributed system.

Install

smg-grpc-servicer on PyPI

Before you install

Low friction: pure Python wheel with four runtime dependencies (smg-grpc-proto, grpcio, grpcio-reflection, grpcio-health-checking). Active maintenance—released 15 days ago with 461 repository stars.

Requires Python >=3.10. Each backend (vLLM, MLX, TokenSpeed, SGLang) must be installed separately; the servicer itself provides only the gRPC bridge.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions.

Quickstart

# For vLLM:
pip install smg-grpc-servicer[vllm]
vllm serve meta-llama/Llama-2-7b-hf --grpc

# For MLX:
pip install smg-grpc-servicer[mlx]
python -m smg_grpc_servicer.mlx --model meta-llama/Llama-2-7b-hf --host 0.0.0.0 --port 50051

Verify before relying

  • Whether gRPC service discovery and reflection work out-of-the-box or require additional configuration.
  • Performance characteristics and latency overhead of the gRPC bridge relative to native backend APIs.
  • Load-balancing and concurrency limits when serving multiple concurrent inference requests.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
smg-grpc-protogrpciogrpcio-reflectiongrpcio-health-checking
MaintenanceActively maintained 15 days since the last release
Last repo commit
First released
Downloads1,118,607 / month, #4,344 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Operating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13

Evidence: smg_grpc_servicer-0.8.0-py3-none-any.whl

Tags

Capabilities
grpc llm inference servervllm grpc servicerllm model serving grpcdistributed llm inferencegrpc language model endpointmlx tokenspeed sglang grpcllm inference bridge
Topics
llm-inferencegrpc-bridgemodel-serving

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “grpc llm inference server”

  • smg-grpc-servicerProvides gRPC servicer implementations that expose LLM inference…
  • smg-grpc-protoProvides pre-compiled Python gRPC stubs for protocol buffers used by…
  • tritonclienttritonclient is a Python client library for communicating with Triton…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also smg-grpc-proto · vllm · vllm-tpu · mlserver · sgl-kernel · sglang · tensorflow-serving-api · jina · sglang-router · codeshield