llguidance
Bindings for the Low-level Guidance (llguidance) Rust library for use within Guidance
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and solves a critical problem in LLM inference—ensuring structured output without sacrificing speed. It is production-ready, widely integrated into major frameworks, and backed by peer-reviewed research. Install if you need deterministic structured output from LLMs in any inference context.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; compiled wheels available for common platforms but installation may require compilation on unsupported architectures.
- Medium install friction due to compiled wheels across multiple platforms and architectures.
- Package is actively maintained with a recent release and no known vulnerabilities.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution for commercial and private projects with minimal attribution requirements.
last release 2026-08-11 (3 days) · last repo commit 2026-08-10 · 839 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 5,546,840 downloads/mo, #2,077 on PyPI
Alternatives
Verify before relying
pip install llguidance
import llguidance
# Use with Guidance library or directly via Rust/C bindings
# See integration examples in vLLM, llama.cpp, or SGLang- Exact performance characteristics (50μs per token claim) for specific tokenizer sizes and grammar complexity in real-world deployments.
- Compatibility matrix and tested versions for each integrated project (vLLM, llama.cpp, SGLang, etc.).
- Whether the Python package is the primary interface or primarily a binding for Rust library use.
What it is and what it does
llguidance is a Rust library with Python bindings that implements constrained decoding for large language models. It computes token masks—sets of valid next tokens—that ensure LLM output conforms to a specified grammar, JSON schema, or regular expression. The library supports multiple grammar formats including JSON schemas, regular expressions, and context-free grammars in Lark-like syntax, and can be used directly from Python, Rust, C, or C++.
The library is designed for performance: mask computation takes approximately 50μs per token for a 128k-token vocabulary with negligible startup cost. It has been integrated into major LLM inference frameworks including vLLM, llama.cpp, SGLang, and Chromium, and powers OpenAI's Structured Output feature. It uses Earley's algorithm for parsing combined with regex derivatives and trie-based token traversal to avoid the startup overhead and memory costs of pre-computed automata approaches.
Use it for
- Enforce JSON schema compliance in LLM outputs for API integrations or data pipelines requiring structured responses.
- Constrain LLM generation to valid regular expressions or domain-specific grammars in real-time inference servers.
- Integrate structured output constraints into vLLM, llama.cpp, or SGLang deployments without significant latency overhead.
- Use within the Guidance Python library to build multi-turn LLM workflows with guaranteed output format compliance.
- Embed in Chromium-based browsers to enforce JSON Schema for the Prompt API's structured output feature.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and solves a critical problem in LLM inference—ensuring structured output without sacrificing speed. It is production-ready, widely integrated into major frameworks, and backed by peer-reviewed research. Install if you need deterministic structured output from LLMs in any inference context.
Install
llguidance on PyPI
Before you install
Medium install friction due to compiled wheels across multiple platforms and architectures. Package is actively maintained with a recent release and no known vulnerabilities. Requires Python 3.10 or later.
Requires Python 3.10 or later; compiled wheels available for common platforms but installation may require compilation on unsupported architectures.
License in practice
MIT license permits unrestricted use, modification, and distribution for commercial and private projects with minimal attribution requirements.
Quickstart
pip install llguidance
import llguidance
# Use with Guidance library or directly via Rust/C bindings
# See integration examples in vLLM, llama.cpp, or SGLang
Verify before relying
- Exact performance characteristics (50μs per token claim) for specific tokenizer sizes and grammar complexity in real-world deployments.
- Compatibility matrix and tested versions for each integrated project (vLLM, llama.cpp, SGLang, etc.).
- Whether the Python package is the primary interface or primarily a binding for Rust library use.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 3 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 5,546,840 / month, #2,077 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: llguidance-1.8.0-cp314-cp314t-macosx_10_12_x86_64.whl; llguidance-1.8.0-cp314-cp314t-macosx_11_0_arm64.whl; llguidance-1.8.0-cp314-cp314t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; llguidance-1.8.0-cp314-cp314t-manylinux_2_17_i686.manylinux2014_i686.whl; llguidance-1.8.0-cp314-cp314t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; llguidance-1.8.0-cp314-cp314t-win32.whl; llguidance-1.8.0-cp314-cp314t-win_amd64.whl; llguidance-1.8.0-cp314-cp314t-win_arm64.whl; llguidance-1.8.0-cp39-abi3-macosx_10_12_x86_64.whl; llguidance-1.8.0-cp39-abi3-macosx_11_0_arm64.whl; llguidance-1.8.0-cp39-abi3-manylinux_2_31_aarch64.whl; llguidance-1.8.0-cp39-abi3-manylinux_2_31_x86_64.whl; llguidance-1.8.0-cp39-abi3-manylinux_2_34_i686.whl; llguidance-1.8.0-cp39-abi3-manylinux_2_39_riscv64.whl; llguidance-1.8.0-cp39-abi3-win32.whl; llguidance-1.8.0-cp39-abi3-win_amd64.whl; llguidance-1.8.0-cp39-abi3-win_arm64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “constrained decoding llm”
- llguidanceEnforces structured output from large language models by computing…
- xgrammarConstrains LLM text generation to follow specified grammars (JSON,…
- lm-format-enforcerConstrains language model token generation to enforce JSON Schema,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also lm-format-enforcer · xgrammar · renderers · graphrag · lark-parser · lm-eval · lark · spark-parser · vllm · ipex-llm