langextract
LangExtract: A library for extracting structured data from language models
Decision gist · record as of 2026-08-14
Yes. LangExtract is actively maintained, has low install friction, carries no known vulnerabilities, and offers a permissive Apache-2.0 license. It solves a concrete problem—grounded structured extraction from unstructured text—with built-in visualization and multi-provider LLM support. Install it if you need to extract and ground structured data from documents with traceability and schema enforcement. Requires Python 3.10+ and an LLM provider (cloud or local).AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10+.
- Cloud model usage (e.g., Gemini) requires an API key; local models via Ollama are also supported.
- Low install friction with a pure-wheel distribution.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions, making it suitable for most production and research applications.
last release 2026-07-02 (43 days) · last repo commit 2026-08-11 · 38,376 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 407,057 downloads/mo, #6,890 on PyPI
Alternatives
Verify before relying
import langextract as lx
prompt = "Extract characters and emotions in order of appearance."
examples = [lx.data.ExampleData(
text="ROMEO. But soft! What light through yonder window breaks?",
extractions=[lx.data.Extraction(
extraction_class="character",
extraction_text="ROMEO",
attributes={"emotional_state": "wonder"}
)]
)]
result = lx.extract(
text_or_documents="Your input text here",
prompt_description=prompt,
examples=examples,
model_id="gemini-3.5-flash"
)- Whether the package handles rate limiting gracefully for high-volume extraction workloads.
- Performance characteristics and throughput for documents exceeding typical LLM context windows.
- Exact behavior when extractions cannot be grounded in source text beyond the documented char_interval = None filtering.
What it is and what it does
LangExtract is a Python library that uses large language models to extract structured information from unstructured text—such as clinical notes, reports, or literary documents—based on user-defined extraction rules and few-shot examples. It maps every extracted entity back to its exact location in the source text, enabling visual verification and traceability. The library handles long documents through optimized chunking and parallel processing, supports multiple LLM providers (Google Gemini, OpenAI, local Ollama), and generates interactive HTML visualizations to review thousands of extracted entities in their original context.
The core workflow involves defining a prompt that describes what to extract, providing high-quality examples to guide the model, running the extraction on your input text, and optionally visualizing results. LangExtract enforces consistent output schemas and detects when the model extracts content from examples rather than the source text. It is designed for domains where precise source grounding and schema-constrained outputs matter—medical records, legal documents, research papers, or any scenario where traceability and structured output are critical.
Use it for
- Extract medications, diagnoses, and clinical findings from medical notes with exact source citations for audit and verification.
- Parse radiology or pathology reports to structure unstructured narrative text into standardized data fields.
- Identify characters, relationships, and emotions from literary texts with interactive visualization for literary analysis.
- Annotate and label large document collections for machine learning training datasets with built-in grounding and conflict detection.
- Extract key entities and attributes from legal contracts or regulatory documents while maintaining source traceability.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
LangExtract is actively maintained, has low install friction, carries no known vulnerabilities, and offers a permissive Apache-2.0 license. It solves a concrete problem—grounded structured extraction from unstructured text—with built-in visualization and multi-provider LLM support. Install it if you need to extract and ground structured data from documents with traceability and schema enforcement. Requires Python 3.10+ and an LLM provider (cloud or local).
Install
langextract on PyPI
Before you install
Low install friction with a pure-wheel distribution. Active maintenance with recent commits and high repository engagement (38376 stars). Requires Python 3.10+. Depends on 17 runtime packages including google-genai, google-cloud-storage, and standard data-processing libraries.
Requires Python 3.10+. Cloud model usage (e.g., Gemini) requires an API key; local models via Ollama are also supported.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions, making it suitable for most production and research applications.
Quickstart
import langextract as lx
prompt = "Extract characters and emotions in order of appearance."
examples = [lx.data.ExampleData(
text="ROMEO. But soft! What light through yonder window breaks?",
extractions=[lx.data.Extraction(
extraction_class="character",
extraction_text="ROMEO",
attributes={"emotional_state": "wonder"}
)]
)]
result = lx.extract(
text_or_documents="Your input text here",
prompt_description=prompt,
examples=examples,
model_id="gemini-3.5-flash"
)
Verify before relying
- Whether the package handles rate limiting gracefully for high-volume extraction workloads.
- Performance characteristics and throughput for documents exceeding typical LLM context windows.
- Exact behavior when extractions cannot be grounded in source text beyond the documented char_interval = None filtering.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 17 packagesabsl-pyaiohttpasync_timeoutexceptiongroupgoogle-genaigoogle-cloud-storageml-collectionsmore-itertoolsnumpypandaspydanticpython-dotenvPyYAMLregexrequeststqdmtyping-extensions |
| Maintenance | Actively maintained 43 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 407,057 / month, #6,890 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: langextract-1.6.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “llm structured information extraction”
- langextractLangExtract uses LLMs to extract and ground structured information…
- trustcallTrustcall helps LLMs reliably generate and update complex JSON…
- MainContentExtractorExtracts main content from HTML pages and outputs it in HTML, text,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also instructor · graphrag · llm · gliner2 · gliner · MainContentExtractor · trustcall · llama-index · gitingest · doclang