$npx skillfedfor your agent

foundry-local-sdk

Foundry Local Manager Python SDK: Control-plane SDK for Foundry Local.

With conditionsPyPI Software DevelopmentReleased Jul 2026298.9K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — foundry_local_sdk-1.2.4-py3-none-any.whl
v1.2.4 · released 2026-07-30 · Python >=3.11 · 10 runtime deps: pydantic, requests, openai, foundry-local-core, onnxruntime-gpu, onnxruntime, onnxruntime-core, onnxruntime-genai-cuda

Yes, if you need local LLM inference with OpenAI-compatible APIs and don't require production-grade stability. The package is actively maintained, has low install friction, permissive licensing, and zero known vulnerabilities. It is in Alpha status (Development Status :: 3), so expect API changes and incomplete documentation; verify that the model catalog and hardware support match your target deployment before committing to it.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.11 or later.
  • Platform-specific native binaries (CUDA-enabled on Linux x86_64, CPU/WebGPU on other platforms) are installed as wheel dependencies; WinML variant available for Windows only.
  • Low friction wheel install with platform-specific native dependencies (onnxruntime variants, foundry-local-core).

License · maintenance · safety

MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal restrictions.

last release 2026-07-30 (15 days) · last repo commit 2026-08-14 · 2,511 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 298,878 downloads/mo, #7,867 on PyPI

Verify before relying

pip install foundry-local-sdk

from foundry_local_sdk import Configuration, FoundryLocalManager

config = Configuration(app_name="MyApp")
FoundryLocalManager.initialize(config)
manager = FoundryLocalManager.instance
model = manager.catalog.get_model("phi-3.5-mini")
model.load()
client = model.get_chat_client()
response = client.complete_chat([{"role": "user", "content": "Why is the sky blue?"}])
print(response.choices[0].message.content)
  • Whether model catalog is pre-populated or requires external download/configuration
  • Typical memory footprint and hardware requirements for common models
  • Latency characteristics for inference on different hardware (CPU vs GPU)
  • Whether streaming chat and embeddings APIs are fully OpenAI-compatible or have known differences
Same gist for agents: .md · .json

What it is and what it does

Foundry Local SDK is a Python control plane for running large language models, embeddings, and speech-to-text inference entirely on your local machine. It wraps a native Foundry Local Core library (via ctypes FFI) and provides a model catalog, download/cache management, and OpenAI-compatible APIs for chat completions (streaming and non-streaming), tool calling, embeddings, and Whisper-based audio transcription. The package selects the appropriate ONNX Runtime variant per platform (CUDA on Linux x86_64, CPU/WebGPU elsewhere) and publishes a separate WinML variant for Windows.

You initialize a singleton manager with a configuration, discover models from the catalog, download and load them into memory, then interact via chat or embedding clients. The SDK handles model lifecycle (load, unload, cache), execution provider discovery and registration, and optional HTTP endpoints for multi-process scenarios. It is actively maintained, recently released, and has no known security vulnerabilities.

Use it for

  • Run private LLM inference on-device without sending data to cloud APIs or paying per-token fees
  • Build chat applications with streaming responses using OpenAI-compatible client code
  • Generate text embeddings locally for semantic search or RAG workflows
  • Transcribe audio files or streams using Whisper models without external services
  • Prototype and test LLM-powered features before deploying to production infrastructure
  • Call functions/tools from LLM responses in a local environment with full control

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need local LLM inference with OpenAI-compatible APIs and don't require production-grade stability.

The package is actively maintained, has low install friction, permissive licensing, and zero known vulnerabilities. It is in Alpha status (Development Status :: 3), so expect API changes and incomplete documentation; verify that the model catalog and hardware support match your target deployment before committing to it.

Install

foundry-local-sdk on PyPI

Before you install

Low friction wheel install with platform-specific native dependencies (onnxruntime variants, foundry-local-core). Active maintenance with recent release (15 days old); requires Python 3.11+.

Requires Python 3.11 or later. Platform-specific native binaries (CUDA-enabled on Linux x86_64, CPU/WebGPU on other platforms) are installed as wheel dependencies; WinML variant available for Windows only.

License in practice

MIT license permits unrestricted use, modification, and distribution with minimal restrictions.

Quickstart

pip install foundry-local-sdk

from foundry_local_sdk import Configuration, FoundryLocalManager

config = Configuration(app_name="MyApp")
FoundryLocalManager.initialize(config)
manager = FoundryLocalManager.instance
model = manager.catalog.get_model("phi-3.5-mini")
model.load()
client = model.get_chat_client()
response = client.complete_chat([{"role": "user", "content": "Why is the sky blue?"}])
print(response.choices[0].message.content)

Verify before relying

  • Whether model catalog is pre-populated or requires external download/configuration
  • Typical memory footprint and hardware requirements for common models
  • Latency characteristics for inference on different hardware (CPU vs GPU)
  • Whether streaming chat and embeddings APIs are fully OpenAI-compatible or have known differences

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.11
Install frictionLow. Pure-Python wheel
Runtime dependencies
10 packages
pydanticrequestsopenaifoundry-local-coreonnxruntime-gpuonnxruntimeonnxruntime-coreonnxruntime-genai-cudaonnxruntime-genaionnxruntime-genai-core
MaintenanceActively maintained 15 days since the last release
Last repo commit
First released
Downloads298,878 / month, #7,867 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaIntended Audience :: DevelopersProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules

Evidence: foundry_local_sdk-1.2.4-py3-none-any.whl

Tags

Capabilities
local AI model inferenceoffline language model runneropenai-compatible chat apimodel management and discoverylocal embeddings generationspeech-to-text transcriptionfoundry local sdk
Topics
local-inferencellm-sdkoffline-ai

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “offline language model runner”

  • foundry-local-sdkA Python SDK for discovering, downloading, loading, and running local…
  • argostranslateArgos Translate is an open-source offline neural machine translation…
  • nvidia-lm-evalEvaluates language models against standardized benchmarks (MMLU,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Software Development packages

typing-extensions Worth it
PyPI · Software Development · released Jul 2026

Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.

PSF-2.0pure Python · 3.9+
1.9Bdownloads / mo
numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
fastapi Worth it
PyPI · Software Development · released Jul 2026

FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.

MITpure Python · 3.10+
568.6Mdownloads / mo
annotated-doc With conditions
PyPI · Software Development · released Jul 2026

Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.

MITpure Python · 3.9+
456.2Mdownloads / mo
typer Worth it
PyPI · Software Development · released Aug 2026

Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.

Install it if you are building CLIs in Python.

MITpure Python · 3.10+
369.3Mdownloads / mo
distlib With conditions
PyPI · Software Development · released Jun 2026

Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.

permissive licensepure Python
323.3Mdownloads / mo

See also azure-ai-inference · foundry-platform-sdk · zai-sdk · rev-ai · SpeechRecognition · azure-ai-projects · agent-framework-foundry-local · langchain-azure-ai · cloudfoundry-client · cfenv