{"enrichment":{"faq":[{"a":"Ollama lets you deploy and serve open-weight language models on your machine with automatic GPU detection. Install Ollama, then use `ollama pull <model-name>` to download a model (e.g., `ollama pull qwen`), and `ollama run <model-name>` to start it. Ollama automatically detects and uses your GPU for acceleration, making inference fast without cloud dependencies.","q":"How do I run local LLM models with Ollama?"},{"a":"Yes. Ollama exposes an OpenAI-compatible API endpoint that lets you integrate local models into applications expecting standard OpenAI interfaces. Once a model is running, you can query it via the local endpoint, making it easy to swap cloud APIs for self-hosted inference without rewriting client code.","q":"Does Ollama provide an OpenAI-compatible API endpoint?"},{"a":"Ollama supports flexible context window configuration for agent sessions. When pulling or running a model, you can adjust parameters to control context length based on your workload requirements. This lets you balance between model capability and memory usage for different agent scenarios.","q":"How do I configure context length for Ollama models?"},{"a":"Yes. Ollama integrates with PenguinHarness for model registration, allowing you to manage local models within the PenguinHarness agent framework. Register your running Ollama endpoint so agents can discover and use your locally-served models.","q":"Can I register Ollama models with PenguinHarness?"},{"a":"Ollama and vLLM are both local model serving options. Ollama emphasizes ease of setup with automatic GPU detection and a simple pull-and-run workflow, while vLLM focuses on high-throughput inference optimization. Choose Ollama for quick deployment; consider vLLM if you need advanced batching or performance tuning.","q":"What's the difference between Ollama and vLLM for local serving?"},{"a":"Ollama automatically detects and uses your GPU for acceleration during setup. Install Ollama on Windows, Mac, or Linux, and it will configure GPU support without manual intervention. Verify GPU usage by checking logs when you run a model\u2014Ollama will report which device (GPU or CPU) is being used.","q":"How do I set up Ollama GPU acceleration?"}],"shadow_tags":["local-inference","self-hosted-llm","gpu-acceleration","model-deployment","openai-compatible","offline-ai","model-serving","context-management"],"summary_rewrite":"Ollama lets you deploy and serve open-weight language models on your machine with automatic GPU detection. It exposes an OpenAI-compatible API endpoint, integrates with PenguinHarness for model registration, and supports flexible context window configuration for agent workloads."},"files":[{"bytes":3501,"path":"packages/skills/skills/ollama/SKILL.md","sha256":"53f822948738d7360af80fa50da6f2d9e963a230fca3b9020ad254d56859a8cb","url":"https://skillfed.io/files/Prism-Shadow/penguin-harness/ollama/c64c4f37/SKILL.md"}],"id":"Prism-Shadow/penguin-harness/ollama","links":{"html":"https://skillfed.io/Prism-Shadow/penguin-harness/ollama","md":"https://skillfed.io/Prism-Shadow/penguin-harness/ollama.md","repo":"https://github.com/Prism-Shadow/penguin-harness"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":24,"language":"TypeScript","last_updated":"2026-07-27","license":"Apache-2.0","name":"ollama","publisher":"Prism-Shadow","stars":205},"relations":{"similar":[{"id":"Prism-Shadow/penguin-harness/vllm"},{"id":"Prism-Shadow/penguin-harness/llamafactory"},{"id":"Prism-Shadow/penguin-harness/penguin-cli"},{"id":"Prism-Shadow/penguin-harness/agenthub-models"},{"id":"Prism-Shadow/penguin-harness/penguin-sdk"},{"id":"BagelHole/DevOps-Security-Agent-Skills/vllm-server"},{"id":"moltis-org/moltis/llama-cpp"},{"id":"Prism-Shadow/penguin-harness/agent-evaluation"},{"id":"Aradotso/trending-skills/dflash-mlx-speculative-decoding"},{"id":"BagelHole/DevOps-Security-Agent-Skills/mac-mini-llm-lab"}]},"slug":{"owner":"Prism-Shadow","repo":"penguin-harness","skill":"ollama"},"version":"c64c4f37"}
