$npx skillfedfor your agent

k8s-incident

Structured runbooks and diagnostic workflows for responding to Kubernetes incidents. Covers pod failures, node health, network connectivity, storage issues, and control plane problems with prioritized troubleshooting steps and emergency actions.

k8s-incident provides runbooks and diagnostic workflows to resolve active Kubernetes outages, pod failures, node issues, and network problems.

AI-generated summary based on this skill's SKILL.md

934 177 MITupdated by rohitg00

Decision gist · record as of 2026-04-08

k8s-incident provides runbooks and diagnostic workflows to resolve active Kubernetes outages, pod failures, node issues, and network problems. Structured runbooks and diagnostic workflows for responding to Kubernetes incidents. Covers pod failures, node health, network connectivity, storage issues, and control plane problems with prioritized troubleshooting steps and emergency actions.

manual: git clone https://github.com/rohitg00/kubectl-mcp-server → cp -r kubectl-mcp-server/kubernetes-skills/claude/k8s-incident ~/.claude/skills/k8s-incident
kubernetes-skills/claude/k8s-incident/SKILL.md · version 171904db

Use it when

  • k8s-incident offers prioritized incident response workflows for cluster outages.
  • k8s-incident includes diagnostics for node health problems.

Verify before relying

Read SKILL.md below before installing (2 files). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

rohitg00/kubectl-mcp-server/k8s-incident · repository language: Python

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

How do I fix a kubernetes pod crash loop backoff?

k8s-incident provides structured runbooks for diagnosing pod crash loops. Start by checking pod events with `kubectl describe pod`, review container logs with `kubectl logs`, and verify image availability, resource limits, and application configuration. The runbook guides you through severity assessment and escalation paths for production emergencies.

What should I do during a k8s cluster down emergency?

k8s-incident offers prioritized incident response workflows for cluster outages. Begin with cluster health checks, assess control plane and node status, triage affected services by severity, and follow the structured runbook for your specific failure mode. The skill includes emergency recovery actions like pod deletion and deployment rollback.

How do I debug kubernetes node not ready issues?

k8s-incident includes diagnostics for node health problems. Inspect node status with `kubectl describe node`, check kubelet logs, verify disk pressure and memory availability, and review network connectivity. The runbook helps you quickly assess whether the issue is local to the node or cluster-wide.

What are the kubernetes incident response procedures?

k8s-incident structures emergency response around rapid triage, severity assessment, and targeted diagnostics. It covers pod failures, node issues, network problems, storage errors, and control plane failures with step-by-step troubleshooting. Each runbook includes documentation templates for post-mortems and incident timelines.

How do I diagnose kubernetes control plane issues?

k8s-incident provides diagnostics for etcd, API server, scheduler, and controller manager failures. Check component status, review control plane logs, verify cluster networking, and assess resource constraints. The skill guides you through identifying whether the issue affects the entire cluster or specific workloads.

How to respond to k8s incidents quickly?

k8s-incident accelerates incident response with structured runbooks and diagnostic workflows. Prioritize triage by assessing cluster health, service impact, and failure type. Execute emergency actions like pod deletion or rollback when needed, and collect comprehensive diagnostics for post-incident analysis and documentation.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

Kubernetes Incident Response

Runbooks and diagnostic workflows for common Kubernetes incidents.

When to Apply

Use this skill when: - User mentions: "incident", "outage", "emergency", "down", "not working" - Operations: emergency response, production issues, service degradation - Keywords: "urgent", "broken", "fix", "restore", "recover"

Priority Rules

Priority Rule Impact Tools
1 Check control plane first CRITICAL get_pods(namespace="kube-system")
2 Assess node health CRITICAL get_nodes
3 Gather events before changes HIGH get_events
4 Document timeline HIGH Manual notes
5 Rollback if safe MEDIUM rollback_deployment

Quick Reference

Incident First Tool Next Steps
Pod failure get_pod_logs(previous=True) describe_pod,

(truncated - see the full file via the links below)

File tree — 2 files
kubernetes-skills/claude/k8s-incident/SKILL.md
kubernetes-skills/claude/k8s-incident/scripts/collect-diagnostics.py

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Respond to active Kubernetes incidents with structured runbooks and diagnostics”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

k8s-deploy
by rohitg00 · rohitg00/kubectl-mcp-server

k8s-deploy provides deployment workflows for Kubernetes using kubectl-mcp-server tools, supporting standard deployments via manifests or Helm alongside progressive delivery through Argo Rollouts and Flagger. It covers canary promotions, blue-green strategies, rolling updates, scaling, and rollback operations across single and multi-cluster environments.

MITfor claude-codeupdated Apr 2026
★ 934repo stars
k8s-autoscaling
by rohitg00 · rohitg00/kubectl-mcp-server

Set up automatic scaling for Kubernetes workloads using HPA for CPU/memory-based scaling, VPA for resource optimization, and KEDA for event-driven scenarios like queue processing and scheduled scaling. Includes tools for detecting installations, managing scaled objects, and troubleshooting common scaling issues.

MITfor claude-codeupdated Apr 2026
★ 934repo stars
k8s-storage
by rohitg00 · rohitg00/kubectl-mcp-server

k8s-storage handles Kubernetes storage provisioning and management through a set of kubectl-backed tools. Create and configure persistent volume claims, storage classes, and persistent volumes; inspect their status and access modes; and troubleshoot storage issues directly from conversation.

MITfor claude-codeupdated Apr 2026
★ 934repo stars
k8s-vind
by rohitg00 · rohitg00/kubectl-mcp-server

k8s-vind lets you provision and operate virtual Kubernetes clusters as isolated workloads inside a host cluster. Use it to spin up ephemeral dev environments, establish tenant isolation, or test configurations without overhead. The skill exposes 14 tools for cluster creation, lifecycle management, connection handling, and resource optimization.

MITfor claude-codeupdated Apr 2026
★ 934repo stars
Systematic Debugging
by mrgoonie · mrgoonie/claudekit-skills

Systematic Debugging enforces a disciplined four-phase process: investigate root cause, analyze patterns, form and test hypotheses, then implement fixes. It stops you from guessing or patching symptoms, ensuring you understand the actual problem before making changes.

no license declared → metadata onlyfor claude-codeupdated Apr 2026
★ 2,186repo stars
issue-manage
by catlog22 · catlog22/Claude-Code-Workflow

Issue Management provides a menu-driven interface for tracking and organizing issues through the `ccw issue` CLI. Create, view, edit, delete, and bulk-update issues while filtering by status, priority, or other criteria—with automatic archiving of completed items.

MITfor claude-codeupdated Jun 2026
★ 2,142repo stars
Tags
incident-responseemergency-proceduresrunbook-automationproduction-outagescluster-diagnosticsfailure-recoverytriage-workflowsoperational-resilience