{"enrichment":{"faq":[{"a":"Watch Skill processes videos locally by extracting scene-aware frames, running OCR on text within them, and transcribing audio via offline Whisper or captions. It builds a searchable index so you can ask follow-up questions about the video content without re-processing. No API keys are required for core features.","q":"How do I watch videos in Claude with the watch skill?"},{"a":"Yes. Watch Skill processes video content offline without uploading files to cloud services. It extracts frames locally, performs OCR on detected text, and transcribes audio using offline Whisper or existing captions\u2014all on your machine with no external API calls required for core functionality.","q":"Can I extract frames from video offline without uploading?"},{"a":"Watch Skill extracts structured data including scene-aware frames, text via OCR, full audio transcription with timestamps, and caption data. It indexes all extracted content so you can search across multiple videos, ask follow-up questions about specific scenes or text, and retrieve results without reprocessing.","q":"What does watch skill video processing extract from files?"},{"a":"Watch Skill builds a persistent indexed database of all extracted content\u2014frames, OCR text, transcripts, and captions\u2014from every video you process. You can then search this index and ask follow-up questions about any video's content, with results returned instantly from the database rather than requiring re-analysis.","q":"How can I search text and captions across multiple videos?"},{"a":"Watch Skill's core features\u2014frame extraction, indexing, and search\u2014require no API keys. Vision and speech-to-text are optional and configurable; you can use offline Whisper for transcription or rely on existing captions, keeping your video processing completely local and private.","q":"Is video OCR and transcription available without an API key?"}],"shadow_tags":["video-processing","offline-first","frame-extraction","speech-to-text","optical-character-recognition","content-indexing","persistent-cache","scene-detection","local-inference","multi-modal-analysis"],"summary_rewrite":"Watch Skill processes videos locally by extracting scene-aware frames, running OCR on text within them, and transcribing audio via offline Whisper or captions. It builds a searchable index so follow-up questions are answered without re-processing the video. No API keys required for core features; vision and STT are optional and configurable."},"files":[{"bytes":5363,"path":"adapters/claude-skill/skills/watch/SKILL.md","sha256":"17b1a4d3caea734de6ca0e82a65881f9766ca8a040fd4b58390da20d9e898167","url":"https://skillfed.io/files/oxbshw/watch-skill/watch/217c7d6d/SKILL.md"}],"id":"oxbshw/watch-skill/watch","links":{"html":"https://skillfed.io/oxbshw/watch-skill/watch","md":"https://skillfed.io/oxbshw/watch-skill/watch.md","repo":"https://github.com/oxbshw/watch-skill"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":36,"language":"Python","last_updated":"2026-07-12","license":"MIT","name":"watch","publisher":"oxbshw","stars":236},"relations":{"categories":["video-analysis","content-extraction","offline-processing"],"similar":[{"id":"oxbshw/watch-skill/watching-videos"},{"id":"bradautomates/claude-video/watch"},{"id":"oxbshw/watch-skill/asking-with-evidence"},{"id":"HUANGCHIHHUNGLeo/claude-real-video/claude-real-video-for-agents"},{"id":"oxbshw/watch-skill/video-memory"},{"id":"cdeistopened/skill-stack/youtube-clip-extractor"},{"id":"intellectronica/agent-skills/youtube-transcript"},{"id":"oxbshw/watch-skill/sharing-results"},{"id":"oxbshw/watch-skill/learning-from-mistakes"},{"id":"HUANGCHIHHUNGLeo/claude-real-video/claude-real-video"}]},"slug":{"owner":"oxbshw","repo":"watch-skill","skill":"watch"},"version":"217c7d6d"}
