{"categories":[{"label":"Libraries","url":"https://skillfed.io/packages/category/software-development-libraries/4"},{"label":"Scientific/Engineering","url":"https://skillfed.io/packages/category/scientific-engineering/2"},{"label":"Python Modules","url":"https://skillfed.io/packages/category/software-development-libraries-python-modules/7"},{"label":"Utilities","url":"https://skillfed.io/packages/category/utilities/3"},{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/3"},{"label":"Mathematics","url":"https://skillfed.io/packages/category/scientific-engineering-mathematics"},{"label":"Image Recognition","url":"https://skillfed.io/packages/category/scientific-engineering-image-recognition"}],"enrichment":{"capability":"nemo-toolkit provides a PyTorch framework for building, training, and deploying speech AI models including automatic speech recognition (ASR), text-to-speech (TTS), and speech-based language models.","skillfed_tags":["speech-recognition","text-to-speech","gpu-accelerated"],"use_cases":["Build automatic speech recognition systems for English or 25+ European languages using pre-trained Parakeet or Canary checkpoints","Deploy streaming ASR with configurable latency (80ms\u20131s) using Nemotron-3.5-ASR-Streaming for real-time transcription","Generate multilingual speech synthesis using MagpieTTS with support for 9 languages including English, Spanish, French, and Mandarin","Fine-tune or customize speech models on your own audio data using the modular PyTorch-based training pipeline","Integrate speech recognition and translation into conversational AI applications using Nemotron VoiceChat or speech LLM components"],"what_it_does":"nemo-toolkit is NVIDIA's framework for speech AI, built on PyTorch to help researchers and developers train and deploy models for automatic speech recognition, text-to-speech synthesis, and speech-based language models. It ships with pre-trained checkpoints (Parakeet, Canary, MagpieTTS, Nemotron-Speech) covering multiple languages and use cases, and supports both offline and streaming inference with configurable latency-accuracy tradeoffs.\n\nThe package installs as a pure-Python wheel over your existing PyTorch and CUDA setup without replacing them. It depends on 15 runtime packages including torch, numba, scikit-learn, huggingface_hub, and tensorboard. Training requires an NVIDIA GPU; inference can run on CPU but GPU is recommended. The repository is actively maintained (last commit 2026-08-14) and carries no known vulnerabilities.","worth_installing":"Yes. nemo-toolkit is actively maintained, carries no known vulnerabilities, and offers low install friction. It is well-suited for researchers and developers building speech AI systems. Requires PyTorch 2.7+, Python 3.10+, and ideally an NVIDIA GPU; if your environment already meets these, installation is straightforward. The Apache 2.0 license permits commercial use. Install if you need to work with ASR, TTS, or speech-based language models; skip if you need only inference on pre-trained models without customization or if you lack GPU access."},"id":"nemo-toolkit","links":{"html":"https://skillfed.io/packages/nemo-toolkit","md":"https://skillfed.io/packages/nemo-toolkit.md","pypi":"https://pypi.org/project/nemo-toolkit/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-07","license_spdx":null,"license_treatment":"permissive","name":"nemo-toolkit","python_support":"supports_current","summary":"NeMo - a toolkit for Conversational AI"},"popularity":{"monthly_downloads":1538107,"position":3792,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"3.0.0"}
