The Nemotron 3 Embed series received broad ecosystem support on its release day: model weights are available on Hugging Face, deployable via NVIDIA NIM microservices, and compatible with the vLLM inference framework. The 32K context window supports long documents, codebases, and multi-turn agent history retrieval, reducing truncation loss.
Multilingual and code retrieval capabilities cover global enterprise data and technical documentation. NVIDIA emphasizes that open weights and recipes give teams full control over inspection, tuning, and deployment of retrieval models, suitable for production-grade RAG, agent retrieval, code retrieval, and agent memory scenarios.