NVIDIA AI Releases Star Elastic: One Checkpoint that Contains 30B, 23B, and 12B Reasoning Models with Zero-Shot Slicing
mai 10, 2026admin NUAI,Committee,Actualités,Uncategorized(0)Training a family of large language models (LLMs) has always come with a painful multiplier:...

NVIDIA AI Releases Orchestrator-8B: A Reinforcement Learning Trained Controller for Efficient Tool and Model Selection
novembre 29, 2025admin NUAI,Committee,Actualités,Uncategorized(0)How can an AI system learn to pick the right model or tool for each...

NVIDIA AI Releases OpenReasoning-Nemotron: A Suite of Reasoning-Enhanced LLMs Distilled from DeepSeek R1 0528
juillet 20, 2025admin NUAI,Committee,Actualités,Uncategorized(0)NVIDIA AI has introduced OpenReasoning-Nemotron, a family of large language models (LLMs) designed to excel...

NVIDIA AI Releases Nemotron-Labs-Diffusion: A Tri-Mode Language Model with 6× Tokens Per Forward Over Qwen3-8B
mai 20, 2026admin NUAI,Committee,Actualités,Uncategorized(0)NVIDIA researchers have released Nemotron-Labs-Diffusion, a language model family that unifies three decoding modes in...

NVIDIA AI Releases Nemotron-Elastic-12B: A Single AI Model that Gives You 6B/9B/12B Variants without Extra Training Cost
novembre 24, 2025admin NUAI,Committee,Actualités,Uncategorized(0)Why are AI dev teams still training and storing multiple large language models for different...

NVIDIA AI Releases Nemotron 3: A Hybrid Mamba Transformer MoE Stack for Long Context Agentic AI
décembre 21, 2025admin NUAI,Committee,Actualités,Uncategorized(0)NVIDIA has released the Nemotron 3 family of open models as part of a full...

NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB
juillet 17, 2026admin NUAI,Committee,Actualités,Uncategorized(0)Embedding models decide which passages an agent ever sees. NVIDIA released Nemotron 3 Embed model...
NVIDIA AI Releases Gated DeltaNet-2: A Linear Attention Layer That Decouples Erase and Write in the Delta Rule
mai 24, 2026admin NUAI,Committee,Actualités,Uncategorized(0)Linear attention replaces the unbounded KV cache of softmax attention with a fixed-size recurrent state...

NVIDIA AI Releases Dynamo Snapshot: A CRIU-Based Fast Startup System for AI Inference on Kubernetes
juin 5, 2026admin NUAI,Committee,Actualités,Uncategorized(0)In production inference deployments, demand fluctuates over time, requiring inference replicas to scale elastically. Cold-starting...
