← Back to the wire6 Oct 2026
The AI WireDispatch No. 055
azure ▸ product· filed 22 Aug 2026 · 2 min read

Microsoft Foundry Adds DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning Models

Microsoft Foundry is adding DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning to its catalog, both optimized for agentic workloads. The models are available through multiple deployment paths, including Azure, Fireworks, and Hugging Face, with varying formats and pricing. Microsoft reports significant benchmark improvements for DeepSeek, though these figures are not independently verified.

Machine-drafted illustration · reviewed by a humanFIG. 01

Listen to this dispatch

Narrated by an AI-generated voice.

Microsoft Foundry is adding two open models to its catalog: DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning. Both target agentic workloads—coding agents, tool use, and workflow automation—and both arrive through multiple deployment paths, reflecting an approach that lets teams choose how they consume models rather than being locked into a single provider.

DeepSeek-V4-Flash-0731 is available direct from Azure and through Fireworks on Foundry. Microsoft reports gains over the previous Flash release on Foundry, including a jump on the DeepSWE benchmark from 7.3 to 54.4—roughly a sevenfold improvement—and a 21-point increase on Terminal Bench, from 61.8 to 82.7. These are Microsoft's figures, not independently verified.

NVIDIA Nemotron 3.5 Lightning is positioned for agentic workloads with tool calling, long-context processing, multilingual support, structured outputs, and multi-step task execution. It's available through Fireworks on Foundry and through the Hugging Face collection, which ships two formats: BF16 and NVFP4. BF16 is a full-precision reference model suited for fine-tuning, distillation, and research, but it demands more accelerator memory. NVFP4 is NVIDIA's recommended format for production inference, offering lower memory usage and faster throughput at some cost to model fidelity. Teams planning heavy customization will likely use BF16; teams prioritizing latency and cost per token will lean toward NVFP4.

Pricing varies by model and deployment type. DeepSeek-V4-Flash-0731 runs $0.150 per million input tokens and $0.310 per million output tokens on US Data Zone Standard. Nemotron Lightning 3.5 is cheaper: $0.060 input, $0.220 output. Cached input tokens cost less on both. Provisioned throughput is billed per deployed unit per hour—$1.00 for Global, $1.10 for US Data Zone—rather than per token consumed.

Teams can start serverless while evaluating a model, then move to reserved capacity as traffic stabilizes. The Hugging Face collection route adds managed compute with dedicated GPU capacity, letting teams pick accelerator family and scaling behavior. Fireworks deployments run inside a Foundry project with Azure governance and access controls, and bring-your-own-weights is supported on the same path. Direct from Azure models are billed through the Azure subscription and covered by Azure SLAs.

More deployment paths mean more choices about format, pricing, and infrastructure. For teams already inside Azure, the appeal is consolidated tooling: discover, evaluate, deploy, and govern from one platform. For teams that simply want a model, the catalog's growing complexity is something to navigate rather than something that disappears.

Read the original at techcommunity.microsoft.com →

End of dispatch
More on the wire
87techcommunity.microsoft.com ▸ azure · filed 2 Oct 2026Federated Agent Factory: Microsoft's Reference Architecture for AI Agents Across Entra Tenants86techcommunity.microsoft.com ▸ models · filed 2 Oct 2026Microsoft adds streaming transcription and two text-to-speech models to Azure AI Foundry85techcommunity.microsoft.com ▸ tooling · filed 2 Oct 2026Open-source demo separates AI model proposals from sandboxed tool execution84techcommunity.microsoft.com ▸ product · filed 27 Sept 2026Azure Container Apps Sandboxes: per-agent microVMs with egress controls