← Back to the wire22 Aug 2026
The AI WireDispatch No. 055
azureproduct· filed 22 Aug 2026 · 2 min read

Microsoft Foundry Adds DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning Models

Microsoft Foundry is adding DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning to its catalog, both optimized for agentic workloads. The models are available through multiple deployment paths, including Azure, Fireworks, and Hugging Face, with varying formats and pricing. Microsoft reports significant benchmark improvements for DeepSeek, though these figures are not independently verified.

Machine-drafted illustration · reviewed by a humanFIG. 01

Listen to this dispatch

Narrated by an AI-generated voice.

Microsoft Foundry is adding two open models to its catalog: DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning. Both target agentic workloads—coding agents, tool use, and workflow automation—and both arrive through multiple deployment paths, reflecting an approach that lets teams choose how they consume models rather than being locked into a single provider.

DeepSeek-V4-Flash-0731 is available direct from Azure and through Fireworks on Foundry. Microsoft reports gains over the previous Flash release on Foundry, including a jump on the DeepSWE benchmark from 7.3 to 54.4—roughly a sevenfold improvement—and a 21-point increase on Terminal Bench, from 61.8 to 82.7. These are Microsoft's figures, not independently verified.

NVIDIA Nemotron 3.5 Lightning is positioned for agentic workloads with tool calling, long-context processing, multilingual support, structured outputs, and multi-step task execution. It's available through Fireworks on Foundry and through the Hugging Face collection, which ships two formats: BF16 and NVFP4. BF16 is a full-precision reference model suited for fine-tuning, distillation, and research, but it demands more accelerator memory. NVFP4 is NVIDIA's recommended format for production inference, offering lower memory usage and faster throughput at some cost to model fidelity. Teams planning heavy customization will likely use BF16; teams prioritizing latency and cost per token will lean toward NVFP4.

Pricing varies by model and deployment type. DeepSeek-V4-Flash-0731 runs $0.150 per million input tokens and $0.310 per million output tokens on US Data Zone Standard. Nemotron Lightning 3.5 is cheaper: $0.060 input, $0.220 output. Cached input tokens cost less on both. Provisioned throughput is billed per deployed unit per hour—$1.00 for Global, $1.10 for US Data Zone—rather than per token consumed.

Teams can start serverless while evaluating a model, then move to reserved capacity as traffic stabilizes. The Hugging Face collection route adds managed compute with dedicated GPU capacity, letting teams pick accelerator family and scaling behavior. Fireworks deployments run inside a Foundry project with Azure governance and access controls, and bring-your-own-weights is supported on the same path. Direct from Azure models are billed through the Azure subscription and covered by Azure SLAs.

More deployment paths mean more choices about format, pricing, and infrastructure. For teams already inside Azure, the appeal is consolidated tooling: discover, evaluate, deploy, and govern from one platform. For teams that simply want a model, the catalog's growing complexity is something to navigate rather than something that disappears.

Read the original at techcommunity.microsoft.com

End of dispatch
More on the wire
57developer.nvidia.comresearch · filed 22 Aug 2026AVO agent architecture achieves perfect ARC-AGI-3 score and produces optimized GPU kernels56github.blogproduct · filed 22 Aug 2026GitHub Copilot cloud agent available in Microsoft Teams public preview54azuretooling · filed 22 Aug 2026MCP Connectors canvas gives Copilot agents managed access to external tools without manual configuration53azureproduct · filed 22 Aug 2026Microsoft expands Foundry model router to additional regions and refreshes model pool