Microsoft Foundry Adds DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning Models
Microsoft Foundry is adding DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning to its catalog, both optimized for agentic workloads. The models are available through multiple deployment paths, including Azure, Fireworks, and Hugging Face, with varying formats and pricing. Microsoft reports significant benchmark improvements for DeepSeek, though these figures are not independently verified.

Listen to this dispatch
Narrated by an AI-generated voice.
Microsoft Foundry is adding two open models to its catalog: DeepSeek-V4-Flash-0731 and NVIDIA Nemotron 3.5 Lightning. Both target agentic workloads—coding agents, tool use, and workflow automation—and both arrive through multiple deployment paths, reflecting an approach that lets teams choose how they consume models rather than being locked into a single provider.
DeepSeek-V4-Flash-0731 is available direct from Azure and through Fireworks on Foundry. Microsoft reports gains over the previous Flash release on Foundry, including a jump on the DeepSWE benchmark from 7.3 to 54.4—roughly a sevenfold improvement—and a 21-point increase on Terminal Bench, from 61.8 to 82.7. These are Microsoft's figures, not independently verified.
NVIDIA Nemotron 3.5 Lightning is positioned for agentic workloads with tool calling, long-context processing, multilingual support, structured outputs, and multi-step task execution. It's available through Fireworks on Foundry and through the Hugging Face collection, which ships two formats: BF16 and NVFP4. BF16 is a full-precision reference model suited for fine-tuning, distillation, and research, but it demands more accelerator memory. NVFP4 is NVIDIA's recommended format for production inference, offering lower memory usage and faster throughput at some cost to model fidelity. Teams planning heavy customization will likely use BF16; teams prioritizing latency and cost per token will lean toward NVFP4.
Pricing varies by model and deployment type. DeepSeek-V4-Flash-0731 runs $0.150 per million input tokens and $0.310 per million output tokens on US Data Zone Standard. Nemotron Lightning 3.5 is cheaper: $0.060 input, $0.220 output. Cached input tokens cost less on both. Provisioned throughput is billed per deployed unit per hour—$1.00 for Global, $1.10 for US Data Zone—rather than per token consumed.
Teams can start serverless while evaluating a model, then move to reserved capacity as traffic stabilizes. The Hugging Face collection route adds managed compute with dedicated GPU capacity, letting teams pick accelerator family and scaling behavior. Fireworks deployments run inside a Foundry project with Azure governance and access controls, and bring-your-own-weights is supported on the same path. Direct from Azure models are billed through the Azure subscription and covered by Azure SLAs.
More deployment paths mean more choices about format, pricing, and infrastructure. For teams already inside Azure, the appeal is consolidated tooling: discover, evaluate, deploy, and govern from one platform. For teams that simply want a model, the catalog's growing complexity is something to navigate rather than something that disappears.
Read the original at techcommunity.microsoft.com →