Changelog
Follow up on the latest improvements and updates.
RSS
DeepSeek-V4-Flash-0731, DeepSeek's latest release optimized for agentic coding and autonomous workflows, is now available through DigitalOcean Inference Engine. A 284B-parameter mixture-of-experts model with only 13B parameters active per request, this model delivers fast inference at a fraction of the compute cost of a dense model of comparable size. A full re-post-training pass focused on autonomous task execution pushes its agentic performance ahead of DeepSeek's own V4-Pro, despite the smaller active parameter count. It excels at repo-scale coding agent workflows, multi-step tool use, and production-grade agentic loops.
Kansas City (MKC1) is now available as our newest DigitalOcean data center region. MKC1 expands DigitalOcean's footprint to 20 data centers across 11 global regions. MKC1 is a fully liquid-cooled data center and runs NVIDIA B300 GPUs.
Available at launch:
- Droplets, custom images, and Droplet Autoscaler
- Kubernetes (DOKS), App Platform, and Functions
- Managed Databases
- Spaces Object Storage (standard and cold tiers), Volumes, NFS, backups and snapshots
- Networking: Load Balancers, VPC, NAT Gateway, Cloud Firewalls, reserved IPs, DNS, and Partner Network Connect
- Cloud Security Posture Management (CSPM), Marketplace, monitoring and alerts, and IAM
MKC1 has achieved an ISO 27001:2022 certification alongside SOC 2 Type II, which is reflected in a SOC 3 Type II report. These certifications can be accessed in our Security Reports and Certifications Center. As with every DigitalOcean region, you get the full platform in one place and on one bill, so your inference, compute, storage, databases, and apps sit together instead of across vendor boundaries.
Kimi K3 is the first open 3T-class model, with a 1M-token context window that ingests entire codebases and long documents without chunking. Its sparse expert architecture (2.8T parameters, 16 of 896 active per token) delivers frontier-scale capability with efficient inference, and native multimodal support lets it reason over text, images, and video together. Tuned for max thinking effort by default, it sustains long-horizon autonomous work like coding, research automation, and multi-hour agentic sessions.
Claude Opus 5, Anthropic's latest flagship model for coding, reasoning, and knowledge work, is now available through DigitalOcean Inference Engine. Designed for demanding agentic workflows, Opus 5 delivers frontier-level performance with improved cost efficiency, offering intelligence close to Claude Fable 5 at half the price. It excels at complex software engineering, scientific research, and business automation tasks, while supporting long-context reasoning, tool use, and structured outputs.
Model synthesis is a new opt-in server-side tool in DigitalOcean Inference Engine. Run multiple models on the same inference request and get a single synthesized response, no custom orchestration required. Built for complex, high-reasoning tasks, it works with the existing Chat Completions, Responses, and Messages APIs. Choose an optimized preset or configure the panel and synthesizer directly.
OpenAI's GPT-5.6 family, Sol (flagship), Terra (balanced), and Luna (fast, cost-efficient), is now available through DigitalOcean Serverless Inference. Sol delivers state-of-the-art performance across coding, knowledge work, cybersecurity, and scientific reasoning, while Terra provides a balanced option for everyday production workloads and Luna offers their fastest, most affordable model in the family. New
max
reasoning and ultra
mode help tackle complex, multi-step tasks.With Prompt Caching, DigitalOcean automatically recognizes repeated token prefixes and currently offers an 80% discount on cached input tokens compared to standard input pricing across supported models. See the pricing page for current rates, applicable terms, and any pricing updates. The result is lower costs and faster time-to-first-token, with no changes to your application code. Caching applies automatically, and every API response includes a
cached_tokens
count so you can verify your savings in real time. It is available today across a broad set of leading open models, with more models coming soon. Teams can now validate any model or inference router configuration on their own data before production. Run structured LLM-as-a-Judge evaluations across catalog models, fine-tuned models, BYOM imports, and router setups without stitching together a separate evaluation stack.
Claude Sonnet 5, Anthropic’s latest model for coding, autonomous agents, and professional work at scale, is now available through DigitalOcean Serverless Inference. Designed for production AI applications, Sonnet 5 can plan, use tools, and complete complex tasks while delivering near-Opus 4.8 performance on leading agentic benchmarks, including SWE-bench, BrowseComp, and OSWorld-Verified, at Sonnet pricing.
DOKS clusters can now authenticate users through an external OpenID Connect provider. Each cluster has its own independent configuration managed via doctl, so dev, staging, and production environments can each enforce distinct access policies. Token issuance and revocation are handled directly from the IdP, so deactivating a user there removes their cluster access without manual credential rotation.
Load More
→