NewSnippets: tell the AI about you and your company once.See how New modelClaude Sonnet 5.5 is now available.Read more

DeepSeek R1-0528: China's AI Breakthrough Challenges Silicon Valley's Dominance

DeepSeek R1-0528 upgrade delivers 87.5% AIME accuracy, distilled 8B models, and enhanced reasoning.

DeepSeek quietly released DeepSeek R1-0528, an upgraded version of its reasoning model that the Chinese startup describes as a “minor trial upgrade”. Yet the improvements are anything but minor, delivering performance gains that position this open-source model dangerously close to proprietary alternatives from OpenAI and Google.

For enterprises, it is further evidence that open-source models are closing in on, and in places matching, the best closed models.

Major Performance Leaps in Mathematical Reasoning

In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training. The results speak volumes about the model’s enhanced capabilities.

For instance, in the AIME 2025 test, the model’s accuracy has increased from 70% in the previous version to 87.5% in the current version. This 17.5-point improvement in mathematical reasoning represents one of the most dramatic performance gains seen in recent AI model updates.

The model has demonstrated outstanding performance across various benchmark evaluations, including mathematics, programming, and general logic. Its overall performance is now approaching that of leading models, such as O3 and Gemini 2.5 Pro.

The Power of Deeper Reasoning

What’s driving these improvements? The model demonstrates deeper chain-of-thought reasoning, using nearly double the tokens per query on challenging problems (averaging 23K tokens of “thinking” vs 12K before). This intensive computational approach allows the model to work through complex problems more thoroughly, resulting in significantly higher accuracy rates.

“DeepSeek’s latest upgrade is sharper on reasoning, stronger on math and code, and closing in on top-tier models like Gemini and O3,” according to Adina Yakefu, AI researcher at Hugging Face. The upgraded model has “major improvements in inference and hallucination reduction”.

Game-Changing Distilled Models

Perhaps even more significant for businesses is DeepSeek’s release of distilled models. DeepSeek also released a smaller, “distilled” version of its new R1, DeepSeek-R1-0528-Qwen3-8B, that DeepSeek claims beats comparably sized models on certain benchmarks. The smaller updated R1, which was built using the Qwen3-8B model Alibaba launched in May as a foundation, performs better than Google’s Gemini 2.5 Flash on AIME 2025.

According to the cloud platform NodeShift, Qwen3-8B requires a GPU with 40GB-80GB of RAM to run (e.g., an Nvidia H100). DeepSeek’s distilled new R1 AI model can run on a single GPU, putting it within reach of hobbyists.

This accessibility democratizes advanced AI reasoning capabilities, allowing smaller organizations to deploy sophisticated models without massive infrastructure investments.

Competitive Positioning Against Industry Leaders

How does DeepSeek R1-0528 stack up against the competition? The upgraded DeepSeek R1 model is just behind OpenAI’s o4-mini and o3 reasoning models on LiveCodeBench, a site that benchmarks models against different metrics.

Notably, o3’s superior scores sometimes rely on running in a costly “high effort” mode (using extended thinking or tool use), which is computationally expensive. R1-0528 nearly matches o3’s performance without such extreme settings.

For coding tasks, the o4-mini-high mode achieved ~69% on the Aider-Polyglot test, essentially on par with DeepSeek R1-0528’s 71-72% on that test.

Open Source Advantages and Commercial Licensing

DeepSeek-R1-0528-Qwen3-8B is available under a permissive MIT license, meaning it can be used commercially without restriction. This licensing model removes barriers for businesses looking to integrate advanced AI reasoning capabilities into their products and services.

The use of DeepSeek-R1 models is also subject to MIT License. DeepSeek-R1 series (including Base and Chat) supports commercial use and distillation, providing unprecedented flexibility for enterprise deployment.

Enhanced Capabilities and Reduced Hallucinations

In May 2025, they released DeepSeek-R1-0528, an upgraded version with better benchmark performance, fewer hallucinations, and new capabilities like function calling and JSON output support. This update introduces several key improvements: improved benchmark performance across both reasoning and factual tasks, enhanced front-end capabilities for smoother interaction in chat platforms, reduced hallucinations, increasing factual reliability, and support for JSON output and function calling.

These improvements make the model more suitable for production environments where reliability and structured output formats are critical business requirements.

Strategic Implications

The economics shift noticeably. Organizations using multi-model AI platforms can now access reasoning capabilities that rival proprietary models at dramatically reduced costs. The availability of multiple model sizes - from 8 billion to 671 billion parameters - allows businesses to optimize their AI infrastructure based on specific use case requirements.

Where OpenAI o1 costs $15 per million input tokens and $60 per million output tokens, DeepSeek Reasoner, which is based on the R1 model, costs $0.55 per million input and $2.19 per million output tokens - representing cost savings of over 90%.

The Future of Open Source AI Leadership

This version shows DeepSeek is not just catching up, it’s competing. The rapid advancement of open-source models like DeepSeek R1-0528 suggests that the future of AI may be more distributed and accessible than many anticipated.

For enterprises, the takeaway is less about one release than about pace. When an open model closes the gap in a single update, the right model for a given task can change from one quarter to the next. Being able to test a new model against your own work, quickly, matters more than any single bet.

StickyPrompts gives your team DeepSeek alongside models from OpenAI, Anthropic and Google in one governed workspace. Run the same prompt on several of them side by side and keep the one that does the job best. Start a free workspace.

All posts
Trial
A $5 balance to start. No card needed.
Every model, one workspace

Put the best model for the job in front of your team.

StickyPrompts gives your team the leading frontier and open-weight models in one governed workspace, with shared prompts, team chats and admin controls. Start free with a $5 trial balance.