Hangzhou, China · 20 January 2025 — DeepSeek, a Chinese AI research company, released DeepSeek-R1, an open-weights large language model engineered for reasoning tasks that achieved performance competitive with OpenAI's o1 across mathematics, coding, and logical inference benchmarks. The release, made freely available under the MIT License, sent shockwaves through the AI industry and triggered significant volatility in technology stocks as investors reassessed the cost trajectory of frontier AI development.
The model's release marked a watershed moment in the ongoing debate between open and proprietary AI development, demonstrating that high-performance reasoning capabilities could be achieved at substantially lower training costs than previously assumed by the industry.
Architecture and Training
DeepSeek-R1 is built upon the DeepSeek-V3-Base model and employs a Mixture-of-Experts (MoE) architecture with 671 billion total parameters, of which approximately 37 billion are activated per forward pass. This selective activation design allows the model to maintain high performance while keeping inference costs manageable — a key factor in its disruptive market impact.
The training methodology diverged significantly from conventional approaches. Rather than relying primarily on supervised fine-tuning (SFT) with large datasets of human-labeled examples, DeepSeek employed large-scale reinforcement learning (RL) as the core training mechanism.
However, R1-Zero exhibited issues including endless repetition, poor readability, and language mixing. To address these problems, the final DeepSeek-R1 model incorporated a multi-stage training pipeline:
- Cold-start data: Minimal human-labeled examples were used to seed the model's reasoning capabilities
- Reasoning-oriented RL: Extensive reinforcement learning to develop chain-of-thought reasoning
- Rejection sampling and SFT: Generating training data from the RL-enhanced model to improve non-reasoning tasks
- Alignment RL: A second RL stage to align with human preferences
Benchmark Performance
DeepSeek-R1 demonstrated competitive or superior performance against leading proprietary models across multiple benchmarks:
| Benchmark | DeepSeek R1 | OpenAI o1-1217 | Claude-3.5-Sonnet | GPT-4o |
|---|---|---|---|---|
| AIME 2024 (Pass@1) | 79.8 | 79.2 | 16.0 | 9.3 |
| MATH-500 (Pass@1) | 97.3 | 96.4 | 78.3 | 74.6 |
| MMLU (Pass@1) | 90.8 | 91.8 | 88.3 | 87.2 |
| LiveCodeBench (Pass@1) | 65.9 | 63.4 | 33.8 | 34.2 |
| Codeforces (Rating) | 2029 | 2061 | 717 | 759 |
The model achieved 79.8% on AIME 2024 (the American Invitational Mathematics Examination) and 97.3% on MATH-500, slightly surpassing OpenAI's o1 on both benchmarks. On coding tasks, it scored 65.9% on LiveCodeBench, again edging out o1's 63.4%.
Distilled Models and Open Source Release
In addition to the full 671B model, DeepSeek released six smaller "distilled" models ranging from 1.5 billion to 70 billion parameters. These distilled models were created by fine-tuning existing open-source architectures (Qwen2.5 and Llama 3 series) using reasoning data generated by DeepSeek-R1.
All models were released under the MIT License, permitting free academic and commercial use, including modification and derivative works such as distillation for training other models. The full model weights were published on HuggingFace, and the accompanying technical paper was posted to arXiv.
Market Impact and Industry Response
The release of DeepSeek-R1 had immediate and dramatic effects on financial markets. On January 27, 2025, NVIDIA's stock price fell by approximately 17%, wiping out nearly $600 billion in market capitalization — the largest single-day loss in U.S. stock market history. The broader technology sector experienced significant sell-offs as investors reconsidered the capital expenditure assumptions underpinning the AI investment cycle.
The market reaction reflected concerns that DeepSeek's approach — achieving frontier-level reasoning performance at a fraction of the training cost reported by major U.S. AI labs — could undermine the competitive moats of companies investing billions in AI infrastructure. DeepSeek reported training costs of approximately $5.6 million for the underlying DeepSeek-V3 model, though this figure covered only the final training run and excluded prior research, infrastructure, and personnel costs.
Implications for AI Development
DeepSeek-R1's release raised several important questions for the AI research community:
- Open vs. proprietary: The model's competitive performance under an open license challenged the assumption that frontier AI capabilities require closed, capital-intensive development.
- RL as a reasoning catalyst: The success of the RL-first training approach suggested that supervised fine-tuning on human-labeled data may be less essential for reasoning capabilities than previously believed.
- Distillation as a scaling strategy: The strong performance of distilled models indicated that reasoning capabilities could be efficiently transferred to smaller, more practical models.
- Geopolitical dimensions: The release by a Chinese company intensified discussions about AI competition, export controls on advanced chips, and the global distribution of AI capabilities.
The model's chain-of-thought reasoning approach — where it explicitly works through problems step by step before producing an answer — also made its internal reasoning process visible to users, a transparency feature that proprietary models like OpenAI's o1 did not fully provide.
Sources
- DeepSeek-AI, "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning," arXiv:2501.12948, January 2025
- DeepSeek-R1 GitHub repository:
- DeepSeek-R1 model weights on HuggingFace:
- International Institute for Strategic Studies (IISS), "DeepSeek's release of an open-weight frontier AI model," Strategic Comments, April 2025
- DeepSeek API documentation:
FIRAT Editorial Board
Institutional Research Desk · Foresight Institute of Research and Translation
The collective editorial and research translation board of FIRAT, synthesising peer-reviewed evidence, policy briefs, and division milestones across our seven foundational research pillars.



