DeepSeek开源模型排行榜
本榜收录 DeepSeek 在本站的 71 个开源模型,按社区累计下载量排序。该系列以推理能力与宽松的开源协议受到关注,覆盖通用对话与推理专用方向,并提供多种参数档位与蒸馏版本,适合需要在本地构建推理能力、或对推理过程可解释性有要求的场景。
数据更新时间:2026-09-22
本榜收录 71 个模型
模型库共 1,408 个
排序依据:累计下载量 ↓
模型总榜
国产大模型榜
海外开源模型榜
通义千问(阿里)
AI2(Allen)
Microsoft
Google
面壁智能
上海AI实验室
智谱AI
DeepSeek
智源研究院
魔搭社区
腾讯混元
Mistral AI
下载量排名 TOP50
魔搭社区累计下载| 排名 | 模型 | 厂商 / 协议 | 参数量 | 下载量 | 点赞 |
|---|---|---|---|---|---|
| 1 | DeepSeek-R1-Distill-Qwen-32B | DeepSeek mit | 32.8B | 274.7万 | 273 |
| 2 | DeepSeek-V3.1-Terminus | DeepSeek mit | 684.5B | 260.0万 | 70 |
| 3 | DeepSeek-R1-Distill-Qwen-1.5B | DeepSeek mit | 1.8B | 222.1万 | 343 |
| 4 | DeepSeek-R1-Distill-Llama-70B | DeepSeek mit | 70.6B | 212.4万 | 141 |
| 5 | DeepSeek-R1-Distill-Qwen-7B | DeepSeek mit | 7.6B | 91.7万 | 447 |
| 6 | DeepSeek-V2-Lite-Chat | DeepSeek other | 15.7B | 79.1万 | 18 |
| 7 | DeepSeek-Coder-V2-Lite-Instruct | DeepSeek other | 15.7B | 66.9万 | 34 |
| 8 | DeepSeek-V2-Chat | DeepSeek other | 235.7B | 62.3万 | 42 |
| 9 | DeepSeek-R1-0528 | DeepSeek mit | 684.5B | 61.1万 | 330 |
| 10 | DeepSeek-V4-Flash | DeepSeek mit | 290.9B | 59.0万 | 400 |
| 11 | DeepSeek-R1-Distill-Qwen-14B | DeepSeek mit | 14.8B | 47.0万 | 180 |
| 12 | DeepSeek-V3.2 | DeepSeek | 685.4B | 43.9万 | 463 |
| 13 | DeepSeek-V3 | DeepSeek | 684.5B | 33.7万 | 266 |
| 14 | DeepSeek-R1-0528-Qwen3-8B | DeepSeek mit | 8.2B | 29.3万 | 130 |
| 15 | DeepSeek-Coder-V2-Instruct | DeepSeek other | 235.7B | 28.9万 | 7 |
| 16 | deepseek-coder-6.7b-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | 6.7B | 25.4万 | 32 |
| 17 | DeepSeek-V4-Pro | DeepSeek mit | 861.6B | 24.2万 | 552 |
| 18 | DeepSeek-R1-Distill-Llama-8B | DeepSeek mit | 8.0B | 15.6万 | 80 |
| 19 | DeepSeek-V2-Lite | DeepSeek other | 15.7B | 13.6万 | 7 |
| 20 | deepseek-moe-16b-chat | DeepSeek other | 16.4B | 12.9万 | 9 |
| 21 | DeepSeek-V4-Flash-0731 | DeepSeek mit | 304.2B | 11.6万 | 517 |
| 22 | deepseek-llm-67b-chat | DeepSeek other | — | 10.9万 | 6 |
| 23 | DeepSeek-V3.1 | DeepSeek mit | 684.5B | 7.0万 | 256 |
| 24 | deepseek-llm-7b-chat | DeepSeek other | — | 6.7万 | 28 |
| 25 | DeepSeek-V4-Flash-DSpark | DeepSeek mit | 165.3B | 6.0万 | 35 |
| 26 | DeepSeek-V3-Base | DeepSeek | 684.5B | 3.5万 | 44 |
| 27 | DeepSeek-V3-0324 | DeepSeek mit | 684.5B | 3.3万 | 297 |
| 28 | DeepSeek-V2 | DeepSeek other | 235.7B | 1.7万 | 7 |
| 29 | DeepSeek-V4-Flash-Base | DeepSeek MIT License | 292.0B | 1.6万 | 20 |
| 30 | DeepSeek-V3.2-Speciale | DeepSeek | 685.4B | 1.6万 | 103 |
| 31 | deepseek-llm-7b-base | DeepSeek other | — | 1.1万 | 8 |
| 32 | deepseek-coder-1.3b-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | 1.3B | 1.1万 | 9 |
| 33 | deepseek-coder-33b-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | 33.3B | 9927 | 12 |
| 34 | deepseek-moe-16b-base | DeepSeek other | 16.4B | 9747 | 2 |
| 35 | DeepSeek-V4-Pro-DSpark | DeepSeek mit | 889.5B | 9550 | 62 |
| 36 | deepseek-coder-5.7bmqa-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | — | 9297 | 0 |
| 37 | deepseek-coder-6.7b-base Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | 6.7B | 8008 | 9 |
| 38 | DeepSeek-V4-Pro-Base | DeepSeek MIT License | 1600.8B | 7924 | 23 |
| 39 | DeepSeek-V4-Pro-0813 | DeepSeek mit | 1650.5B | 7903 | 310 |
| 40 | deepseek-vl-7b-base | DeepSeek other | 7.3B | 6801 | 6 |
| 41 | deepseek-vl-1.3b-base | DeepSeek other | 2.0B | 6372 | 4 |
| 42 | deepseek-coder-33b-base Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | 33.3B | 6361 | 8 |
| 43 | deepseek-coder-1.3b-base Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. | DeepSeek other | — | 6310 | 7 |
| 44 | DeepSeek-Coder-V2-Lite-Base | DeepSeek other | 15.7B | 5273 | 1 |
| 45 | DeepSeek-V2.5-1210 | DeepSeek other | 235.7B | 4541 | 6 |
| 46 | DeepSeek-V2-Chat-0628 | DeepSeek | 235.7B | 4149 | 3 |
| 47 | DeepSeek-Math-V2 | DeepSeek apache-2.0 | 685.4B | 3313 | 35 |
| 48 | DeepSeek-Prover-V1.5-RL | DeepSeek other | 6.9B | 2660 | 2 |
| 49 | deepseek-coder-7b-instruct-v1.5 | DeepSeek other | 6.9B | 2576 | 4 |
| 50 | deepseek-llm-67b-base | DeepSeek other | — | 2076 | 5 |
榜单说明
排序依据:模型在魔搭社区(ModelScope)的累计下载量降序,同等下载量按社区点赞数排序。数据每日同步一次,最近同步时间:2026-09-22 23:31。 下载量受模型体积、上手门槛与发布时长影响,小参数模型因部署成本低往往下载量更高,选型时建议结合参数量、开源协议与任务类型综合判断。 每个模型的 README 说明、下载命令与详情,见模型详情页。
常见问题
推理专用模型和通用模型怎么选?
需要展示解题步骤、做数学与代码推理时选推理专用版本;日常对话、信息抽取与低成本高并发场景选通用版本更划算。
蒸馏版本与小参数版本有什么区别?
蒸馏版是把大模型的推理能力迁移到小模型,响应更快但复杂推理上限较低;按业务对精度与延迟的要求取舍。