DeepSeek开源模型排行榜

本榜收录 DeepSeek 在本站的 71 个开源模型,按社区累计下载量排序。该系列以推理能力与宽松的开源协议受到关注,覆盖通用对话与推理专用方向,并提供多种参数档位与蒸馏版本,适合需要在本地构建推理能力、或对推理过程可解释性有要求的场景。

数据更新时间:2026-09-22 本榜收录 71 个模型 模型库共 1,408 排序依据:累计下载量 ↓
模型总榜 国产大模型榜 海外开源模型榜 通义千问(阿里) AI2(Allen) Microsoft Google 面壁智能 上海AI实验室 智谱AI DeepSeek 智源研究院 魔搭社区 腾讯混元 Mistral AI

下载量排名 TOP50

魔搭社区累计下载
排名 模型 厂商 / 协议 参数量 下载量 点赞
1 DeepSeek-R1-Distill-Qwen-32B DeepSeek mit 32.8B 274.7万 273
2 DeepSeek-V3.1-Terminus DeepSeek mit 684.5B 260.0万 70
3 DeepSeek-R1-Distill-Qwen-1.5B DeepSeek mit 1.8B 222.1万 343
4 DeepSeek-R1-Distill-Llama-70B DeepSeek mit 70.6B 212.4万 141
5 DeepSeek-R1-Distill-Qwen-7B DeepSeek mit 7.6B 91.7万 447
6 DeepSeek-V2-Lite-Chat DeepSeek other 15.7B 79.1万 18
7 DeepSeek-Coder-V2-Lite-Instruct DeepSeek other 15.7B 66.9万 34
8 DeepSeek-V2-Chat DeepSeek other 235.7B 62.3万 42
9 DeepSeek-R1-0528 DeepSeek mit 684.5B 61.1万 330
10 DeepSeek-V4-Flash DeepSeek mit 290.9B 59.0万 400
11 DeepSeek-R1-Distill-Qwen-14B DeepSeek mit 14.8B 47.0万 180
12 DeepSeek-V3.2 DeepSeek 685.4B 43.9万 463
13 DeepSeek-V3 DeepSeek 684.5B 33.7万 266
14 DeepSeek-R1-0528-Qwen3-8B DeepSeek mit 8.2B 29.3万 130
15 DeepSeek-Coder-V2-Instruct DeepSeek other 235.7B 28.9万 7
16 deepseek-coder-6.7b-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 6.7B 25.4万 32
17 DeepSeek-V4-Pro DeepSeek mit 861.6B 24.2万 552
18 DeepSeek-R1-Distill-Llama-8B DeepSeek mit 8.0B 15.6万 80
19 DeepSeek-V2-Lite DeepSeek other 15.7B 13.6万 7
20 deepseek-moe-16b-chat DeepSeek other 16.4B 12.9万 9
21 DeepSeek-V4-Flash-0731 DeepSeek mit 304.2B 11.6万 517
22 deepseek-llm-67b-chat DeepSeek other 10.9万 6
23 DeepSeek-V3.1 DeepSeek mit 684.5B 7.0万 256
24 deepseek-llm-7b-chat DeepSeek other 6.7万 28
25 DeepSeek-V4-Flash-DSpark DeepSeek mit 165.3B 6.0万 35
26 DeepSeek-V3-Base DeepSeek 684.5B 3.5万 44
27 DeepSeek-V3-0324 DeepSeek mit 684.5B 3.3万 297
28 DeepSeek-V2 DeepSeek other 235.7B 1.7万 7
29 DeepSeek-V4-Flash-Base DeepSeek MIT License 292.0B 1.6万 20
30 DeepSeek-V3.2-Speciale DeepSeek 685.4B 1.6万 103
31 deepseek-llm-7b-base DeepSeek other 1.1万 8
32 deepseek-coder-1.3b-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 1.3B 1.1万 9
33 deepseek-coder-33b-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 33.3B 9927 12
34 deepseek-moe-16b-base DeepSeek other 16.4B 9747 2
35 DeepSeek-V4-Pro-DSpark DeepSeek mit 889.5B 9550 62
36 deepseek-coder-5.7bmqa-instruct Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 9297 0
37 deepseek-coder-6.7b-base Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 6.7B 8008 9
38 DeepSeek-V4-Pro-Base DeepSeek MIT License 1600.8B 7924 23
39 DeepSeek-V4-Pro-0813 DeepSeek mit 1650.5B 7903 310
40 deepseek-vl-7b-base DeepSeek other 7.3B 6801 6
41 deepseek-vl-1.3b-base DeepSeek other 2.0B 6372 4
42 deepseek-coder-33b-base Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 33.3B 6361 8
43 deepseek-coder-1.3b-base Deepseek Coder comprises a series of code language models trained on both 87% code and 13% natural language in English and Chinese, with each model pre-trained on 2T tokens. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various benchmarks. DeepSeek other 6310 7
44 DeepSeek-Coder-V2-Lite-Base DeepSeek other 15.7B 5273 1
45 DeepSeek-V2.5-1210 DeepSeek other 235.7B 4541 6
46 DeepSeek-V2-Chat-0628 DeepSeek 235.7B 4149 3
47 DeepSeek-Math-V2 DeepSeek apache-2.0 685.4B 3313 35
48 DeepSeek-Prover-V1.5-RL DeepSeek other 6.9B 2660 2
49 deepseek-coder-7b-instruct-v1.5 DeepSeek other 6.9B 2576 4
50 deepseek-llm-67b-base DeepSeek other 2076 5

榜单说明

排序依据:模型在魔搭社区(ModelScope)的累计下载量降序,同等下载量按社区点赞数排序。数据每日同步一次,最近同步时间:2026-09-22 23:31。 下载量受模型体积、上手门槛与发布时长影响,小参数模型因部署成本低往往下载量更高,选型时建议结合参数量、开源协议与任务类型综合判断。 每个模型的 README 说明、下载命令与详情,见模型详情页。

常见问题

推理专用模型和通用模型怎么选?
需要展示解题步骤、做数学与代码推理时选推理专用版本;日常对话、信息抽取与低成本高并发场景选通用版本更划算。
蒸馏版本与小参数版本有什么区别?
蒸馏版是把大模型的推理能力迁移到小模型,响应更快但复杂推理上限较低;按业务对精度与延迟的要求取舍。