Skip to content

Latest commit

 

History

History
183 lines (176 loc) · 16 KB

File metadata and controls

183 lines (176 loc) · 16 KB

支持的模型

本页面由脚本自动生成。数据源: chitu/config/models/*.yaml。更新命令: python3 script/generate_supported_models_docs.py

开源模型

名称 支持工具调用(内含约束解码) 用法(启动赤兔时追加下列参数) 获取方法
DeepSeek-R1 models=DeepSeek-R1 https://huggingface.co/deepseek-ai/DeepSeek-R1
DeepSeek-R1-Distill-Llama-70B models=DeepSeek-R1-Distill-Llama-70B https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B
DeepSeek-R1-Distill-Qwen-14B models=DeepSeek-R1-Distill-Qwen-14B https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
DeepSeek-R1-FP4 models=DeepSeek-R1-FP4 https://huggingface.co/nvidia/DeepSeek-R1-FP4
DeepSeek-R1-Q4_K_M models=DeepSeek-R1-Q4_K_M https://huggingface.co/bartowski/DeepSeek-R1-GGUF
DeepSeek-R1-bf16 models=DeepSeek-R1-bf16 https://huggingface.co/opensourcerelease/DeepSeek-R1-bf16
DeepSeek-R1-mix-fp4 models=DeepSeek-R1-fp4-all https://www.modelscope.cn/models/qingcheng-ai/DeepSeek-R1-fp4
DeepSeek-R1-mix-fp4 models=DeepSeek-R1-fp4-mix https://www.modelscope.cn/models/qingcheng-ai/DeepSeek-R1-mix
DeepSeek-V3 models=DeepSeek-V3 https://huggingface.co/deepseek-ai/DeepSeek-V3
DeepSeek-V3.1 models=DeepSeek-V3.1 https://huggingface.co/deepseek-ai/DeepSeek-V3.1
DeepSeek-V3.1-Terminus models=DeepSeek-V3.1-Terminus https://huggingface.co/deepseek-ai/DeepSeek-V3.1-Terminus
DeepSeek-V3.2 models=DeepSeek-V3.2 https://huggingface.co/deepseek-ai/DeepSeek-V3.2
DeepSeek-V3.2-Exp models=DeepSeek-V3.2-Exp https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp
DeepSeek-V3.2-Exp-kv-fp8 models=DeepSeek-V3.2-Exp-kv-fp8 https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp
DeepSeek-V3.2-kv-fp8 models=DeepSeek-V3.2-kv-fp8 https://huggingface.co/deepseek-ai/DeepSeek-V3.2
DeepSeek-V4-Flash-Base models=DeepSeek-V4-Flash-Base https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Base
DeepSeek-V4-Flash-FP8 models=DeepSeek-V4-Flash-FP8 https://huggingface.co/sgl-project/DeepSeek-V4-Flash-FP8
DeepSeek-V4-Pro models=DeepSeek-V4-Pro https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
glm-4-32b models=GLM-4-32B-0414 https://modelscope.cn/models/ZhipuAI/GLM-4-32B-0414
glm-4-9b models=GLM-4-9B-0414 https://www.modelscope.cn/models/ZhipuAI/GLM-4-9B-0414
GLM-4.5 models=GLM-4.5 https://huggingface.co/zai-org/GLM-4.5
GLM-4.5-Air models=GLM-4.5-Air https://huggingface.co/zai-org/GLM-4.5-Air
GLM-4.5V models=GLM-4.5V https://huggingface.co/zai-org/GLM-4.5V
GLM-4.6 models=GLM-4.6 https://huggingface.co/zai-org/GLM-4.6
GLM-4.6V models=GLM-4.6V https://huggingface.co/zai-org/GLM-4.6V
GLM-4.7 models=GLM-4.7 https://huggingface.co/zai-org/GLM-4.7
GLM-4.7-FP8 models=GLM-4.7-FP8 https://huggingface.co/zai-org/GLM-4.7-FP8
GLM-4.7-Flash models=GLM-4.7-Flash https://huggingface.co/zai-org/GLM-4.7-Flash
GLM-5 models=GLM-5 https://huggingface.co/zai-org/GLM-5
GLM-5-FP8 models=GLM-5-FP8 https://huggingface.co/zai-org/GLM-5-FP8
GLM-5-FP8-kv models=GLM-5-FP8-kv https://huggingface.co/zai-org/GLM-5-FP8
GLM-5-W8A8 models=GLM-5-W8A8 https://modelscope.cn/models/metax-tech/GLM-5-W8A8/
GLM-5.1 models=GLM-5.1 https://huggingface.co/zai-org/GLM-5.1
GLM-5.1-Channel-INT4-w4a8 models=GLM-5.1-Channel-INT4-w4a8 https://modelscope.cn/models/hygon/GLM-5.1-Channel-INT4-w4a8
GLM-5.1-FP8 models=GLM-5.1-FP8 https://huggingface.co/zai-org/GLM-5.1-FP8
GLM-5.1-W8A8 models=GLM-5.1-W8A8 https://modelscope.cn/models/metax-tech/GLM-5.1-W8A8
GLM-5.2 models=GLM-5.2 https://huggingface.co/zai-org/GLM-5.2
GLM-5.2-FP8 models=GLM-5.2-FP8 https://huggingface.co/zai-org/GLM-5.2-FP8
GLM-5.2-W8A8 models=GLM-5.2-W8A8 https://modelscope.cn/models/metax-tech/GLM-5.2-W8A8
glm-z1-32b models=GLM-Z1-32B-0414 https://modelscope.cn/models/ZhipuAI/GLM-Z1-32B-0414/
glm-z1-9b models=GLM-Z1-9B-0414 https://modelscope.cn/models/ZhipuAI/GLM-Z1-9B-0414
Kimi-K2-Instruct models=Kimi-K2-Instruct https://huggingface.co/moonshotai/Kimi-K2-Instruct
Kimi-K2.5 models=Kimi-K2.5 https://huggingface.co/moonshotai/Kimi-K2.5
Kimi-K2.6 models=Kimi-K2.6 https://huggingface.co/moonshotai/Kimi-K2.6
Kimi-K2.7-Code models=Kimi-K2.7-Code https://huggingface.co/moonshotai/Kimi-K2.7-Code
LLaDA2.1-flash models=LLaDA2.1-flash https://huggingface.co/inclusionAI/LLaDA2.0-flash
LLaDA2.1-mini models=LLaDA2.1-mini https://huggingface.co/inclusionAI/LLaDA2.0-mini
Llama-3-8B-QServe models=Llama-3-8B-QServe https://huggingface.co/mit-han-lab/Llama-3-8B-QServe
Llama-3-8B-QServe-g128 models=Llama-3-8B-QServe-g128 https://huggingface.co/mit-han-lab/Llama-3-8B-QServe-g128
Llama-3.3-70B-Instruct models=Llama-3.3-70B-Instruct https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct
Meta-Llama-3-8B-Instruct models=Meta-Llama-3-8B-Instruct https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct (The default checkpoint. Not the "original" one)
Meta-Llama-3-8B-Instruct-original models=Meta-Llama-3-8B-Instruct-original https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct (Please use its "original" checkpoint)
MiniMax-M3 models=MiniMax-M3 https://huggingface.co/MiniMaxAI/MiniMax-M3
MiniMax-M3-MXFP8 models=MiniMax-M3-MXFP8 https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8
Mixtral-8x7B-Instruct-v0.1 models=Mixtral-8x7B-Instruct-v0.1 https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1
QwQ-32B models=QwQ-32B https://huggingface.co/Qwen/QwQ-32B
QwQ-32B-AWQ models=QwQ-32B-AWQ https://huggingface.co/Qwen/QwQ-32B-AWQ
QwQ-32B-FP8 models=QwQ-32B-FP8 https://huggingface.co/qingcheng-ai/QWQ-32B-FP8
QwQ-32B-GPTQ models=QwQ-32B-GPTQ https://huggingface.co/clowman/QwQ-32B-GPTQ-Int8
QwQ-32B-fp4 models=QwQ-32B-fp4 https://huggingface.co/qingcheng-ai/QwQ-32B-fp4
Qwen2-72B-Instruct models=Qwen2-72B-Instruct https://huggingface.co/Qwen/Qwen2-72B-Instruct
Qwen2-7B-Instruct models=Qwen2-7B-Instruct https://huggingface.co/Qwen/Qwen2-7B-Instruct
Qwen2.5-0.5B models=Qwen2.5-0.5B https://huggingface.co/Qwen/Qwen2.5-0.5B
Qwen2.5-0.5B-Instruct models=Qwen2.5-0.5B-Instruct https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct
Qwen2.5-1.5B models=Qwen2.5-1.5B https://huggingface.co/Qwen/Qwen2.5-1.5B
Qwen2.5-1.5B-Instruct models=Qwen2.5-1.5B-Instruct https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
Qwen2.5-32B models=Qwen2.5-32B https://huggingface.co/Qwen/Qwen2.5-32B
Qwen2.5-32B-Instruct models=Qwen2.5-32B-Instruct https://huggingface.co/Qwen/Qwen2.5-32B-Instruct
Qwen2.5-3B models=Qwen2.5-3B https://huggingface.co/Qwen/Qwen2.5-3B
Qwen2.5-3B-Instruct models=Qwen2.5-3B-Instruct https://huggingface.co/Qwen/Qwen2.5-3B-Instruct
Qwen2.5-7B models=Qwen2.5-7B https://huggingface.co/Qwen/Qwen2.5-7B
Qwen2.5-7B-Instruct models=Qwen2.5-7B-Instruct https://huggingface.co/Qwen/Qwen2.5-7B-Instruct
Qwen2.5-VL-32B-Instruct models=Qwen2.5-VL-32B-Instruct https://huggingface.co/Qwen/Qwen2.5-VL-32B-Instruct
Qwen2.5-VL-7B-Instruct models=Qwen2.5-VL-7B-Instruct https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct
Qwen3-0.6B models=Qwen3-0.6B https://huggingface.co/Qwen/Qwen3-0.6B
Qwen3-1.7B models=Qwen3-1.7B https://huggingface.co/Qwen/Qwen3-1.7B
Qwen3-14B models=Qwen3-14B https://huggingface.co/Qwen/Qwen3-14B
Qwen3-14B-FP8 models=Qwen3-14B-FP8 https://huggingface.co/Qwen/Qwen3-14B-FP8
Qwen3-235B-A22B models=Qwen3-235B-A22B https://huggingface.co/Qwen/Qwen3-235B-A22B
Qwen3-235B-A22B-Instruct models=Qwen3-235B-A22B-Instruct https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507
Qwen3-235B-A22B-fp4 models=Qwen3-235B-A22B-fp4 https://huggingface.co/nvidia/Qwen3-235B-A22B-NVFP4
Qwen3-235B-A22B-fp8 models=Qwen3-235B-A22B-fp8 https://huggingface.co/Qwen/Qwen3-235B-A22B-FP8
Qwen3-235B-A22B-fp8 models=Qwen3-235B-A22B-fp8-kv https://huggingface.co/Qwen/Qwen3-235B-A22B-FP8
Qwen3-30B-A3B models=Qwen3-30B-A3B https://huggingface.co/Qwen/Qwen3-30B-A3B
Qwen3-30B-A3B-fp4 models=Qwen3-30B-A3B-fp4 https://huggingface.co/nvidia/Qwen3-30B-A3B-NVFP4
Qwen3-30B-A3B-fp8 models=Qwen3-30B-A3B-fp8 https://huggingface.co/Qwen/Qwen3-30B-A3B-FP8
Qwen3-30B-A3B-fp8-kv models=Qwen3-30B-A3B-fp8-kv https://huggingface.co/Qwen/Qwen3-30B-A3B-FP8
Qwen3-32B models=Qwen3-32B https://huggingface.co/Qwen/Qwen3-32B
Qwen3-32B-FP8 models=Qwen3-32B-FP8 https://huggingface.co/Qwen/Qwen3-32B-FP8
Qwen3-32B-fp4 models=Qwen3-32B-fp4 https://huggingface.co/qingcheng-ai/Qwen3-32B-fp4
Qwen3-4B models=Qwen3-4B https://huggingface.co/Qwen/Qwen3-4B
Qwen3-8B models=Qwen3-8B https://huggingface.co/Qwen/Qwen3-8B
Qwen3-8B-fp4 models=Qwen3-8B-fp4 https://huggingface.co/qingcheng-ai/Qwen3-8B-fp4
Qwen3-Coder-30B-A3B-Instruct models=Qwen3-Coder-30B-A3B-Instruct https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct
Qwen3-Coder-30B-A3B-Instruct-fp8 models=Qwen3-Coder-30B-A3B-Instruct-fp8 https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8
Qwen3-Coder-480B-A35B-Instruct models=Qwen3-Coder-480B-A35B-Instruct https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct
Qwen3-Coder-480B-A35B-Instruct-fp8 models=Qwen3-Coder-480B-A35B-Instruct-fp8 https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
Qwen3-Coder-Next models=Qwen3-Coder-Next https://huggingface.co/Qwen/Qwen3-Coder-Next
Qwen3-Coder-Next-FP8 models=Qwen3-Coder-Next-FP8 https://huggingface.co/Qwen/Qwen3-Coder-Next-FP8
Qwen3-Next-80B-A3B-Instruct models=Qwen3-Next-80B-A3B-Instruct https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
Qwen3-Next-80B-A3B-Instruct-FP8 models=Qwen3-Next-80B-A3B-Instruct-FP8 https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct-FP8
Qwen3-Next-80B-A3B-Thinking models=Qwen3-Next-80B-A3B-Thinking https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking
Qwen3-VL-235B-A22B-Instruct models=Qwen3-VL-235B-A22B-Instruct https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct
Qwen3-VL-8B-Instruct models=Qwen3-VL-8B-Instruct https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct
Qwen3.5-0.8B models=Qwen3.5-0.8B https://huggingface.co/Qwen/Qwen3.5-0.8B
Qwen3.5-122B-A10B models=Qwen3.5-122B-A10B https://huggingface.co/Qwen/Qwen3.5-122B-A10B
Qwen3.5-122B-A10B-FP8 models=Qwen3.5-122B-A10B-FP8 https://huggingface.co/Qwen/Qwen3.5-122B-A10B-FP8
Qwen3.5-27B models=Qwen3.5-27B https://huggingface.co/Qwen/Qwen3.5-27B
Qwen3.5-27B-FP8 models=Qwen3.5-27B-FP8 https://huggingface.co/Qwen/Qwen3.5-27B-FP8
Qwen3.5-2B models=Qwen3.5-2B https://huggingface.co/Qwen/Qwen3.5-2B
Qwen3.5-35B-A3B models=Qwen3.5-35B-A3B https://huggingface.co/Qwen/Qwen3.5-35B-A3B
Qwen3.5-35B-A3B-FP8 models=Qwen3.5-35B-A3B-FP8 https://huggingface.co/Qwen/Qwen3.5-35B-A3B-FP8
Qwen3.5-397B-A17B models=Qwen3.5-397B-A17B https://huggingface.co/Qwen/Qwen3.5-397B-A17B
Qwen3.5-397B-A17B-FP8 models=Qwen3.5-397B-A17B-FP8 https://huggingface.co/Qwen/Qwen3.5-397B-A17B-FP8
Qwen3.5-4B models=Qwen3.5-4B https://huggingface.co/Qwen/Qwen3.5-4B
Qwen3.5-9B models=Qwen3.5-9B https://huggingface.co/Qwen/Qwen3.5-9B
Qwen3.6-27B models=Qwen3.6-27B https://huggingface.co/Qwen/Qwen3.6-27B
Qwen3.6-27B-FP8 models=Qwen3.6-27B-FP8 https://huggingface.co/Qwen/Qwen3.6-27B-FP8
Qwen3.6-35B-A3B models=Qwen3.6-35B-A3B https://huggingface.co/Qwen/Qwen3.6-35B-A3B
Qwen3.6-35B-A3B-FP8 models=Qwen3.6-35B-A3B-FP8 https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8
Seed-OSS-36B-Instruct models=Seed-OSS-36B-Instruct https://huggingface.co/ByteDance-Seed/Seed-OSS-36B-Instruct
glm-4-9b-chat models=glm-4-9b-chat https://huggingface.co/zai-org/glm-4-9b-chat-hf
gpt-oss-120b-BF16 models=gpt-oss-120b-BF16 https://huggingface.co/unsloth/gpt-oss-120b-BF16
gpt-oss-20b-BF16 models=gpt-oss-20b-BF16 https://huggingface.co/unsloth/gpt-oss-20b-BF16

赤兔-pro 模型

以下模型随赤兔-pro提供,请联系 solution@chitu.ai 进行商务咨询。

名称 支持工具调用(内含约束解码) 用法(启动赤兔时追加下列参数)
DeepSeek-R1-Distill-Llama-70B-ascend-int8 models=DeepSeek-R1-Distill-Llama-70B-ascend-int8
DeepSeek-R1-Distill-Qwen-14B-fp8 models=DeepSeek-R1-Distill-Qwen-14B-fp8
DeepSeek-R1-ascend-int8 models=DeepSeek-R1-int8-ascend
DeepSeek-R1-MXFP4 models=DeepSeek-R1-mxfp4
DeepSeek-R1-w4a8-hygon models=DeepSeek-R1-w4a8-hygon
DeepSeek-V3-ascend-int8 models=DeepSeek-V3-int8-ascend
DeepSeek-V3.1-Terminus-ascend-int8 models=DeepSeek-V3.1-Terminus-int8-ascend
DeepSeek-V4-Flash models=DeepSeek-V4-Flash
GLM-4.5-Air-qc-fp8 models=GLM-4.5-Air-qc-fp8
GLM-4.5-qc-fp8 models=GLM-4.5-qc-fp8
QwQ-32B-simple-w8a8 models=QwQ-32B-simple-w8a8
QwQ-32B-simple-w8a8-muxi models=QwQ-32B-simple-w8a8-muxi
Qwen2.5-3B-Mix models=Qwen2.5-3B-Mix
Qwen2.5-72B-Instruct-ascend-int8 models=Qwen2.5-72B-Instruct-ascend-int8
Qwen2.5-VL-32B-Instruct-ascend-int8 models=Qwen2.5-VL-32B-Instruct-ascend-int8
Qwen3-14B-QServe-g128 models=Qwen3-14B-QServe-g128
Qwen3-14B-ascend-int8 models=Qwen3-14B-ascend-int8
Qwen3-14B-fp4 models=Qwen3-14B-fp4
Qwen3-14B-mixq-mix models=Qwen3-14B-mixq-mix
Qwen3-14B-mixq-w8a8 models=Qwen3-14B-mixq-w8a8
Qwen3-14B-w4-g128-symm-a8 models=Qwen3-14B-w4-g128-symm-a8
Qwen3-235B-A22B-Instruct-ascend-int8 models=Qwen3-235B-A22B-Instruct-ascend-int8
Qwen3-235B-A22B-ascend-int8 models=Qwen3-235B-A22B-ascend-int8
Qwen3-235B-A22B-mxfp4 models=Qwen3-235B-A22B-mxfp4
Qwen3-30B-A3B-mix-fp4-fp8 models=Qwen3-30B-A3B-mix-fp4-fp8
Qwen3-30B-A3B-mix-fp4-fp8 models=Qwen3-30B-A3B-mix-fp4-fp8-merged
Qwen3-30B-A3B-mxfp4 models=Qwen3-30B-A3B-mxfp4
Qwen3-32B-QServe-w4a8-g128 models=Qwen3-32B-QServe-w4a8-g128
Qwen3-32B-ascend-int8 models=Qwen3-32B-ascend-int8
Qwen3-32B-fp4 models=Qwen3-32B-fp4-merged
Qwen3-32B-mixq-mix models=Qwen3-32B-mixq-mix
Qwen3-32B-mxfp4 models=Qwen3-32B-mxfp4
Qwen3-32B-w4-g128-symm-a8 models=Qwen3-32B-w4-g128-symm-a8
Qwen3-4B-fp4 models=Qwen3-4B-fp4
Qwen3-4B-mxfp4 models=Qwen3-4B-mxfp4
Qwen3-8B-ascend-int8 models=Qwen3-8B-ascend-int8
Qwen3-Coder-480B-A35B-Instruct-int8 models=Qwen3-Coder-480B-A35B-Instruct-int8
Qwen3.5-122B-A10B-mxfp4 models=Qwen3.5-122B-A10B-mxfp4
Qwen3.5-27B-mxfp4 models=Qwen3.5-27B-mxfp4
Qwen3.5-35B-A3B-mxfp4 models=Qwen3.5-35B-A3B-mxfp4
Qwen3.5-397B-A17B-mxfp4 models=Qwen3.5-397B-A17B-mxfp4
Qwen3.6-27B-mxfp4 models=Qwen3.6-27B-mxfp4
Qwen3.6-35B-A3B-mxfp4 models=Qwen3.6-35B-A3B-mxfp4