Alibaba Cloud Model Studio · API · 1,000,000 context tokens · 2026-09-20
Official source and availability ↗LLM & SLM
Choosing a language model: large or small?
Compare models, inference software and memory requirements for your workload, with published results and practical guides.
Running a language model on your own computer or server can support chat, coding, translation or answers grounded in documents. Data confidentiality, control over service dependencies and predictable costs are common reasons to consider local deployment. The range of models and execution methods makes the choice less straightforward: which model is sufficient for the task, what fits the available hardware, and when does upgrading the infrastructure make a useful difference?
This guide brings those decisions together. Its interactive tables compare small and large language models, related specialist models, hardware requirements and runtime software, with links to the supporting sources. The accompanying articles explain the practical advantages and limits of each approach, connecting the task and required quality to the resources and budget needed to deliver it.

Where should I start?
Open a table to compare options; read the articles in each path to understand the concepts and choices.
Interactive guide
Start with the key questions; add deployment details later to refine the plan.
Once you know the model and its memory requirements, compare suitable GPUs and servers.
Sources and coverage
Specifications come from model cards, configurations and versioned software documentation. Quality and speed results retain their reporter and source conditions; this site has not run independent model benchmarks.
Candidates cover distinct roles: text generation, coding, images and documents, embedding and reranking. Model size and a multilingual label do not replace evidence for the selected task and language.
The hardware table and most current speed reports focus on NVIDIA. CPU, Apple Silicon and AMD execution depends on the software and backend; the guide lacks matched speed results for ranking all these platforms. Software routes and supported hardware
Memory estimates combine weight-file bytes, KV and a reserve for each device. The conventional KV formula is not applied to hybrid, MLA or unknown architectures. Assumptions are available in the memory section.
Data reviewed: · Report a correction through the résumé contact links
Model catalog
Downloadable-weight models: size, architecture, context and licence. API-only services are not listed here.
Test shown in the result column: MIRACL · nDCG@10 · en
This selection applies to result columns across tables. Execution conditions are available in each model’s details.
Results grouped by test ←Advanced filters 17 controls
Row comparison and conditions
Comparison rule: Side-by-side display is available; calculations require compatible units and version-specific attribution.
Models available through APIs · 4
Self-hostable weights were not verified in this review; these releases are excluded from local deployment suggestions.
Alibaba Cloud Model Studio · API · 1,000,000 context tokens · 2026-09-20
Official source and availability ↗Alibaba Cloud Model Studio · API · 8,192 input tokens · 2026-09-20
Official source and availability ↗Alibaba Cloud Model Studio · API · 8,192 input tokens · 2026-09-20
Official source and availability ↗| Compare | Details | Input → output | Main uses | Licence | Download and run | ||||
|---|---|---|---|---|---|---|---|---|---|
| BGE · BAAI | approximately 0.569 billion · Dense | Text ← Vector | 8,192 tokens | Hybrid document retrieval in RAG | MIT ↗ | ||||
| BGE · BAAI | approximately 0.568 billion · Dense | Text ← Structured data | 8,192 tokens | Reorder retrieved documents | Apache-2.0 ↗ | ||||
| Aya / Cohere · Cohere Labs | approximately 32 billion · Dense | Text ← Text | 131,072 tokens | Multilingual writing and chat with long documents | CC-BY-NC-4.0 + Cohere Acceptable Use Policy ↗ Commercial use subject to licence | ||||
| Aya / Cohere · Cohere Labs | approximately 8 billion · Dense | Text ← Text | 8,192 tokens | Multilingual writing and rewriting assistant | CC-BY-NC-4.0 + Cohere Acceptable Use Policy ↗ Commercial use subject to licence | ||||
| Aya / Cohere · Cohere Labs | approximately 3.35 billion · Dense | Text ← Text | 8,192 tokens | Local chat across languages | CC-BY-NC-4.0 + Cohere Acceptable Use Policy ↗ Commercial use subject to licence | ||||
| SmolLM · Hugging Face | approximately 1.7 billion · Dense | Text ← Text | 8,192 tokens | Prototype a small assistant with function calling | Apache-2.0 ↗ | ||||
| SmolLM · Hugging Face | approximately 0.135 billion · Dense | Text ← Text | 8,192 tokens | Explore instruction following with a very small model | Apache-2.0 ↗ | ||||
| SmolLM · Hugging Face | approximately 0.36 billion · Dense | Text ← Text | 8,192 tokens | Short-text rewriting and summarization | Apache-2.0 ↗ | ||||
| SmolLM · Hugging Face | approximately 3 billion · Dense | Text ← Text | 65,536 tokens Up to 131,072 tokens with Configure YaRN and increase max_position_embeddings | Small assistant with selectable thinking mode | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 0.6 billion · Dense | Text ← Text | 32,768 tokens | Prototype chat with the smallest Qwen3 | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 1.7 billion · Dense | Text ← Text | 32,768 tokens | Small text assistant with reasoning control | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 14.8 billion · Dense | Text ← Text | 32,768 tokens Up to 131,072 tokens with YaRN configuration | Text generation and analysis with dense Qwen3 | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 30.5 billion · 3.3 billion active · Mixture of experts (MoE) | Text ← Text | 32,768 tokens Up to 131,072 tokens with YaRN configuration | Tool-oriented assistant with MoE architecture | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 32.8 billion · Dense | Text ← Text | 32,768 tokens Up to 131,072 tokens with YaRN configuration | Multi-step analysis with a larger dense Qwen3 | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 4 billion · Dense | Text ← Text | 32,768 tokens Up to 131,072 tokens with YaRN configuration | General assistant at four billion parameters | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 8.2 billion · Dense | Text ← Text | 32,768 tokens Up to 131,072 tokens with YaRN configuration | Chat and text tasks with thinking control | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 30.5 billion · 3.3 billion active · Mixture of experts (MoE) | Text ← Text | 262,144 tokens Up to 1,048,576 tokens with YaRN configuration | Repository editing with a tool-using assistant | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 0.6 billion · Dense | Text ← Vector | 32,768 tokens | Vector retrieval with the smallest Qwen3 Embedding | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 4 billion · Dense | Text ← Vector | 32,768 tokens | Multilingual indexing with adjustable vectors | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 8 billion · Dense | Text ← Vector | 32,768 tokens | Document retrieval with the largest listed Qwen3 Embedding | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 0.6 billion · Dense | Text ← Structured data | 32,768 tokens | Lighter reranking in the Qwen3 family | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 4 billion · Dense | Text ← Structured data | 32,768 tokens | Instruction-aware reranking of candidate documents | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 8 billion · Dense | Text ← Structured data | 32,768 tokens | Reranking with the 8B Qwen3 variant | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 8 billion nominal · Dense Total parameter count unverified | Text, Image, Video ← Text | 262,144 tokens | Read images and documents with Qwen3-VL | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 2 billion · approximately 2 billion Language component · Dense · hybrid attention | Text, Image, Video ← Text | 262,144 tokens | Multimodal prototyping with small Qwen3.5 | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 35 billion · 3 billion active · approximately 35 billion Language component · Mixture of experts (MoE) · hybrid attention | Text, Image, Video ← Text | 262,144 tokens Up to 1,010,000 tokens with Context extension with YaRN | Multimodal agent with MoE Qwen3.5 | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 4 billion · approximately 4 billion Language component · Dense · hybrid attention | Text, Image, Video ← Text | 262,144 tokens Up to 1,010,000 tokens with Context extension with YaRN | Text and image processing at four billion parameters | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 9 billion · approximately 9 billion Language component · Dense · hybrid attention | Text, Image, Video ← Text | 262,144 tokens Up to 1,010,000 tokens with Context extension with YaRN | Document and image assistant with dense Qwen3.5 | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 27 billion · approximately 27 billion Language component · Hybrid · hybrid attention | Text, Image, Video ← Text | 262,144 tokens Up to 1,000,000 tokens with Context extension with YaRN | Long workflows with a text–image agent | Apache-2.0 ↗ | ||||
| OLMo · Allen Institute for AI | approximately 7 billion · Dense | Text ← Text | 65,536 tokens | Instruction-following research with inspectable training | Apache-2.0 ↗ | ||||
| DeepSeek · DeepSeek | approximately 8 billion · Dense | Text ← Text | 131,072 tokens | Explore reasoning distilled from R1-0528 | MIT ↗ | ||||
| DeepSeek · DeepSeek | approximately 70 billion · Dense | Text ← Text | 131,072 tokens | Distilled reasoning on a Llama 70B base | MIT + underlying Llama 3.3 terms ↗ | ||||
| DeepSeek · DeepSeek | approximately 1.5 billion · Dense | Text ← Text | 131,072 tokens | Explore reasoning limits in a very small distilled model | MIT ↗ | ||||
| DeepSeek · DeepSeek | approximately 14 billion · Dense | Text ← Text | 131,072 tokens | Problem solving with a mid-sized R1 distillation | MIT ↗ | ||||
| DeepSeek · DeepSeek | approximately 32 billion · Dense | Text ← Text | 131,072 tokens | Analysis and coding with a dense 32B distillation | MIT ↗ | ||||
| DeepSeek · DeepSeek | approximately 7 billion · Dense | Text ← Text | 131,072 tokens | Problem solving with the 7B R1 distillation | MIT ↗ | ||||
| DeepSeek · DeepSeek | 685.397 billion stored elements · Mixture of experts (MoE) Count stored in the checkpoint, including additional components; not active parameters per generated token. | Text ← Text | 163,840 tokens | Connect reasoning and tools in agent workflows | MIT ↗ | ||||
| DeepSeek · DeepSeek | 8 billion Active during prefill · 16 billion Active during decode · approximately 552 billion Main backbone · approximately 196 billion Engram conditional memory · Mixture of experts (MoE) · hybrid attention Total parameter count unverified | Text, Image ← Text | 1,000,000 tokens | Multimodal agent for very long inputs | MIT ↗ | ||||
| Gemma · Google DeepMind | approximately 12 billion · Dense | Text, Image ← Text | 131,072 tokens | Ask about images and text with mid-sized Gemma 3 | gemma ↗ | ||||
| Gemma · Google DeepMind | approximately 1 billion · Dense | Text ← Text | 32,768 tokens | Small text assistant from the Gemma 3 family | gemma ↗ | ||||
| Gemma · Google DeepMind | approximately 27 billion · Dense | Text, Image ← Text | 131,072 tokens | Text and image understanding with the largest listed Gemma 3 | gemma ↗ | ||||
| Gemma · Google DeepMind | approximately 4 billion · Dense | Text, Image ← Text | 131,072 tokens | Getting started with images in the Gemma 3 family | gemma ↗ | ||||
| Gemma · Google DeepMind | approximately 26 billion nominal · approximately 25.2 billion Language component in the publisher's table · Mixture of experts (MoE) Total parameter count unverified | Text, Image, Video ← Text | 262,144 tokens | Multimodal reasoning with MoE Gemma 4 | Apache-2.0 ↗ | ||||
| Gemma · Google DeepMind | approximately 2.3 billion Effective count excluding the embedding table · approximately 5.1 billion Language component including embeddings · approximately 0.3 billion Audio encoder · Dense Total parameter count unverified | Text, Image, Audio, Video ← Text | 131,072 tokens | Local text, image and audio processing with small Gemma 4 | Apache-2.0 ↗ | ||||
| Granite · IBM | approximately 2 billion · Dense | Text ← Text | 131,072 tokens | Generate answers from retrieved documents | Apache-2.0 ↗ | ||||
| E5 · intfloat / multilingual E5 authors | 0.118 billion · Dense | Text ← Vector | 512 tokens | Multilingual retrieval with low-dimensional vectors | MIT ↗ | ||||
| Llama · Meta | approximately 70 billion · Dense | Text ← Text | 131,072 tokens | General assistant with large Llama 3.1 | llama3.1 ↗ | ||||
| Llama · Meta | approximately 8 billion · Dense | Text ← Text | 131,072 tokens | Chat and text work with Llama 3.1 8B | llama3.1 ↗ | ||||
| Llama · Meta | approximately 1 billion · Dense | Text ← Text | 131,072 tokens | Local rewriting and summarization with small Llama | llama3.2 ↗ | ||||
| Llama · Meta | approximately 3 billion · Dense | Text ← Text | 131,072 tokens | On-device text assistant with Llama 3B | llama3.2 ↗ | ||||
| Phi · Microsoft | approximately 3.8 billion · Dense | Text ← Text | 131,072 tokens | Analysis and logic under tighter resource constraints | MIT ↗ | ||||
| Mistral · Mistral AI | approximately 24 billion · Dense | Text, Image ← Text | 262,144 tokens | Multi-file editing and repository-search agent | Apache-2.0 ↗ | ||||
| Mistral · Mistral AI | approximately 3.4 billion Language component · Dense Total parameter count unverified | Text, Image ← Text | 262,144 tokens Published weights are FP8; execution capacity depends on precision and context memory. | Multimodal assistant for edge deployment | Apache-2.0 ↗ | ||||
| Mistral · Mistral AI | approximately 7 billion · Dense | Text ← Text | 32,768 tokens | Text assistant with Mistral function-calling format | Apache-2.0 ↗ | ||||
| Mistral · Mistral AI | approximately 24 billion · Dense | Text, Image ← Text | 131,072 tokens | Multilingual text and image assistant | Apache-2.0 ↗ | ||||
| Nemotron · NVIDIA | approximately 9 billion · Hybrid · hybrid attention | Text ← Text | 131,072 tokens | Document-grounded answers with a reasoning budget | nvidia-open-model-license ↗ | ||||
| gpt-oss · OpenAI | approximately 117 billion · 5.1 billion active · Mixture of experts (MoE) | Text ← Text | 131,072 tokens With the YaRN settings supplied with the published model | Adjustable-reasoning agent with the larger gpt-oss | Apache-2.0 ↗ | ||||
| gpt-oss · OpenAI | approximately 21 billion · 3.6 billion active · Mixture of experts (MoE) | Text ← Text | 131,072 tokens With the YaRN settings supplied with the published model | Local reasoning with the smaller gpt-oss | Apache-2.0 ↗ | ||||
| GLM · Z.ai | approximately 30 billion · 3 billion active · Mixture of experts (MoE) | Text ← Text | 202,752 tokens | Coding agent with preserved reasoning | MIT ↗ | ||||
| Kimi · Moonshot AI | approximately 1,000 billion nominal · 32 billion active · Mixture of experts (MoE) Total parameter count unverified | Text ← Text | 131,072 tokens | Text agent for coding and tool-based workflows | Modified MIT ↗ Commercial use subject to licence | ||||
| Kimi · Moonshot AI | approximately 1,000 billion nominal · 32 billion active · Mixture of experts (MoE) Total parameter count unverified | Text ← Text | 262,144 tokens | Reasoning agent for long tool chains and analysis | Modified MIT ↗ Commercial use subject to licence | ||||
| Kimi · Moonshot AI | approximately 1,000 billion nominal · 32 billion active · Mixture of experts (MoE) Total parameter count unverified | Text, Image, Video ← Text | 262,144 tokens | Multimodal agent for coding from visual designs and document analysis | Modified MIT ↗ Commercial use subject to licence | ||||
| MiniMax · MiniMax | approximately 230 billion nominal · 10 billion active · Mixture of experts (MoE) Total parameter count unverified | Text ← Text | 196,608 tokens | Software development assistant for tool-based workflows | Modified MIT ↗ Commercial use subject to licence | ||||
| MiniMax · MiniMax | approximately 230 billion nominal · 10 billion active · Mixture of experts (MoE) Total parameter count unverified | Text ← Text | 196,608 tokens | Tool-oriented model for coding and multi-step planning | Modified MIT ↗ Commercial use subject to licence | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 235 billion · 22 billion active · Mixture of experts (MoE) | Text ← Text | 262,144 tokens | Large direct-answer assistant with long context | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 480 billion · 35 billion active · Mixture of experts (MoE) | Text ← Text | 262,144 tokens | Coding agent for large, multi-file repositories | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 27 billion · Dense · hybrid attention | Text, Image, Video ← Text | 262,144 tokens | Multimodal assistant for analysis, code and documents | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 0.8 billion · Dense · hybrid attention | Text, Image, Video ← Text | 262,144 tokens | Small model for prototyping and task specialization | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 122 billion · 10 billion active · Mixture of experts (MoE) | Text, Image, Video ← Text | 262,144 tokens | Multimodal MoE model for analysis and tool use | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 397 billion · 17 billion active · Mixture of experts (MoE) | Text, Image, Video ← Text | 262,144 tokens | Large multimodal assistant for complex problems | Apache-2.0 ↗ | ||||
| Llama · Meta | approximately 70 billion nominal · Dense Total parameter count unverified | Text ← Text | 131,072 tokens | Multilingual text assistant for answers and enterprise tasks | llama3.3 ↗ Commercial use subject to licence | ||||
| Phi · Microsoft | approximately 14 billion nominal · Dense Total parameter count unverified | Text ← Text | 16,384 tokens | English model for mathematics, logic and text generation | MIT ↗ | ||||
| SmolVLM · Hugging Face | approximately 2.2 billion nominal · Dense Total parameter count unverified | Text, Image, Video ← Text | 8,192 tokens | Small model for questions about images and video | Apache-2.0 ↗ | ||||
| MiniLM · Sentence Transformers | 0.023 billion · Dense | Text ← Vector | 256 tokens | Lightweight English embeddings for search and clustering | Apache-2.0 ↗ | ||||
| E5 · intfloat / multilingual E5 authors | 0.278 billion · Dense | Text ← Vector | 512 tokens | Multilingual embeddings with separate query and document prefixes | MIT ↗ | ||||
| E5 · intfloat / multilingual E5 authors | 0.56 billion · Dense | Text ← Vector | 512 tokens | Multilingual embeddings for semantic search | MIT ↗ | ||||
| BGE · BAAI | 0.033 billion · Dense | Text ← Vector | 512 tokens | Small English embedding model for document retrieval | MIT ↗ | ||||
| BGE · BAAI | 0.278 billion · Dense | Text ← Structured data | 512 tokens | English and Chinese reranker for search results | MIT ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 7.61 billion · Dense | Text ← Text | 32,768 tokens | Small coding assistant for explaining and correcting code | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 14.7 billion · Dense | Text ← Text | 32,768 tokens | Coding assistant for generation, explanation and debugging | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | approximately 32.5 billion · Dense | Text ← Text | 32,768 tokens | Coding assistant with more capacity for difficult problems | Apache-2.0 ↗ | ||||
| EmbeddingGemma · Google DeepMind | approximately 0.3 billion nominal · Dense Total parameter count unverified | Text ← Vector | 2,048 tokens | Small multilingual embedding model for on-device execution | gemma ↗ Commercial use subject to licence | ||||
| Jina · Jina AI | approximately 0.57 billion nominal · Dense Total parameter count unverified | Text ← Vector | 8,192 tokens | Multilingual embedding model with task-specific adapters | CC-BY-NC-4.0 ↗ Commercial use subject to licence | ||||
| Jina · Jina AI | 0.278 billion · Dense | Text ← Structured data | 1,024 tokens | Multilingual reranker for longer documents | CC-BY-NC-4.0 ↗ Commercial use subject to licence | ||||
| Mixedbread · Mixedbread | 0.335 billion · Dense | Text ← Vector | 512 tokens | English embedding model for retrieval with query instructions | Apache-2.0 ↗ | ||||
| Nomic · Nomic AI | 0.137 billion · Dense | Text ← Vector | 8,192 tokens | English embedding model with long context and reducible dimensions | Apache-2.0 ↗ | ||||
| ModernBERT · Answer.AI / LightOn | 0.15 billion · Dense | Text ← Structured data | 8,192 tokens | Base encoder for training classification and entity extraction | Apache-2.0 ↗ | ||||
| Qwen · Qwen | 1.54 billion · Dense | Text ← Text | 32,768 tokens | Code completion and fill-in-the-middle | apache-2.0 ↗ | ||||
| StarCoder2 · bigcode | 3 billion · Dense | Text ← Text | 16,384 tokens Local attention with a 4,096-token window. | Code completion with sliding-window attention | bigcode-openrail-m ↗ Commercial use subject to licence | ||||
| Qwen · Qwen | 80 billion · 3 billion active · Mixture of experts (MoE) · hybrid attention | Text ← Text | 262,144 tokens | Coding agent | apache-2.0 ↗ | ||||
| E5 · intfloat | 0.56 billion · Dense | Text ← Vector | 512 tokens | Multilingual retrieval with task instructions | mit ↗ | ||||
| Qwen · Qwen | 4 billion · Dense | Text ← Text | 262,144 tokens | Small direct-answer assistant | apache-2.0 ↗ | ||||
| Tooka · PartAI | 0.123 billion · Dense | Text ← Vector | 512 tokens | Persian text embeddings | — | ||||
| Tooka · PartAI | 0.353 billion · Dense | Text ← Vector | 512 tokens | Persian text embeddings | — | ||||
| ParsBERT · HooshvareLab | Dense Parameter count not recorded | Text ← Structured data | 512 tokens | Training base for Persian tasks | — | ||||
| Salamandra · BSC-LT | 2.253 billion · Dense | Text ← Text | 8,192 tokens | Iberian-language text generation | Apache-2.0 ↗ | ||||
| Salamandra · BSC-LT | 7.768 billion · Dense | Text ← Text | 8,192 tokens | Iberian-language text generation | Apache-2.0 ↗ | ||||
| MiniCPM · openbmb | 2.517 billion · Dense | Text ← Text | 131,072 tokens | Reasoning with a compact model | Apache-2.0 ↗ | ||||
| LFM · LiquidAI | 1.17 billion · Hybrid | Text ← Text | 32,768 tokens | On-device information extraction | LFM Open License 1.0 ↗ Commercial use subject to licence | ||||
| Granite · ibm-granite | 3.66 billion · Dense | Text ← Text | 131,072 tokens | Reasoning with a compact model | Apache-2.0 ↗ | ||||
| GLM · Z.ai | 744 billion · 40 billion active · Mixture of experts (MoE) | Text ← Text | 202,752 tokens | Coding and tool use | MIT ↗ | ||||
| GLM · Z.ai | 753.864 billion stored elements · Mixture of experts (MoE) Count stored in the checkpoint, including additional components; not active parameters per generated token. | Text ← Text | 202,752 tokens | Coding and tool use | MIT ↗ | ||||
| GLM · Z.ai | 753.33 billion stored elements · Mixture of experts (MoE) Count stored in the checkpoint, including additional components; not active parameters per generated token. | Text ← Text | 1,048,576 tokens | Coding and tool use | MIT ↗ | ||||
| GLM · Z.ai | Mixture of experts (MoE) Parameter count not recorded | Text ← Text | 1,048,576 tokens | Coding and tool use | GLM-5.3 License ↗ Commercial use subject to licence | ||||
| Kimi · Moonshot AI | 1,000 billion · 32 billion active · Mixture of experts (MoE) | Text, Image, Video ← Text | 262,144 tokens | Coding and tool use | Modified MIT ↗ Commercial use subject to licence | ||||
| Kimi · Moonshot AI | 1,000 billion · 32 billion active · Mixture of experts (MoE) | Text, Image, Video ← Text | 262,144 tokens | Coding and tool use | Modified MIT ↗ Commercial use subject to licence | ||||
| Kimi · Moonshot AI | 2,800 billion · 104 billion active · Mixture of experts (MoE) | Text, Image, Video ← Text | 1,048,576 tokens | Coding and tool use | Kimi K3 License ↗ Commercial use subject to licence | ||||
| Qwen · Qwen | 2 billion · Dense | Text, Image, Video ← Vector | 32,768 tokens | Multimodal retrieval | Apache-2.0 ↗ | ||||
| Qwen · Qwen | 8 billion · Dense | Text, Image, Video ← Vector | 32,768 tokens | Multimodal retrieval | Apache-2.0 ↗ | ||||
| Qwen · Qwen | 2 billion · Dense | Text, Image, Video ← Structured data | 32,768 tokens | Multimodal reranking | Apache-2.0 ↗ | ||||
| Qwen · Qwen | 8 billion · Dense | Text, Image, Video ← Structured data | 32,768 tokens | Multimodal reranking | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | 27.781 billion · approximately 27 billion Language component; publisher rounded count · Dense · hybrid attention | Text, Image, Video ← Text | 262,144 tokens Up to 1,010,000 tokens with Requires YaRN configuration; native context is 262,144 tokens. | Text and vision model with hybrid attention | Apache-2.0 ↗ | ||||
| Qwen · Qwen / Alibaba Cloud | 35.952 billion · 3 billion active · approximately 35 billion Language component; publisher rounded count · Mixture of experts (MoE) · hybrid attention | Text, Image, Video ← Text | 262,144 tokens Up to 1,010,000 tokens with Requires YaRN configuration; native context is 262,144 tokens. | Text and vision model with hybrid attention | Apache-2.0 ↗ | ||||
| HY-MT · Tencent | 1.8 billion · Dense | Text ← Text | — | Dedicated text translation | Tencent HY Community License ↗ Commercial use subject to licence | ||||
| HY-MT · Tencent | 7 billion · Dense | Text ← Text | — | Dedicated text translation | Tencent HY Community License ↗ Commercial use subject to licence | ||||
| HY-MT · Tencent | 1.8 billion · Dense | Text ← Text | — | Specialized translation | Apache-2.0 ↗ | ||||
| HY-MT · Tencent | 7 billion · Dense | Text ← Text | — | Specialized translation | Apache-2.0 ↗ | ||||
| HY-MT · Tencent | 30 billion · 3 billion active · Mixture of experts (MoE) | Text ← Text | — | Specialized translation | Apache-2.0 ↗ | ||||
| Gemma · Google | 4 billion · Dense | Text, Image ← Text | — | Specialized translation | Gemma terms ↗ Commercial use subject to licence | ||||
| Gemma · Google | 12 billion · Dense | Text, Image ← Text | — | Specialized translation | Gemma terms ↗ Commercial use subject to licence | ||||
| Gemma · Google | 27 billion · Dense | Text, Image ← Text | — | Specialized translation | Gemma terms ↗ Commercial use subject to licence | ||||
| MADLAD · Google | 3 billion · Dense | Text ← Text | — | Specialized translation | Apache-2.0 ↗ | ||||
| NLLB · Meta | 0.6 billion · Dense | Text ← Text | — | Specialized translation | CC-BY-NC-4.0 ↗ Commercial use subject to licence | ||||
| Qwen-Image · Qwen / Alibaba | approximately 7 billion DiT image generator · approximately 8 billion Qwen3-VL encoder · Other Total parameter count unverified | Text, Image ← Image | — | Image generation and editing | Qwen Research License — non-commercial ↗ Commercial use subject to licence | ||||
| MiMo · Xiaomi MiMo | 1,020 billion · 42 billion active · Mixture of experts (MoE) · hybrid attention | Text, Image, Video, Audio ← Text | 1,048,576 tokens | Coding, software agents, text and image analysis | MIT ↗ |
“—” means no information is recorded, not that the model lacks the capability.
Model fit for the task
Models for conversation, programming, search and document tasks.
Test shown in the result column: MIRACL · nDCG@10 · en
This selection applies to result columns across tables. Execution conditions are available in each model’s details.
Results grouped by test ←Which model should I consider?
Documented starting points, not a quality ranking.
Spanish retrieval
E5-Small is a smaller baseline; Large-Instruct scores higher in this Spanish MIRACL report. The Qwen results here are multilingual aggregates, not Spanish-specific scores.
When queries and documents are both Spanish; identify cross-language retrieval separately.
The 51.2-to-53.7 difference establishes neither final-answer quality nor coverage of every Spanish variety.
English retrieval
Treat Large and Large-Instruct as separate variants: Large scores higher on English MIRACL in this report. For a reranking step, use Qwen’s comparison with a common candidate pool.
When target-language evidence matters more than a multilingual average.
The Instruct label does not guarantee an improvement on every task.
Adding reranking
Start the size comparison with 0.6B. Consider 4B and 8B when their task-specific gains justify a second-stage resource budget; 8B does not lead on every metric.
When relevant documents reach the candidate pool but rank poorly; candidate count is part of the design.
Reranking cannot recover a document absent from its input pool.
Coding problems or code edits?
For an instruction-following assistant, distinguish repository editing from coding problems. StarCoder2-3B serves code completion/FIM; HumanEval does not measure chat or editing-agent quality.
Once the interface is known: editor completion, multi-file changes, or solving a programming problem.
MoE active parameters do not determine resident weight memory.
A small model for a bounded task
LFM is a small candidate for extraction and bounded workflows; MiniCPM and Granite also have reasoning and tool-use evidence. Actual weight size can differ from the rounded size in a model name.
When task scope, output format and memory constraints are defined.
A strong small-model score establishes neither lower latency nor laptop feasibility at maximum context.
Selecting a Spanish-language generator
Salamandra supplies direct evidence for several Spanish tasks. Keep it alongside general multilingual candidates; the available data does not identify a best Spanish chatbot.
When Spanish evidence must be distinguished from a general multilingual claim.
Language coverage is not evidence for every region or professional domain.
Intent and text classification
Use MassiveIntent and MTOP evidence for intent classification. An embedding model is one component of the classifier; the result also depends on the classifier and its training data.
For a defined set of labels; the published evidence supports an initial shortlist.
An embedding model or bare ParsBERT is not a ready-to-use chatbot.
A 3B model with two thinking modes. For short replies, choose non-thinking mode and an explicit output budget.
en: Publisher declarationSmolLM3-3B — model card · Publisher report
https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Revision / commit
- a07cc9a04f16550a088caea529712d1d335b0ac1
- Relevant source section
- README.md: model description / architecture / intended use; matching lines 26, 33, 345
- Model announcement counts are rounded; safetensors.total counts stored elements, not necessarily unique parameters.
SmolLM3-3B — Hub metadata · Publisher report
https://huggingface.co/api/models/HuggingFaceTB/SmolLM3-3B?blobs=true ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Revision / commit
- a07cc9a04f16550a088caea529712d1d335b0ac1
- Relevant source section
- $.sha; $.id; $.pipeline_tag; $.cardData; $.safetensors; $.siblings
- A file inventory and tensor count do not equal runtime memory or unique parameters.
SmolLM3-3B — parameter specification · Publisher report
https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Revision / commit
- a07cc9a04f16550a088caea529712d1d335b0ac1
- Relevant source section
- README.md: parameter count / named model variant; matching lines 14, 35, 59, 90, 151, 160, 210, 216, 221, 225, 246, 261
- Model announcement counts are rounded; safetensors.total counts stored elements, not necessarily unique parameters.
SmolLM3-3B — license declaration · Publisher report
https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Revision / commit
- a07cc9a04f16550a088caea529712d1d335b0ac1
- Relevant source section
- README.md: license / underlying-model terms; Hub $.cardData.license; matching lines 3, 31, 381, 382
SmolLM3-3B — configuration · Publisher report
https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/config.json ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Revision / commit
- a07cc9a04f16550a088caea529712d1d335b0ac1
- Relevant source section
- config.json: architectures, model_type, text_config, max_position_embeddings, quantization_config, torch_dtype/dtype
SmolLM3 release · Publisher report
https://huggingface.co/blog/smollm3 ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Relevant source section
- Publication date July 8 2025; SmolLM3-3B
- Release date of the named variants, not repository creation or data extraction.
SmolLM3-3B — declared context · Publisher report
https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Relevant source section
- Key features; Long context processing: trained at 64k; YaRN for 128k
- Documented capacity of the named variant; separate from a test's output length or a hosting API's limit
HuggingFaceTB/SmolLM3-3B — model card · Publisher report
https://huggingface.co/HuggingFaceTB/SmolLM3-3B/raw/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Relevant source section
- Description; model details; usage
HuggingFaceTB/SmolLM3-3B — repository metadata · Publisher report
https://huggingface.co/api/models/HuggingFaceTB/SmolLM3-3B?blobs=true ↗- Publisher / author
- Hugging Face
- Accessed
- 2026-09-15
- Relevant source section
- cardData; sha; safetensors; siblings
Advanced filters 3 controls
Row comparison and conditions
Comparison rule: Models may differ; task, language, dataset, benchmark version, metric and unit must be matched or disclosed.
| Compare | Details | Role in the system | Primary use | Distinguishing feature and rationale | Relevant usage condition | Recommendation basis | Get started | |
|---|---|---|---|---|---|---|---|---|
| Document retrieval | Hybrid document retrieval in RAG | Dense, sparse and multi-vector outputs in one model; over 100 languages, 8,192-token inputs and 1,024-dimensional dense vectors. | Use FlagEmbedding for all three outputs together; GGUF or Ollama output depends on the backend. | Guide's analytical recommendation | ||||
| Document reranking | Reorder retrieved documents | Reads the query and document together to score relevance; complements BGE-M3 retrieval. | Retrieve a limited candidate set first. A sigmoid score is not answer correctness probability; the output is not an embedding. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction | Multilingual writing and chat with long documents | Aya Expanse 32B has a 131,072-token context; Persian appears among the publisher's 23 declared languages. | Research release under a noncommercial license; context length does not establish uniform accuracy throughout a document. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction | Multilingual writing and rewriting assistant | Aya Expanse 8B combines an 8,192-token context with multilingual preference training; it is also a candidate for Persian writing evaluation. | Observe this variant's noncommercial license and context cap; 32B results do not transfer to it. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming | Local chat across languages | Tiny Aya Global provides broad language coverage; it is distinct from the regional Earth, Fire and Water variants. | Evaluate Persian examples for Persian use; regional or base variants cannot be substituted for this checkpoint without review. | Guide's analytical recommendation | ||||
| Tool calling, Text generation, Document-grounded generation, Structured extraction, Programming | Prototype a small assistant with function calling | The largest SmolLM2 variant listed supports a function-calling format as well as instruction following; the two smaller variants do not inherit that feature. | Primarily an English model. The host application executes functions and must validate arguments. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction | Explore instruction following with a very small model | The 135M-parameter SmolLM2 variant is suitable for simple text-generation prototypes and exploring very small models' limits. | Keep inputs short and tasks narrow; do not expect broad knowledge or the 1.7B variant's tool calling. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction | Short-text rewriting and summarization | SmolLM2 360M uses SFT and DPO for instruction tuning; its model card includes a CPU execution example. | English-focused; small model size does not establish speed or Persian quality. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Small assistant with selectable thinking mode | SmolLM3 3B offers direct-answer and reasoning modes, a training context of 65,536 tokens and published training details. | Longer context requires YaRN. The card names six native languages including German, while repository tags list eight different languages; neither list includes Persian. Requires Transformers 4.53 or later., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Prototype chat with the smallest Qwen3 | Dense Qwen3 0.6B offers thinking and non-thinking modes; an official Q8_0 file is available. | Limit the output budget; the family name does not establish the problem-solving ability of larger variants., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Small text assistant with reasoning control | Qwen3-1.7B is among the sub-2B options listed; its message template can select direct-answer mode. | In thinking mode, reasoning tokens count toward cost and output length. This record's official GGUF is Q8_0., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Text generation and analysis with dense Qwen3 | Qwen3-14B is a dense model with two answer modes; official GGUF variants offer a choice of weight precision. | Q4_K_M and Q8_0 packages are available; their speed and quality have not been tested by this guide., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming | Tool-oriented assistant with MoE architecture | Qwen3-30B-A3B uses selected experts and the Qwen-Agent pattern for connecting tools. | Active parameters do not replace total weight size in memory estimates; execution requires compatible message templates and parsers., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Reasoning, Programming, Text generation, Document-grounded generation, Structured extraction, Tool calling | Multi-step analysis with a larger dense Qwen3 | Unlike 30B-A3B, Qwen3-32B is dense; reasoning mode can be chosen according to question complexity. | Compare direct-answer and thinking modes under matched output budgets and conditions; this overview is not a quality ranking., This is a general model; the specialized Coder has different training and tool formats., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation, Publisher-documented use | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | General assistant at four billion parameters | Qwen3-4B switches thinking mode within one checkpoint; official GGUF files offer several quantizations. | Native context is 32,768 tokens; extending it requires context-extension settings., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Chat and text tasks with thinking control | Qwen3-8B is a general dense model with 8.2B declared parameters; its message template selects direct answers or reasoning. | Official Q4_K_M and Q8_0 files are available for local use; this guide has not measured their quality or runtime memory., Retrieval, document insertion and citation checks belong to the system; this relationship is not a RAG or Persian benchmark., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Programming, Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning | Repository editing with a tool-using assistant | Qwen3-Coder-30B-A3B-Instruct is tuned for code generation, repository editing and coding-agent workflows, with a native context of 262,144 tokens. | This is a non-thinking variant. Use Coder's tool-calling format; it is distinct from general Qwen3-30B-A3B., The host application supplies file, testing and execution tools., This checkpoint produces direct answers only; it has no separate thinking mode. | Guide's analytical recommendation, Publisher-documented use | ||||
| Document retrieval | Vector retrieval with the smallest Qwen3 Embedding | The 0.6B variant produces instruction-aware embeddings with up to 1,024 dimensions and 32,768-token inputs. | Add the task instruction to the query; index documents using the same model and dimension settings. | Guide's analytical recommendation | ||||
| Document retrieval | Multilingual indexing with adjustable vectors | Qwen3-Embedding-4B outputs up to 2,560 dimensions and uses query instructions to adapt representations to a task. | Dimension reduction requires checking retrieval quality on the intended collection; an MTEB score cannot substitute for that check. | Guide's analytical recommendation | ||||
| Document retrieval | Document retrieval with the largest listed Qwen3 Embedding | Qwen3-Embedding-8B supports up to 4,096 dimensions and 32,768 tokens of context; text goes in and vectors come out. | Account for indexing and vector storage separately from answer generation; this model does not write the final answer. | Guide's analytical recommendation | ||||
| Document reranking | Lighter reranking in the Qwen3 family | Qwen3-Reranker 0.6B is tuned to score query–document relevance with task instructions. | Requires the reranker template and yes/no scoring; do not use the chat or embedding-generation path. | Guide's analytical recommendation | ||||
| Document reranking | Instruction-aware reranking of candidate documents | Qwen3-Reranker-4B is the middle option; the query, retrieval instruction and document all contribute to relevance scoring. | Pass only candidate documents to this stage; more query–document pairs increase second-stage cost. | Guide's analytical recommendation | ||||
| Document reranking | Reranking with the 8B Qwen3 variant | Qwen3-Reranker-8B is the largest reranker in this collection; it scores query–document relevance. | To choose between this and 4B, compare retrieval quality and latency on your documents; size alone is not a ranking. | Guide's analytical recommendation | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming | Read images and documents with Qwen3-VL | Qwen3-VL 8B Instruct combines text, image and video inputs for visual understanding and document tasks. | Use the multimodal processor and message template; Thinking and Instruct checkpoints behave differently. | Guide's analytical recommendation | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Multimodal prototyping with small Qwen3.5 | Small Qwen3.5-2B is presented for prototyping and task-specific fine-tuning; it accepts visual inputs alongside text. | Larger family variants' quality does not transfer to this 2B model., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Tool calling, Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming | Multimodal agent with MoE Qwen3.5 | The 35B-A3B variant uses an integrated text–image foundation and hybrid architecture, offering a candidate for multi-step workflows. | Hosted Qwen3.5-Flash has different tools and default context; API specifications do not transfer to local weights., Use the multimodal settings and image limits documented for these weights., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation, Publisher-documented use | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Text and image processing at four billion parameters | Qwen3.5-4B belongs to the integrated text–image generation; it is a smaller-than-9B candidate for document information extraction. | Scanned documents require the multimodal path; the text context limit does not determine how many images can be processed., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Document and image assistant with dense Qwen3.5 | Qwen3.5 9B combines text and images in a dense architecture. | Compare against 4B with the same document data and image settings; distinguish language and vision parameter counts., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Tool calling, Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming | Long workflows with a text–image agent | Qwen3.8-27B is a dense text–image model with thinking control; its announcement emphasizes coding, research and multi-step work. | Promised cloud-service features, such as built-in tools and default context, are not guarantees for local weights., Image dimensions and frame sampling are part of the usage conditions., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation, Publisher-documented use | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Instruction-following research with inspectable training | Olmo 3 7B Instruct includes Dolma 3 and Dolci data; published training details distinguish it for reproducible research. | This checkpoint is Instruct; do not attribute Olmo Think results to it. Requires Transformers 4.57 or later. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Explore reasoning distilled from R1-0528 | Post-trained from Qwen3-8B-Base using DeepSeek-R1-0528 reasoning chains; distinct from standard Qwen3-8B. | Full R1-0528 benchmark results do not belong to this 8B variant., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming | Distilled reasoning on a Llama 70B base | This checkpoint distills R1 onto a Llama base and is the largest distilled variant listed. | Both the Llama base license and distillation publisher's terms matter; estimate weight size and reasoning-chain length separately., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming | Explore reasoning limits in a very small distilled model | The 1.5B variant distills R1 onto Qwen2.5-Math, allowing exploration of reasoning behavior transferred to a small model. | Long responses can become repetitive., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming | Problem solving with a mid-sized R1 distillation | DeepSeek-R1-Distill-Qwen-14B uses Qwen2.5-14B and R1-generated data; its weights are not the full R1 model. | Reported numbers apply to this variant and evaluation settings; avoid unmatched comparisons with direct-answer modes., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming | Analysis and coding with a dense 32B distillation | The Qwen-32B variant distills R1 onto a Qwen2.5 base. | Declared reasoning ability is not tool calling or automatic code execution; the application must provide tools and answer validation., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming | Problem solving with the 7B R1 distillation | Based on Qwen2.5-Math-7B; distinct from the later R1-0528 distillation onto Qwen3. | Keep this repository's template and tokenizer; comparison with the 8B variant requires a shared evaluation., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming | Connect reasoning and tools in agent workflows | DeepSeek-V3.2 combines DSA sparse attention with agent-oriented post-training; its chat template differs from earlier versions. | Use V3.2-specific templates and execution paths; Speciale results do not transfer to this model., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Tool calling, Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming | Multimodal agent for very long inputs | DeepSeek-V4.1-Flash processes text and images with CED architecture; active parameter counts differ between prefill and decode. | Active parameters are not a single fixed count; KV compression and a specialized execution path are part of this architecture., Use the CED-specific execution path; this relationship is not an OCR evaluation result., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation, Publisher-documented use | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming | Ask about images and text with mid-sized Gemma 3 | Gemma 3 12B IT is an instruction-tuned multimodal variant with 131,072 tokens of context, connecting image inputs to text answers. | Requires the vision processor and Gemma 3 template; maximum context does not establish constant long-document comprehension. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming | Small text assistant from the Gemma 3 family | Unlike larger variants in the generation, Gemma 3 1B IT is text-only and has 32,768 tokens of context. | Scanned images first require external OCR; the multimodal features of 4B and above do not apply to this variant. | Guide's analytical recommendation | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming | Text and image understanding with the largest listed Gemma 3 | Gemma 3 27B IT is this generation's larger dense variant with visual input and multilingual coverage. | Downloading weights requires accepting the Gemma license. | Guide's analytical recommendation | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming | Getting started with images in the Gemma 3 family | Gemma 3 4B IT is the smallest multimodal model of this generation listed here; unlike 1B, it can read images alongside text. | Download the 4B weights and processor together; validate extracted document numbers in the text output. | Guide's analytical recommendation | ||||
| Reasoning, Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Multimodal reasoning with MoE Gemma 4 | Gemma 4 26B-A4B is an expert model with thinking control and 262,144 tokens of context; its architecture differs from small E2B. | This variant has no audio input; active parameters differ from all model weights., Requires the image processor path; E2B audio support does not transfer to this model., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation, Publisher-documented use | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Local text, image and audio processing with small Gemma 4 | Gemma 4 E2B is presented for on-device execution and also supports audio input. E2B is an effective count, not all weights. | Use total counts and actual files for memory calculations; audio and image paths need their matching processors., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Document-grounded generation, Structured extraction, Text generation, Programming, Tool calling | Generate answers from retrieved documents | Granite 3.3 2B Instruct is a small business-oriented model supporting RAG, summarization, text extraction and thinking mode. | The system must retrieve documents and place them in the prompt; the model is not a search engine. Persian is not among its 12 declared languages., Validate output against the application's schema; requesting a format alone does not guarantee valid JSON. | Guide's analytical recommendation, Publisher-documented use | ||||
| Document retrieval | Multilingual retrieval with low-dimensional vectors | Multilingual E5 Small converts text to 384-dimensional vectors with a 512-token input cap, making it a candidate for short document chunks. | query: and passage: prefixes are required even outside English; split long text before indexing. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | General assistant with large Llama 3.1 | Llama 3.1 70B Instruct supports multilingual chat and text tasks with 131,072 tokens of context. | This variant takes text input; weight use is subject to the Llama license. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Chat and text work with Llama 3.1 8B | Llama 3.1 8B is tuned for general assistance and instruction following; it differs from the base model of the same size. | Keep the Instruct template and tokenizer; long context alone does not guarantee document-grounded answer quality. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Local rewriting and summarization with small Llama | Llama 3.2 1B Instruct is this generation's small text model, presented for narrow text tasks on local devices. | Does not accept images directly; Llama 3.2 Vision models are separate products. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | On-device text assistant with Llama 3B | Llama 3.2 3B Instruct is tuned for chat, rewriting and summarization, with more parameters than the 1B variant. | Compare quality and latency on the intended task; neither the 1B nor the 3B text model accepts images. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Analysis and logic under tighter resource constraints | Phi-4-mini-instruct emphasizes reasoning data and instruction following, with 131,072 tokens of context. | Distinct from mini-reasoning and multimodal-instruct; compare mathematical or logical examples using this exact variant name. | Guide's analytical recommendation | ||||
| Programming, Tool calling, Text generation, Document-grounded generation, Structured extraction, Image understanding | Multi-file editing and repository-search agent | Devstral Small 2 targets software engineering and tool-assisted code inspection and editing; this generation also accepts images. | This repository's Instruct weights are FP8; use the Mistral format and dependencies for this revision, not generic chat-model settings., Agent tools and Mistral templates must match the model revision., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation, Publisher-documented use | ||||
| Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Multimodal assistant for edge deployment | Ministral 3 3B Instruct is the family's small text–image model; the language component has 3.4B parameters and the vision encoder 0.4B. | Official weights are FP8; 3B in the model name does not count every language and vision parameter. | Guide's analytical recommendation | ||||
| Tool calling, Text generation, Document-grounded generation, Structured extraction | Text assistant with Mistral function-calling format | Mistral 7B Instruct v0.3 adds the third tokenizer version and function calling to this text model. | Use the tool format and v3 tokenizer with this checkpoint; this variant has no image input. | Guide's analytical recommendation | ||||
| Image understanding, Structured extraction, Text generation, Document-grounded generation, Programming, Tool calling | Multilingual text and image assistant | Mistral Small 3.1 24B Instruct adds image processing and long context to Small; Persian appears in the publisher's language list. | Tools and JSON output need a matching template and parser; do not mix Base and Instruct results., Configure the parser and schema validation in the application layer. | Guide's analytical recommendation, Publisher-documented use | ||||
| Document-grounded generation, Text generation, Structured extraction, Reasoning, Programming, Tool calling | Document-grounded answers with a reasoning budget | Nemotron Nano 9B v2 combines Mamba-2 with attention; reasoning mode and its token budget are controllable. | Retrieved documents must be supplied externally; Persian is not among the six declared languages, and the backend must support the hybrid architecture., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning | Adjustable-reasoning agent with the larger gpt-oss | gpt-oss-120b is an MoE model with 117B total and 5.1B active parameters; expert weights are released in MXFP4. | Correct execution requires harmony format; the host supplies browser or Python tools. Publisher memory claims are not site measurements., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Reasoning, Text generation, Document-grounded generation, Structured extraction, Tool calling | Local reasoning with the smaller gpt-oss | gpt-oss-20b has 21B total and 3.6B active parameters, targets local or specialized use and offers three reasoning levels. | Requires harmony format and an MXFP4-compatible backend; do not substitute the 20B name for actual weight counts in calculations., For simple answers, disable thinking where supported or limit the output budget. | Guide's analytical recommendation | ||||
| Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming | Coding agent with preserved reasoning | GLM-4.7-Flash is a 30B-A3B-class MoE model; its card describes preserving thinking between turns for multi-step agents. | vLLM and SGLang paths depend on specific development versions., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Text agent for coding and tool-based workflows | — | — | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Reasoning agent for long tool chains and analysis | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding | Multimodal agent for coding from visual designs and document analysis | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Software development assistant for tool-based workflows | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Tool-oriented model for coding and multi-step planning | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Large direct-answer assistant with long context | — | This checkpoint produces direct answers only; it has no separate thinking mode. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Coding agent for large, multi-file repositories | — | This checkpoint produces direct answers only; it has no separate thinking mode. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding | Multimodal assistant for analysis, code and documents | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding | Small model for prototyping and task specialization | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding | Multimodal MoE model for analysis and tool use | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding | Large multimodal assistant for complex problems | — | For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling | Multilingual text assistant for answers and enterprise tasks | — | — | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming | English model for mathematics, logic and text generation | — | — | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Image understanding | Small model for questions about images and video | — | Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. | Guide's analytical recommendation | ||||
| Document retrieval | Lightweight English embeddings for search and clustering | — | Mean token pooling with attention mask and L2 normalization | Publisher-documented use | ||||
| Document retrieval | Multilingual embeddings with separate query and document prefixes | — | Masked mean pooling and L2; query: for queries and passage: for documents | Publisher-documented use | ||||
| Document retrieval | Multilingual embeddings for semantic search | — | Masked mean pooling and L2; query: for queries and passage: for documents | Publisher-documented use | ||||
| Document retrieval | Small English embedding model for document retrieval | — | CLS pooling and normalization; retrieval instruction on queries only | Publisher-documented use | ||||
| Document reranking | English and Chinese reranker for search results | — | Cross-encoder on query–document pairs; relevance score | Publisher-documented use | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming | Small coding assistant for explaining and correcting code | — | — | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming | Coding assistant for generation, explanation and debugging | — | — | Guide's analytical recommendation | ||||
| Text generation, Document-grounded generation, Structured extraction, Programming | Coding assistant with more capacity for difficult problems | — | — | Guide's analytical recommendation | ||||
| Document retrieval | Small multilingual embedding model for on-device execution | — | Over 100 languages; Matryoshka dimension reduction down to 128 | Publisher-documented use | ||||
| Document retrieval | Multilingual embedding model with task-specific adapters | — | retrieval.query / retrieval.passage adapters; mean pooling and L2 | Publisher-documented use | ||||
| Document reranking | Multilingual reranker for longer documents | — | Multilingual reranking with inputs up to 1,024 tokens; noncommercial license | Publisher-documented use | ||||
| Document retrieval | English embedding model for retrieval with query instructions | — | Query instruction: Represent this sentence for searching relevant passages: | Publisher-documented use | ||||
| Document retrieval | English embedding model with long context and reducible dimensions | — | search_query: prefix for queries and search_document: for documents; normalization and dimension reduction | Publisher-documented use | ||||
| Structured extraction | Base encoder for training classification and entity extraction | — | Bidirectional encoder; token output or trained task head | Publisher-documented use | ||||
| Code completion / FIM | Code completion and fill-in-the-middle | Base coding model with about 1.54B parameters; for FIM and code continuation, not instruction chat. | Use this model's completion / FIM format; do not substitute an instruction model's chat template. | Publisher-documented use | ||||
| Code completion / FIM | Code completion with sliding-window attention | Base coding model; 16,384-token input with a 4,096-token attention window. For completion, not a chat assistant. | Use this model's completion / FIM format; do not substitute an instruction model's chat template. | Publisher-documented use | ||||
| Programming | Coding agent | MoE model with 80B total and 3B active parameters, hybrid attention and direct answers; weight memory comes from the full model. | — | Publisher-documented use | ||||
| Document retrieval | Multilingual retrieval with task instructions | 1,024-dimensional embeddings; add a one-sentence instruction to queries and pass documents without it. | Query format: Instruct: … Query: …; documents without instructions. Masked mean pooling and L2 normalization; maximum 512 tokens. | Publisher-documented use | ||||
| Text generation | Small direct-answer assistant | Instruction-tuned 4B variant with a 262,144-token text limit; this checkpoint has no thinking mode. | — | Publisher-documented use | ||||
| Document retrieval | Persian text embeddings | Small Tooka-SBERT-V2 variant with 768-dimensional vectors; a Persian-focused candidate for retrieval and text similarity. | The retrieved metadata does not specify a usage license. | Publisher-documented use | ||||
| Document retrieval | Persian text embeddings | Large Tooka-SBERT-V2 variant with 1,024-dimensional vectors; PTEB results cannot be ranked directly against other benchmarks. | The retrieved metadata does not specify a usage license. | Publisher-documented use | ||||
| Structured extraction | Training base for Persian tasks | Base ParsBERT for Persian text understanding; classification and NER require task heads and training. Not a ready-made retrieval embedding model. | Base weights alone are not a ready-to-use classifier or NER model., The retrieved metadata does not specify a usage license. | Publisher-documented use | ||||
| Text generation | Iberian-language text generation | An Iberian-language reference candidate with published Spanish task results. The 2B instruction model is a language-specific comparison point; these results do not establish a universal winner. | — | Publisher-documented use | ||||
| Text generation | Iberian-language text generation | An Iberian-language reference candidate with published Spanish task results. The 7B instruction model is a language-specific comparison point; these results do not establish a universal winner. | — | Publisher-documented use | ||||
| Reasoning | Reasoning with a compact model | A compact candidate for reasoning and tool workflows, with about 2.52 billion parameters reported in repository metadata. Separate the developer’s evaluations from relayed Artificial Analysis scores. This package establishes no Spanish or Persian task quality. | — | Publisher-documented use | ||||
| Structured extraction | On-device information extraction | A compact candidate for extraction and bounded on-device tasks. The publisher advises against programming and knowledge-intensive use. Spanish is declared, but this package has no task-level Spanish quality score. It uses a custom license. | — | Publisher-documented use | ||||
| Reasoning | Reasoning with a compact model | A Granite reasoning model with thinking and non-thinking modes. The 3B label is a size name; repository metadata reports about 3.66 billion parameters. Keep the native 128K context separate from the claimed extension to 512K. | — | Publisher-documented use | ||||
| Document retrieval | Multimodal retrieval | Text, image and video retrieval with 64–2048 dimensional vectors and a task context of 32K tokens. | — | Guide's analytical recommendation | ||||
| Document retrieval | Multimodal retrieval | Text, image and video retrieval with 64–4096 dimensional vectors and a task context of 32K tokens. | — | Guide's analytical recommendation | ||||
| Document reranking | Multimodal reranking | Scores query relevance for text, images and video; reranks retrieved candidates with up to 32K task tokens. | — | Guide's analytical recommendation | ||||
| Document reranking | Multimodal reranking | Scores query relevance for text, images and video; reranks retrieved candidates with up to 32K task tokens. | — | Guide's analytical recommendation | ||||
| Text generation, Image understanding, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Text and vision model with hybrid attention | Qwen3.6-27B accepts text, images and video. Official BF16 weights are recorded; hybrid-attention runtime memory needs engine-specific measurement. | — | Guide's analytical recommendation | ||||
| Text generation, Image understanding, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling | Text and vision model with hybrid attention | Qwen3.6-35B-A3B accepts text, images and video. Official BF16 weights are recorded; hybrid-attention runtime memory needs engine-specific measurement. | — | Guide's analytical recommendation | ||||
| Text generation | Dedicated text translation | Dedicated translation model with terminology and contextual prompts; not selected for general chat or document Q&A. | — | Guide's analytical recommendation | ||||
| Text generation | Dedicated text translation | Dedicated translation model with terminology and contextual prompts; not selected for general chat or document Q&A. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Second-generation Hy-MT translator with translation instructions and terminology control. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Second-generation Hy-MT translator with translation instructions and terminology control. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Second-generation Hy-MT translator with translation instructions and terminology control. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Text and image-text translator with a documented 2K input limit and a translation-specific template. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Text and image-text translator with a documented 2K input limit and a translation-specific template. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Text and image-text translator with a documented 2K input limit and a translation-specific template. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | T5 translator with broad language coverage; the target-language prefix is part of its input. | — | Guide's analytical recommendation | ||||
| Text generation | Specialized translation | Research translation baseline, noncommercial license and training inputs up to 512 tokens. | — | Guide's analytical recommendation | ||||
| Programming | Coding, software agents, text and image analysis | Xiaomi’s multimodal model for coding, tool use and long inputs; the RL checkpoint has downloadable weights with MIT license metadata. | About 1.02T total and 42B active parameters; weight files occupy 534.1 GiB. This checkpoint does not fit on one 24 or 48 GB GPU., Artificial Analysis Intelligence Index v4.3.2: 46 for the MiMo-V2.6-Pro service; this is neither a Persian evaluation nor a local RL-checkpoint test. | Guide's analytical recommendation |
“—” means no information is recorded, not that the model lacks the capability.
Hardware feasibility
Memory for weights, KV and runtime under the selected settings.
Calculated from specifications. Context includes history, input and output. Active requests differ from daily users.
GPU configurations · 3 Selected configuration
GPU specifications in the GPU guide →Calculation method and scenario limits
Actual weight file size + KV cache + runtime reserve. GGUF scenarios use FP16 KV cache. The default GPU reserve is 2 GiB per card; the CPU reserve is 4 GiB. These reserves are planning assumptions.
MoE calculations include all weights. Nominal GPU capacity is used as a GiB budget; deployment planning should use the device's reported free memory. The sum of several GPUs' memory is not unified memory.
SmolLM2 FP32 scenarios use FP32 for both weights and KV cache. Scenarios above the selected file's context limit are excluded, including Aya Expanse 32B with the selected GGUF file's 8,192-token limit.
Advanced filters 1 controls
Row comparison and conditions
Comparison rule: Hardware cells share a row's scenario. Comparing rows requires matching non-hardware conditions.
| Compare | Details | Model | Weight version | RTX 4090 24GB | RTX 5090 32GB | RTX 6000 Ada 48GB | |||
|---|---|---|---|---|---|---|---|---|---|
| Q8_0 | 3.47 GiB | 0.6 GiB | 0.88 GiB | ||||||
| Q8_0 | 4.58 GiB | 1.71 GiB | 0.88 GiB | ||||||
| Q4_K_M | 5.45 GiB | 2.33 GiB | 1.13 GiB | ||||||
| Q8_0 | 7.11 GiB | 3.99 GiB | 1.13 GiB | ||||||
| Q4_K_M | 7.81 GiB | 4.68 GiB | 1.13 GiB | ||||||
| Q8_0 | 11.24 GiB | 8.11 GiB | 1.13 GiB | ||||||
| Q4_K_M | 11.63 GiB | 8.38 GiB | 1.25 GiB | ||||||
| Q8_0 | 17.87 GiB | 14.62 GiB | 1.25 GiB | ||||||
| Q4_K_M | 20.03 GiB | 17.28 GiB | 0.75 GiB | ||||||
| Q8_0 | 33 GiB | 30.25 GiB | 0.75 GiB | ||||||
| Q4_K_M | 22.4 GiB | 18.4 GiB | 2 GiB | ||||||
| Q8_0 | 36.43 GiB | 32.43 GiB | 2 GiB | ||||||
| Q4_K_M | 44.1 GiB | 39.6 GiB | 2.5 GiB | ||||||
| Q8_0 | 74.33 GiB | 69.83 GiB | 2.5 GiB | ||||||
| Q4_K_M | 11.87 GiB | 8.37 GiB | 1.5 GiB | ||||||
| Q8_0 | 18.12 GiB | 14.62 GiB | 1.5 GiB | ||||||
| Q4_K_M | 22.49 GiB | 18.49 GiB | 2 GiB | ||||||
| Q8_0 | 36.43 GiB | 32.43 GiB | 2 GiB | ||||||
| Q4_K_M | 6.8 GiB | 4.36 GiB | 0.44 GiB | ||||||
| Q8_0 | 9.98 GiB | 7.54 GiB | 0.44 GiB | ||||||
| Q4_K_M | 44.1 GiB | 39.6 GiB | 2.5 GiB | ||||||
| Q8_0 | 74.33 GiB | 69.83 GiB | 2.5 GiB | ||||||
| Q4_K_M | 7.58 GiB | 4.58 GiB | 1 GiB | ||||||
| Q8_0 | 10.95 GiB | 7.95 GiB | 1 GiB | ||||||
| Q4_K_M | 21.69 GiB | 18.44 GiB | 1.25 GiB | ||||||
| Q8_0 | 35.22 GiB | 31.97 GiB | 1.25 GiB | ||||||
| Q4_K_M | 7.71 GiB | 4.71 GiB | 1 GiB | ||||||
| Q8_0 | 10.95 GiB | 7.95 GiB | 1 GiB | ||||||
| Q4_K_M | 4.48 GiB | 0.98 GiB | 1.5 GiB | ||||||
| Q4_K_M | 2.27 GiB | 0.1 GiB | 0.18 GiB | ||||||
| Q8_0 | 2.31 GiB | 0.13 GiB | 0.18 GiB | ||||||
| Q8_0 | 2.67 GiB | 0.36 GiB | 0.31 GiB | ||||||
| Q4_K_M | 20.03 GiB | 17.28 GiB | 0.75 GiB | ||||||
| Q8_0 | 33 GiB | 30.25 GiB | 0.75 GiB | ||||||
| Q4_K_M | 7.81 GiB | 4.68 GiB | 1.13 GiB | ||||||
| Q8_0 | 11.24 GiB | 8.11 GiB | 1.13 GiB | ||||||
| Q4_K_M | 3.26 GiB | 1.04 GiB | 0.22 GiB | ||||||
| Q8_0 | 3.98 GiB | 1.76 GiB | 0.22 GiB | ||||||
| Q4_K_M | 4.06 GiB | 1.44 GiB | 0.63 GiB | ||||||
| Q8_0 | 5.13 GiB | 2.51 GiB | 0.63 GiB | ||||||
| Q4_K_M | 5.32 GiB | 2.32 GiB | 1 GiB | ||||||
| Q8_0 | 6.8 GiB | 3.8 GiB | 1 GiB | ||||||
| Q4_K_M | 7.07 GiB | 4.07 GiB | 1 GiB | ||||||
| Q8_0 | 10.17 GiB | 7.17 GiB | 1 GiB | ||||||
| Q4_K_M | 136.32 GiB | 132.85 GiB | 1.47 GiB | ||||||
| Q8_0 | 236.24 GiB | 232.77 GiB | 1.47 GiB | ||||||
| Q4_K_M | 274.8 GiB | 270.86 GiB | 1.94 GiB | ||||||
| Q8_0 | 479.24 GiB | 475.3 GiB | 1.94 GiB | ||||||
| Q4_K_M | 11.99 GiB | 8.43 GiB | 1.56 GiB | ||||||
| Q8_0 | 18.07 GiB | 14.51 GiB | 1.56 GiB | ||||||
| Q4_K_M | 6.8 GiB | 4.36 GiB | 0.44 GiB | ||||||
| Q8_0 | 9.98 GiB | 7.54 GiB | 0.44 GiB | ||||||
| Q4_K_M | 11.87 GiB | 8.37 GiB | 1.5 GiB | ||||||
| Q8_0 | 18.12 GiB | 14.62 GiB | 1.5 GiB | ||||||
| Q4_K_M | 22.49 GiB | 18.49 GiB | 2 GiB | ||||||
| Q8_0 | 36.43 GiB | 32.43 GiB | 2 GiB |
“—” means no information is recorded, not that the model lacks the capability.
Inference and serving software
Software comparison
Compare software for local inference and model serving.
Advanced filters 36 controls
Row comparison and conditions
Comparison rule: Versions can be displayed together. Interface comparisons keep the backend fixed; full-stack comparisons may vary it.
| Compare | Details | What is it for? | Practical advantage | Selection condition | Get started | ||
|---|---|---|---|---|---|---|---|
| Inference engine / library, API server, Model manager | Local setup, development and small services | Model downloads, management and API setup in one tool | Local and cloud execution are separate; OLLAMA_NO_CLOUD=1 disables cloud paths. | Getting started ↗ | |||
| Inference engine / library, API server | Language model API with concurrent requests | Continuous batching and KV-cache management; OpenAI-compatible API | GGUF requires the separately installed vllm-gguf-plugin; support for this format is experimental. | Getting started ↗ | |||
| Inference engine / library, API server | Language and multimodal model serving on one GPU or a cluster | Text generation, Continuous batching, Prefix caching, Structured output Tool parsers and quantization kernels depend on the model architecture. Starting with 0.5.20, CUDA 12 packages and images are no longer published; 0.5.19 is the last release for this path. Previously published images remain available. Responses storage is off by default. Retrieval, previous_response_id and background requests require --enable-response-store; this option cannot be enabled in prefill/decode-disaggregated (PD) deployments. | Tool parsers and quantization kernels depend on the model architecture. Starting with 0.5.20, CUDA 12 packages and images are no longer published; 0.5.19 is the last release for this path. Previously published images remain available. Responses storage is off by default. Retrieval, previous_response_id and background requests require --enable-response-store; this option cannot be enabled in prefill/decode-disaggregated (PD) deployments. | Getting started ↗ | |||
| Inference engine / library, API server | Multi-user serving and workloads with shared prefixes | Batching and prefix caching; workload-specific tuning | Tool parsers and quantization kernels depend on model architecture; read each benchmark row's version and settings with its result. | Getting started ↗ | |||
| Inference engine / library, API server | GGUF on CPU, GPU and combined RAM/VRAM | Control over quantization, layer partitioning and CPU execution | Partial CPU execution is distinct from AirLLM's layer-wise loading. | Getting started ↗ | |||
| Inference engine / library, API server | Specialized deployment on NVIDIA GPUs | Model, kernel and serving optimization | Support paths and settings vary by release and architecture. | Getting started ↗ | |||
| API server, Deployment manager | Manage model services and inference pipelines | Model management, scheduling and specialist backend integration | Release 2.72.0 does not include a TensorRT-LLM backend container. | Getting started ↗ | |||
| Inference engine / library, API server | Document and query embedding service | Batching and an embedding API; CPU and GPU images | Embedding support does not imply reranking support for every architecture. | Getting started ↗ | |||
| Inference engine / library | Run large models with limited VRAM | Load and unload layers through host memory and storage | The main goal is lower GPU memory use; concurrent chat requires a benchmark of that workload. | Getting started ↗ | |||
| Inference engine / library, API server | Research, direct model control and custom paths | Direct access to model and processing logic | A simple Transformers path's speed does not represent every engine for that model. | Getting started ↗ | |||
| Model gateway, API server | Manage multiple providers and backends | Shared API, routing and usage management | Weight execution capability comes from the backend. | Getting started ↗ | |||
| Chat interface | Chat panel connected to model backends | User chat interface and runtime connection | Model memory and speed belong to the connected backend. | Getting started ↗ | |||
| Inference engine / library, API server | Maintenance of text-generation serving for supported models | Text generation, Model sharding across GPUs (conditional), Streaming output, Continuous batching The repository has been archived and read-only since March 21, 2026. | The repository has been archived and read-only since March 21, 2026. | Getting started ↗ | |||
| Chat interface, Model manager, API server | Download models, chat and start a local API | Model loading and unloading, Streaming output, Structured output, Concurrent requests Model capabilities and formats depend on the selected runtime. | Model capabilities and formats depend on the selected runtime. | Getting started ↗ | |||
| Inference engine / library | When host memory and GPU execution are part of the deployment design. | Heterogeneous CPU/GPU execution, including optimized paths for supported MoE models. | Performance depends on the supported model, CPU, memory and configuration. | Getting started ↗ | |||
| Inference engine / library | For implementing and evaluating the retrieval layer. | Embeddings for semantic retrieval and CrossEncoder scoring or reranking. | Library support does not establish task or language quality. | Getting started ↗ | |||
| Inference engine / library | BGE family and reranker execution | BGE-M3 dense, sparse and multi-vector outputs | The three BGE-M3 outputs use different indexing and consumption paths; dense output alone cannot replace all three. | Getting started ↗ | |||
MLX LM · v0.31.3 | Inference engine / library | Local generation and quantization of compatible models on Apple silicon | Text generation (conditional), Streaming output (conditional), Prefix caching (conditional), Model sharding across GPUs (conditional) A direct route for local generation and quantization on Apple silicon with compatible models. Do not double-count unified memory as RAM plus VRAM. The README’s macOS 15 requirement concerns large-model memory wiring, not every MLX LM feature. | A direct route for local generation and quantization on Apple silicon with compatible models. Do not double-count unified memory as RAM plus VRAM. The README’s macOS 15 requirement concerns large-model memory wiring, not every MLX LM feature. | Getting started ↗ |
“—” means no information is recorded, not that the model lacks the capability.
Evaluation results
Evidence for model selection
How selection and comparison work
On Spanish MIRACL, nDCG@10 rises from 51.2 for Small to 53.7 for Large-Instruct. On English, Large scores 52.9 versus 51.5 for Large-Instruct; Persian scores are 59.0 and 59.4. A multilingual average cannot replace the target-language comparison.
Two displayed numbers are not automatically rankable. Check the model variant, benchmark version, language, generation mode and protocol. A reported score difference is neither a quality ratio nor a statistical-significance claim.
MIRACL · nDCG@10· English
| Model | Score | Difference from first row | Source and settings |
|---|---|---|---|
| 48 points / 100 | — | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=1
| |
| 51.2 points / 100 | 3.2 | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=2
| |
| 52.9 points / 100 | 4.9 | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=3
| |
| 51.5 points / 100 | 3.5 | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=4
|
MIRACL · Recall@100· English
| Model | Score | Difference from first row | Source and settings |
|---|---|---|---|
| 85.3 points / 100 | — | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=5
| |
| 86.4 points / 100 | 1.1 | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=6
| |
| 87.6 points / 100 | 2.3 | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=7
| |
| 88.2 points / 100 | 2.9 | Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher reportAppendix / detailed MIRACL results / language=en / column=8
|
Quality in published evaluations
MATH-500 · pass@1 · English · %
| Select for comparison | Model | Score and scale | Mode / subset | Reporter | Source |
|---|---|---|---|---|---|
| 94.5 %0..100 as printed | Reasoningtest | DeepSeekPublisher report | Conditions and sourceMATH-500 · pass@1: 94.5%Publisher report · DeepSeek· Reasoning· en
DeepSeek-R1-Distill-Llama-70B — publisher evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B/blob/b1c0b44b4369b597ad119a196caf79a9c40e141e/README.md ↗
MATH-500 dataset metadata · Technical documentationhttps://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
DeepSeek-R1: distilled model evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
| ||
| 83.9 %0..100 as printed | Reasoningtest | DeepSeekPublisher report | Conditions and sourceMATH-500 · pass@1: 83.9%Publisher report · DeepSeek· Reasoning· en
DeepSeek-R1-Distill-Qwen-1.5B — publisher evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B/blob/ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562/README.md ↗
MATH-500 dataset metadata · Technical documentationhttps://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
DeepSeek-R1: distilled model evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
| ||
| 93.9 %0..100 as printed | Reasoningtest | DeepSeekPublisher report | Conditions and sourceMATH-500 · pass@1: 93.9%Publisher report · DeepSeek· Reasoning· en
DeepSeek-R1-Distill-Qwen-14B — publisher evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B/blob/1df8507178afcc1bef68cd8c393f61a886323761/README.md ↗
MATH-500 dataset metadata · Technical documentationhttps://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
DeepSeek-R1: distilled model evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
| ||
| 94.3 %0..100 as printed | Reasoningtest | DeepSeekPublisher report | Conditions and sourceMATH-500 · pass@1: 94.3%Publisher report · DeepSeek· Reasoning· en
DeepSeek-R1-Distill-Qwen-32B — publisher evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/blob/711ad2ea6aa40cfca18895e8aca02ab92df1a746/README.md ↗
MATH-500 dataset metadata · Technical documentationhttps://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
DeepSeek-R1: distilled model evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
| ||
| 92.8 %0..100 as printed | Reasoningtest | DeepSeekPublisher report | Conditions and sourceMATH-500 · pass@1: 92.8%Publisher report · DeepSeek· Reasoning· en
DeepSeek-R1-Distill-Qwen-7B — publisher evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B/blob/916b56a44061fd5cd7d6a8fb632557ed4f724f60/README.md ↗
MATH-500 dataset metadata · Technical documentationhttps://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
DeepSeek-R1: distilled model evaluation · Publisher reporthttps://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
|
Small and specialized complementary models
Compare specialized roles, input limits, vector dimensions and results for the selected language.
Test shown in the result column: MIRACL · nDCG@10 · en
This selection applies to result columns across tables. Execution conditions are available in each model’s details.
Results grouped by test ←Advanced filters 6 controls
Row comparison and conditions
Comparison rule: Metric and unit must describe the same task; document/s is never implicitly converted to token/s.
| Compare | Details | Exact task | Output type / dimensions | Languages and selected-language evidence | Licence | Result for the selected language | Download and setup | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Dense, sparse and multi-vector retrieval | approximately 0.569 billion | 8,192 tokens | 1,024-dimensional dense vector + token weights + multi-vector output | multilingual en: No recorded evidence | MIT ↗ | — | 3.81 GiB / one million vectors | ||||
| Query–document pair reranking | approximately 0.568 billion | 8,192 tokens | Relevance score; optional sigmoid maps to 0–1 | multilingual en: Published result | Apache-2.0 ↗ | — | Does not store fixed vectors | ||||
| Semantic retrieval and text embeddings | approximately 0.6 billion | 32,768 tokens | Dense vector; 1,024 default / maximum dimensions | multilingual en: Published result | Apache-2.0 ↗ | — | 3.81 GiB / one million vectors | ||||
| Semantic retrieval and text embeddings | approximately 4 billion | 32,768 tokens | Dense vector; 2,560 default / maximum dimensions | multilingual en: No recorded evidence | Apache-2.0 ↗ | — | 9.54 GiB / one million vectors | ||||
| Semantic retrieval and text embeddings | approximately 8 billion | 32,768 tokens | Dense vector; 4,096 default / maximum dimensions | multilingual en: No recorded evidence | Apache-2.0 ↗ | — | 15.26 GiB / one million vectors | ||||
| Query–document pair reranking | approximately 0.6 billion | 32,768 tokens | Text-pair relevance score; no embedding output | multilingual en: Published result | Apache-2.0 ↗ | — | Does not store fixed vectors | ||||
| Query–document pair reranking | approximately 4 billion | 32,768 tokens | Text-pair relevance score; no embedding output | multilingual en: Published result | Apache-2.0 ↗ | — | Does not store fixed vectors | ||||
| Query–document pair reranking | approximately 8 billion | 32,768 tokens | Text-pair relevance score; no embedding output | multilingual en: Published result | Apache-2.0 ↗ | — | Does not store fixed vectors | ||||
| Text retrieval and similarity | 0.118 billion | 512 tokens | 384-dimensional vector | multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result | MIT ↗ | 1.43 GiB / one million vectors | |||||
| Text retrieval and similarity | 0.023 billion | 256 tokens | 384-dimensional vector | en en: Publisher declaration | Apache-2.0 ↗ | — | 1.43 GiB / one million vectors | ||||
| Text retrieval and similarity | 0.278 billion | 512 tokens | 768-dimensional vector | multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result | MIT ↗ | 2.86 GiB / one million vectors | |||||
| Text retrieval and similarity | 0.56 billion | 512 tokens | 1,024-dimensional vector | multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result | MIT ↗ | 3.81 GiB / one million vectors | |||||
| Text retrieval and similarity | 0.033 billion | 512 tokens | 384-dimensional vector | en en: Publisher declaration | MIT ↗ | — | 1.43 GiB / one million vectors | ||||
| Search result reranking | 0.278 billion | 512 tokens | Query–document relevance score | en, zh en: Publisher declaration | MIT ↗ | — | Does not store fixed vectors | ||||
| Text retrieval and similarity | approximately 0.3 billion nominal Total parameter count unverified | 2,048 tokens | 768-dimensional vector | multilingual en: No recorded evidence | gemma ↗ Commercial use subject to licence | — | 2.86 GiB / one million vectors | ||||
| Text retrieval and similarity | approximately 0.57 billion nominal Total parameter count unverified | 8,192 tokens | 1,024-dimensional vector | multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result | CC-BY-NC-4.0 ↗ Commercial use subject to licence | — | 3.81 GiB / one million vectors | ||||
| Search result reranking | 0.278 billion | 1,024 tokens | Query–document relevance score | multilingual en: No recorded evidence | CC-BY-NC-4.0 ↗ Commercial use subject to licence | — | Does not store fixed vectors | ||||
| Text retrieval and similarity | 0.335 billion | 512 tokens | 1,024-dimensional vector | en en: Publisher declaration | Apache-2.0 ↗ | — | 3.81 GiB / one million vectors | ||||
| Text retrieval and similarity | 0.137 billion | 8,192 tokens | 768-dimensional vector | en en: Publisher declaration | Apache-2.0 ↗ | — | 2.86 GiB / one million vectors | ||||
| Specialization for classification and named entity recognition | 0.15 billion | 8,192 tokens | Token representations; task head after training | en en: Publisher declaration | Apache-2.0 ↗ | — | Does not store fixed vectors | ||||
| Text retrieval and similarity | 0.56 billion | 512 tokens | Dense vector | af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result | mit ↗ | 3.81 GiB / one million vectors | |||||
| Text retrieval and similarity | 0.123 billion | 512 tokens | Dense vector | fa en: No recorded evidence | — | — | 2.86 GiB / one million vectors | ||||
| Text retrieval and similarity | 0.353 billion | 512 tokens | Dense vector | fa en: No recorded evidence | — | — | 3.81 GiB / one million vectors | ||||
| Base for training classification and entity recognition | — | 512 tokens | Token representations; the task head requires training | fa en: No recorded evidence | — | — | Does not store fixed vectors | ||||
| Text, image and video retrieval | 2 billion | 32,768 tokens | Vector with up to 2048 dimensions | — en: No recorded evidence | Apache-2.0 ↗ | — | 7.63 GiB / one million vectors | ||||
| Text, image and video retrieval | 8 billion | 32,768 tokens | Vector with up to 4096 dimensions | — en: No recorded evidence | Apache-2.0 ↗ | — | 15.26 GiB / one million vectors | ||||
| Reranker | 2 billion | 32,768 tokens | — | — en: No recorded evidence | Apache-2.0 ↗ | — | Does not store fixed vectors | ||||
| Reranker | 8 billion | 32,768 tokens | — | — en: No recorded evidence | Apache-2.0 ↗ | — | Does not store fixed vectors |
“—” means no information is recorded, not that the model lacks the capability.
Guide articles
Model selection, memory, runtime software and quality evaluation.
RAG, CAG, KAG, fine-tuning and instruction tuning: how do they differ, and which should you choose?
When a model gives an unsatisfactory answer, does it need more training, or simply access to the right information? Through practical analogies, this article compares RAG, CAG and KAG with fine-tuning and instruction tuning: retrieving documents, caching knowledge, reasoning over relationships and changing model behavior. A comparison table and real project scenarios help distinguish missing knowledge from unsuitable behavior before committing to training, and identify the method or combination that fits the task.
RAG, CAG, KAG, fine-tuning and instruction tuning: how do they differ, and which should you choose?
When a model gives an unsatisfactory answer, does it need more training, or simply access to the right information? Through practical analogies, this article compares RAG, CAG and KAG with fine-tuning and instruction tuning: retrieving documents, caching knowledge, reasoning over relationships and changing model behavior. A comparison table and real project scenarios help distinguish missing knowledge from unsuitable behavior before committing to training, and identify the method or combination that fits the task.
INT8 or FP8: what your GPU can actually run
An eight-bit model format does not define its execution path. Examine kernels, memory and output quality before choosing hardware for language-model inference.
INT8 or FP8: what your GPU can actually run
An eight-bit model format does not define its execution path. Examine kernels, memory and output quality before choosing hardware for language-model inference.
Ollama, vLLM, SGLang or llama.cpp: choosing an inference engine
How scheduling, KV cache, model formats and operational controls change the choice of inference engine, with a worked memory example and a reproducible comparison method.
Ollama, vLLM, SGLang or llama.cpp: choosing an inference engine
How scheduling, KV cache, model formats and operational controls change the choice of inference engine, with a worked memory example and a reproducible comparison method.
Why the fastest GPU does not necessarily deliver the fastest response
More compute and a higher token rate do not always mean a faster response. This article examines time to first token versus completion time, memory and concurrency, the division of work between GPUs and LPUs, and communication costs—so infrastructure choices reflect the capacity to serve requests at the required quality and latency.
Why the fastest GPU does not necessarily deliver the fastest response
More compute and a higher token rate do not always mean a faster response. This article examines time to first token versus completion time, memory and concurrency, the division of work between GPUs and LPUs, and communication costs—so infrastructure choices reflect the capacity to serve requests at the required quality and latency.