LLM & SLM

Choosing a language model: large or small?

Compare models, inference software and memory requirements for your workload, with published results and practical guides.

Running a language model on your own computer or server can support chat, coding, translation or answers grounded in documents. Data confidentiality, control over service dependencies and predictable costs are common reasons to consider local deployment. The range of models and execution methods makes the choice less straightforward: which model is sufficient for the task, what fits the available hardware, and when does upgrading the infrastructure make a useful difference?

This guide brings those decisions together. Its interactive tables compare small and large language models, related specialist models, hardware requirements and runtime software, with links to the supporting sources. The accompanying articles explain the practical advantages and limits of each approach, connecting the task and required quality to the resources and budget needed to deliver it.

Language models of different sizes connected to chat, coding and document applications
Using this collection

Where should I start?

Open a table to compare options; read the articles in each path to understand the concepts and choices.

Interactive guide

Start with the key questions; add deployment details later to refine the plan.

Which task do you want to improve?

The task determines which model capabilities to test.

Organizational model guide

Choose one task for a first trial

Choose a letter summary, invoice amount extraction or policy lookup. Define a correct output with the person doing the work, then return to the first question.

Your supplied information

Read next

For the technical team

Test the same candidates above with recorded versions and settings.

Weight, runtime software and provider terms are separate. Check the linked weight license for this use; runtime and hosted-service conditions apply independently.

All answers and their provenance

2026-09-20.2 · 2026-09-20 — Qwen3.6 / Qwen3.7

For the contractor

Quote a pilot using these same models, budgets and criteria. Separate setup, operations and support, and specify maintenance ownership and load-test acceptance.

125 Model 17 Software 7 interactive tables 4 guide articles Data last updated:

Once you know the model and its memory requirements, compare suitable GPUs and servers.

Sources and coverage

Specifications come from model cards, configurations and versioned software documentation. Quality and speed results retain their reporter and source conditions; this site has not run independent model benchmarks.

Candidates cover distinct roles: text generation, coding, images and documents, embedding and reranking. Model size and a multilingual label do not replace evidence for the selected task and language.

The hardware table and most current speed reports focus on NVIDIA. CPU, Apple Silicon and AMD execution depends on the software and backend; the guide lacks matched speed results for ranking all these platforms. Software routes and supported hardware

Memory estimates combine weight-file bytes, KV and a reserve for each device. The conventional KV formula is not applied to hybrid, MLA or unknown architectures. Assumptions are available in the memory section.

Data reviewed: · Report a correction through the résumé contact links

Data view

Model catalog

Downloadable-weight models: size, architecture, context and licence. API-only services are not listed here.

Test shown in the result column: MIRACL · nDCG@10 · en

This selection applies to result columns across tables. Execution conditions are available in each model’s details.

Results grouped by test ←
125 results out of 125 rows
Downloadable package
Documented setup path
Family
Model type
Model stage
Declared model size(billion)
Advanced filters 17 controls
Artifact publisher
Active parameters(billion)
Parameter architecture
Input
Output
Application
Maximum context length(token)
Release status
Release date
Review date
All rows
Row comparison and conditions

Comparison rule: Side-by-side display is available; calculations require compatible units and version-specific attribution.

Use × beside a heading to hide its column. Open row details for sources and conditions.

Hidden columns:
Models available through APIs · 4

Self-hostable weights were not verified in this review; these releases are excluded from local deployment suggestions.

Language model catalog
CompareDetails
Input → output
Main uses
Licence
Download and run
Official page ↗
BGE · BAAI approximately 0.569 billion · Dense Text ← Vector 8,192 tokens Hybrid document retrieval in RAG MIT
Official page ↗
BGE · BAAI approximately 0.568 billion · Dense Text ← Structured data 8,192 tokens Reorder retrieved documents Apache-2.0
Official page ↗
Aya / Cohere · Cohere Labs approximately 32 billion · Dense Text ← Text 131,072 tokens Multilingual writing and chat with long documents CC-BY-NC-4.0 + Cohere Acceptable Use Policy Commercial use subject to licence
Official page ↗
Aya / Cohere · Cohere Labs approximately 8 billion · Dense Text ← Text 8,192 tokens Multilingual writing and rewriting assistant CC-BY-NC-4.0 + Cohere Acceptable Use Policy Commercial use subject to licence
Official page ↗
Aya / Cohere · Cohere Labs approximately 3.35 billion · Dense Text ← Text 8,192 tokens Local chat across languages CC-BY-NC-4.0 + Cohere Acceptable Use Policy Commercial use subject to licence
Official page ↗
SmolLM · Hugging Face approximately 1.7 billion · Dense Text ← Text 8,192 tokens Prototype a small assistant with function calling Apache-2.0
Official page ↗
SmolLM · Hugging Face approximately 0.135 billion · Dense Text ← Text 8,192 tokens Explore instruction following with a very small model Apache-2.0
Official page ↗
SmolLM · Hugging Face approximately 0.36 billion · Dense Text ← Text 8,192 tokens Short-text rewriting and summarization Apache-2.0
Official page ↗
SmolLM · Hugging Face approximately 3 billion · Dense Text ← Text 65,536 tokens Up to 131,072 tokens with Configure YaRN and increase max_position_embeddings Small assistant with selectable thinking mode Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 0.6 billion · Dense Text ← Text 32,768 tokens Prototype chat with the smallest Qwen3 Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 1.7 billion · Dense Text ← Text 32,768 tokens Small text assistant with reasoning control Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 14.8 billion · Dense Text ← Text 32,768 tokens Up to 131,072 tokens with YaRN configuration Text generation and analysis with dense Qwen3 Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 30.5 billion · 3.3 billion active · Mixture of experts (MoE) Text ← Text 32,768 tokens Up to 131,072 tokens with YaRN configuration Tool-oriented assistant with MoE architecture Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 32.8 billion · Dense Text ← Text 32,768 tokens Up to 131,072 tokens with YaRN configuration Multi-step analysis with a larger dense Qwen3 Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 4 billion · Dense Text ← Text 32,768 tokens Up to 131,072 tokens with YaRN configuration General assistant at four billion parameters Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 8.2 billion · Dense Text ← Text 32,768 tokens Up to 131,072 tokens with YaRN configuration Chat and text tasks with thinking control Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 30.5 billion · 3.3 billion active · Mixture of experts (MoE) Text ← Text 262,144 tokens Up to 1,048,576 tokens with YaRN configuration Repository editing with a tool-using assistant Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 0.6 billion · Dense Text ← Vector 32,768 tokens Vector retrieval with the smallest Qwen3 Embedding Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 4 billion · Dense Text ← Vector 32,768 tokens Multilingual indexing with adjustable vectors Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 8 billion · Dense Text ← Vector 32,768 tokens Document retrieval with the largest listed Qwen3 Embedding Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 0.6 billion · Dense Text ← Structured data 32,768 tokens Lighter reranking in the Qwen3 family Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 4 billion · Dense Text ← Structured data 32,768 tokens Instruction-aware reranking of candidate documents Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 8 billion · Dense Text ← Structured data 32,768 tokens Reranking with the 8B Qwen3 variant Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 8 billion nominal · Dense Total parameter count unverified Text, Image, Video ← Text 262,144 tokens Read images and documents with Qwen3-VL Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 2 billion · approximately 2 billion Language component · Dense · hybrid attention Text, Image, Video ← Text 262,144 tokens Multimodal prototyping with small Qwen3.5 Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 35 billion · 3 billion active · approximately 35 billion Language component · Mixture of experts (MoE) · hybrid attention Text, Image, Video ← Text 262,144 tokens Up to 1,010,000 tokens with Context extension with YaRN Multimodal agent with MoE Qwen3.5 Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 4 billion · approximately 4 billion Language component · Dense · hybrid attention Text, Image, Video ← Text 262,144 tokens Up to 1,010,000 tokens with Context extension with YaRN Text and image processing at four billion parameters Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 9 billion · approximately 9 billion Language component · Dense · hybrid attention Text, Image, Video ← Text 262,144 tokens Up to 1,010,000 tokens with Context extension with YaRN Document and image assistant with dense Qwen3.5 Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 27 billion · approximately 27 billion Language component · Hybrid · hybrid attention Text, Image, Video ← Text 262,144 tokens Up to 1,000,000 tokens with Context extension with YaRN Long workflows with a text–image agent Apache-2.0
Official page ↗
OLMo · Allen Institute for AI approximately 7 billion · Dense Text ← Text 65,536 tokens Instruction-following research with inspectable training Apache-2.0
Official page ↗
DeepSeek · DeepSeek approximately 8 billion · Dense Text ← Text 131,072 tokens Explore reasoning distilled from R1-0528 MIT
Official page ↗
DeepSeek · DeepSeek approximately 70 billion · Dense Text ← Text 131,072 tokens Distilled reasoning on a Llama 70B base MIT + underlying Llama 3.3 terms
Official page ↗
DeepSeek · DeepSeek approximately 1.5 billion · Dense Text ← Text 131,072 tokens Explore reasoning limits in a very small distilled model MIT
Official page ↗
DeepSeek · DeepSeek approximately 14 billion · Dense Text ← Text 131,072 tokens Problem solving with a mid-sized R1 distillation MIT
Official page ↗
DeepSeek · DeepSeek approximately 32 billion · Dense Text ← Text 131,072 tokens Analysis and coding with a dense 32B distillation MIT
Official page ↗
DeepSeek · DeepSeek approximately 7 billion · Dense Text ← Text 131,072 tokens Problem solving with the 7B R1 distillation MIT
Official page ↗
DeepSeek · DeepSeek 685.397 billion stored elements · Mixture of experts (MoE) Count stored in the checkpoint, including additional components; not active parameters per generated token. Text ← Text 163,840 tokens Connect reasoning and tools in agent workflows MIT
Official page ↗
DeepSeek · DeepSeek 8 billion Active during prefill · 16 billion Active during decode · approximately 552 billion Main backbone · approximately 196 billion Engram conditional memory · Mixture of experts (MoE) · hybrid attention Total parameter count unverified Text, Image ← Text 1,000,000 tokens Multimodal agent for very long inputs MIT
SAFETENSORS · Official ↗
Official page ↗
Gemma · Google DeepMind approximately 12 billion · Dense Text, Image ← Text 131,072 tokens Ask about images and text with mid-sized Gemma 3 gemma
Official page ↗
Gemma · Google DeepMind approximately 1 billion · Dense Text ← Text 32,768 tokens Small text assistant from the Gemma 3 family gemma
Official page ↗
Gemma · Google DeepMind approximately 27 billion · Dense Text, Image ← Text 131,072 tokens Text and image understanding with the largest listed Gemma 3 gemma
Official page ↗
Gemma · Google DeepMind approximately 4 billion · Dense Text, Image ← Text 131,072 tokens Getting started with images in the Gemma 3 family gemma
Official page ↗
Gemma · Google DeepMind approximately 26 billion nominal · approximately 25.2 billion Language component in the publisher's table · Mixture of experts (MoE) Total parameter count unverified Text, Image, Video ← Text 262,144 tokens Multimodal reasoning with MoE Gemma 4 Apache-2.0
Official page ↗
Gemma · Google DeepMind approximately 2.3 billion Effective count excluding the embedding table · approximately 5.1 billion Language component including embeddings · approximately 0.3 billion Audio encoder · Dense Total parameter count unverified Text, Image, Audio, Video ← Text 131,072 tokens Local text, image and audio processing with small Gemma 4 Apache-2.0
Official page ↗
Granite · IBM approximately 2 billion · Dense Text ← Text 131,072 tokens Generate answers from retrieved documents Apache-2.0
Official page ↗
E5 · intfloat / multilingual E5 authors 0.118 billion · Dense Text ← Vector 512 tokens Multilingual retrieval with low-dimensional vectors MIT
Official page ↗
Llama · Meta approximately 70 billion · Dense Text ← Text 131,072 tokens General assistant with large Llama 3.1 llama3.1
Official page ↗
Llama · Meta approximately 8 billion · Dense Text ← Text 131,072 tokens Chat and text work with Llama 3.1 8B llama3.1
Official page ↗
Llama · Meta approximately 1 billion · Dense Text ← Text 131,072 tokens Local rewriting and summarization with small Llama llama3.2
Official page ↗
Llama · Meta approximately 3 billion · Dense Text ← Text 131,072 tokens On-device text assistant with Llama 3B llama3.2
Official page ↗
Phi · Microsoft approximately 3.8 billion · Dense Text ← Text 131,072 tokens Analysis and logic under tighter resource constraints MIT
Official page ↗
Mistral · Mistral AI approximately 24 billion · Dense Text, Image ← Text 262,144 tokens Multi-file editing and repository-search agent Apache-2.0
Official page ↗
Mistral · Mistral AI approximately 3.4 billion Language component · Dense Total parameter count unverified Text, Image ← Text 262,144 tokens Published weights are FP8; execution capacity depends on precision and context memory. Multimodal assistant for edge deployment Apache-2.0
Official page ↗
Mistral · Mistral AI approximately 7 billion · Dense Text ← Text 32,768 tokens Text assistant with Mistral function-calling format Apache-2.0
Official page ↗
Mistral · Mistral AI approximately 24 billion · Dense Text, Image ← Text 131,072 tokens Multilingual text and image assistant Apache-2.0
Official page ↗
Nemotron · NVIDIA approximately 9 billion · Hybrid · hybrid attention Text ← Text 131,072 tokens Document-grounded answers with a reasoning budget nvidia-open-model-license
Official page ↗
gpt-oss · OpenAI approximately 117 billion · 5.1 billion active · Mixture of experts (MoE) Text ← Text 131,072 tokens With the YaRN settings supplied with the published model Adjustable-reasoning agent with the larger gpt-oss Apache-2.0
Official page ↗
gpt-oss · OpenAI approximately 21 billion · 3.6 billion active · Mixture of experts (MoE) Text ← Text 131,072 tokens With the YaRN settings supplied with the published model Local reasoning with the smaller gpt-oss Apache-2.0
Official page ↗
GLM · Z.ai approximately 30 billion · 3 billion active · Mixture of experts (MoE) Text ← Text 202,752 tokens Coding agent with preserved reasoning MIT
Official page ↗
Kimi · Moonshot AI approximately 1,000 billion nominal · 32 billion active · Mixture of experts (MoE) Total parameter count unverified Text ← Text 131,072 tokens Text agent for coding and tool-based workflows Modified MIT Commercial use subject to licence
Official page ↗
Kimi · Moonshot AI approximately 1,000 billion nominal · 32 billion active · Mixture of experts (MoE) Total parameter count unverified Text ← Text 262,144 tokens Reasoning agent for long tool chains and analysis Modified MIT Commercial use subject to licence
Official page ↗
Kimi · Moonshot AI approximately 1,000 billion nominal · 32 billion active · Mixture of experts (MoE) Total parameter count unverified Text, Image, Video ← Text 262,144 tokens Multimodal agent for coding from visual designs and document analysis Modified MIT Commercial use subject to licence
Official page ↗
MiniMax · MiniMax approximately 230 billion nominal · 10 billion active · Mixture of experts (MoE) Total parameter count unverified Text ← Text 196,608 tokens Software development assistant for tool-based workflows Modified MIT Commercial use subject to licence
Official page ↗
MiniMax · MiniMax approximately 230 billion nominal · 10 billion active · Mixture of experts (MoE) Total parameter count unverified Text ← Text 196,608 tokens Tool-oriented model for coding and multi-step planning Modified MIT Commercial use subject to licence
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 235 billion · 22 billion active · Mixture of experts (MoE) Text ← Text 262,144 tokens Large direct-answer assistant with long context Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 480 billion · 35 billion active · Mixture of experts (MoE) Text ← Text 262,144 tokens Coding agent for large, multi-file repositories Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 27 billion · Dense · hybrid attention Text, Image, Video ← Text 262,144 tokens Multimodal assistant for analysis, code and documents Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 0.8 billion · Dense · hybrid attention Text, Image, Video ← Text 262,144 tokens Small model for prototyping and task specialization Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 122 billion · 10 billion active · Mixture of experts (MoE) Text, Image, Video ← Text 262,144 tokens Multimodal MoE model for analysis and tool use Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 397 billion · 17 billion active · Mixture of experts (MoE) Text, Image, Video ← Text 262,144 tokens Large multimodal assistant for complex problems Apache-2.0
Official page ↗
Llama · Meta approximately 70 billion nominal · Dense Total parameter count unverified Text ← Text 131,072 tokens Multilingual text assistant for answers and enterprise tasks llama3.3 Commercial use subject to licence
Official page ↗
Phi · Microsoft approximately 14 billion nominal · Dense Total parameter count unverified Text ← Text 16,384 tokens English model for mathematics, logic and text generation MIT
Official page ↗
SmolVLM · Hugging Face approximately 2.2 billion nominal · Dense Total parameter count unverified Text, Image, Video ← Text 8,192 tokens Small model for questions about images and video Apache-2.0
Official page ↗
MiniLM · Sentence Transformers 0.023 billion · Dense Text ← Vector 256 tokens Lightweight English embeddings for search and clustering Apache-2.0
Official page ↗
E5 · intfloat / multilingual E5 authors 0.278 billion · Dense Text ← Vector 512 tokens Multilingual embeddings with separate query and document prefixes MIT
Official page ↗
E5 · intfloat / multilingual E5 authors 0.56 billion · Dense Text ← Vector 512 tokens Multilingual embeddings for semantic search MIT
Official page ↗
BGE · BAAI 0.033 billion · Dense Text ← Vector 512 tokens Small English embedding model for document retrieval MIT
Official page ↗
BGE · BAAI 0.278 billion · Dense Text ← Structured data 512 tokens English and Chinese reranker for search results MIT
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 7.61 billion · Dense Text ← Text 32,768 tokens Small coding assistant for explaining and correcting code Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 14.7 billion · Dense Text ← Text 32,768 tokens Coding assistant for generation, explanation and debugging Apache-2.0
Official page ↗
Qwen · Qwen / Alibaba Cloud approximately 32.5 billion · Dense Text ← Text 32,768 tokens Coding assistant with more capacity for difficult problems Apache-2.0
Official page ↗
EmbeddingGemma · Google DeepMind approximately 0.3 billion nominal · Dense Total parameter count unverified Text ← Vector 2,048 tokens Small multilingual embedding model for on-device execution gemma Commercial use subject to licence
Official page ↗
Jina · Jina AI approximately 0.57 billion nominal · Dense Total parameter count unverified Text ← Vector 8,192 tokens Multilingual embedding model with task-specific adapters CC-BY-NC-4.0 Commercial use subject to licence
Official page ↗
Jina · Jina AI 0.278 billion · Dense Text ← Structured data 1,024 tokens Multilingual reranker for longer documents CC-BY-NC-4.0 Commercial use subject to licence
Official page ↗
Mixedbread · Mixedbread 0.335 billion · Dense Text ← Vector 512 tokens English embedding model for retrieval with query instructions Apache-2.0
Official page ↗
Nomic · Nomic AI 0.137 billion · Dense Text ← Vector 8,192 tokens English embedding model with long context and reducible dimensions Apache-2.0
Official page ↗
ModernBERT · Answer.AI / LightOn 0.15 billion · Dense Text ← Structured data 8,192 tokens Base encoder for training classification and entity extraction Apache-2.0
Official page ↗
Qwen · Qwen 1.54 billion · Dense Text ← Text 32,768 tokens Code completion and fill-in-the-middle apache-2.0
SAFETENSORS · Official ↗
Official page ↗
StarCoder2 · bigcode 3 billion · Dense Text ← Text 16,384 tokens Local attention with a 4,096-token window. Code completion with sliding-window attention bigcode-openrail-m Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen 80 billion · 3 billion active · Mixture of experts (MoE) · hybrid attention Text ← Text 262,144 tokens Coding agent apache-2.0
SAFETENSORS · Official ↗
Official page ↗
E5 · intfloat 0.56 billion · Dense Text ← Vector 512 tokens Multilingual retrieval with task instructions mit
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen 4 billion · Dense Text ← Text 262,144 tokens Small direct-answer assistant apache-2.0
Official page ↗
Tooka · PartAI 0.123 billion · Dense Text ← Vector 512 tokens Persian text embeddings
SAFETENSORS · Official ↗
Official page ↗
Tooka · PartAI 0.353 billion · Dense Text ← Vector 512 tokens Persian text embeddings
SAFETENSORS · Official ↗
Official page ↗
ParsBERT · HooshvareLab Dense Parameter count not recorded Text ← Structured data 512 tokens Training base for Persian tasks
PYTORCH · Official ↗
Official page ↗
Salamandra · BSC-LT 2.253 billion · Dense Text ← Text 8,192 tokens Iberian-language text generation Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
Salamandra · BSC-LT 7.768 billion · Dense Text ← Text 8,192 tokens Iberian-language text generation Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
MiniCPM · openbmb 2.517 billion · Dense Text ← Text 131,072 tokens Reasoning with a compact model Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
LFM · LiquidAI 1.17 billion · Hybrid Text ← Text 32,768 tokens On-device information extraction LFM Open License 1.0 Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Granite · ibm-granite 3.66 billion · Dense Text ← Text 131,072 tokens Reasoning with a compact model Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
GLM · Z.ai 744 billion · 40 billion active · Mixture of experts (MoE) Text ← Text 202,752 tokens Coding and tool use MIT
SAFETENSORS · Official ↗
Official page ↗
GLM · Z.ai 753.864 billion stored elements · Mixture of experts (MoE) Count stored in the checkpoint, including additional components; not active parameters per generated token. Text ← Text 202,752 tokens Coding and tool use MIT
SAFETENSORS · Official ↗
Official page ↗
GLM · Z.ai 753.33 billion stored elements · Mixture of experts (MoE) Count stored in the checkpoint, including additional components; not active parameters per generated token. Text ← Text 1,048,576 tokens Coding and tool use MIT
SAFETENSORS · Official ↗
Official page ↗
GLM · Z.ai Mixture of experts (MoE) Parameter count not recorded Text ← Text 1,048,576 tokens Coding and tool use GLM-5.3 License Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Kimi · Moonshot AI 1,000 billion · 32 billion active · Mixture of experts (MoE) Text, Image, Video ← Text 262,144 tokens Coding and tool use Modified MIT Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Kimi · Moonshot AI 1,000 billion · 32 billion active · Mixture of experts (MoE) Text, Image, Video ← Text 262,144 tokens Coding and tool use Modified MIT Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Kimi · Moonshot AI 2,800 billion · 104 billion active · Mixture of experts (MoE) Text, Image, Video ← Text 1,048,576 tokens Coding and tool use Kimi K3 License Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen 2 billion · Dense Text, Image, Video ← Vector 32,768 tokens Multimodal retrieval Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen 8 billion · Dense Text, Image, Video ← Vector 32,768 tokens Multimodal retrieval Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen 2 billion · Dense Text, Image, Video ← Structured data 32,768 tokens Multimodal reranking Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen 8 billion · Dense Text, Image, Video ← Structured data 32,768 tokens Multimodal reranking Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen / Alibaba Cloud 27.781 billion · approximately 27 billion Language component; publisher rounded count · Dense · hybrid attention Text, Image, Video ← Text 262,144 tokens Up to 1,010,000 tokens with Requires YaRN configuration; native context is 262,144 tokens. Text and vision model with hybrid attention Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
Qwen · Qwen / Alibaba Cloud 35.952 billion · 3 billion active · approximately 35 billion Language component; publisher rounded count · Mixture of experts (MoE) · hybrid attention Text, Image, Video ← Text 262,144 tokens Up to 1,010,000 tokens with Requires YaRN configuration; native context is 262,144 tokens. Text and vision model with hybrid attention Apache-2.0
SAFETENSORS · Official ↗
Official page ↗
HY-MT · Tencent 1.8 billion · Dense Text ← Text Dedicated text translation Tencent HY Community License Commercial use subject to licence
Official page ↗
HY-MT · Tencent 7 billion · Dense Text ← Text Dedicated text translation Tencent HY Community License Commercial use subject to licence
Official page ↗
HY-MT · Tencent 1.8 billion · Dense Text ← Text Specialized translation Apache-2.0
Official page ↗
HY-MT · Tencent 7 billion · Dense Text ← Text Specialized translation Apache-2.0
Official page ↗
HY-MT · Tencent 30 billion · 3 billion active · Mixture of experts (MoE) Text ← Text Specialized translation Apache-2.0
Official page ↗
Gemma · Google 4 billion · Dense Text, Image ← Text Specialized translation Gemma terms Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Gemma · Google 12 billion · Dense Text, Image ← Text Specialized translation Gemma terms Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
Gemma · Google 27 billion · Dense Text, Image ← Text Specialized translation Gemma terms Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
MADLAD · Google 3 billion · Dense Text ← Text Specialized translation Apache-2.0
Official page ↗
NLLB · Meta 0.6 billion · Dense Text ← Text Specialized translation CC-BY-NC-4.0 Commercial use subject to licence
PYTORCH · Official ↗
Official page ↗
Qwen-Image · Qwen / Alibaba approximately 7 billion DiT image generator · approximately 8 billion Qwen3-VL encoder · Other Total parameter count unverified Text, Image ← Image Image generation and editing Qwen Research License — non-commercial Commercial use subject to licence
SAFETENSORS · Official ↗
Official page ↗
MiMo · Xiaomi MiMo 1,020 billion · 42 billion active · Mixture of experts (MoE) · hybrid attention Text, Image, Video, Audio ← Text 1,048,576 tokens Coding, software agents, text and image analysis MIT
SAFETENSORS · Official ↗

“—” means no information is recorded, not that the model lacks the capability.

For RAG, examine retrieval, reranking and answer generation separately.
Data view

Model fit for the task

Models for conversation, programming, search and document tasks.

Test shown in the result column: MIRACL · nDCG@10 · en

This selection applies to result columns across tables. Execution conditions are available in each model’s details.

Results grouped by test ←
Which model should I consider?

Documented starting points, not a quality ranking.

Spanish retrieval

E5-Small is a smaller baseline; Large-Instruct scores higher in this Spanish MIRACL report. The Qwen results here are multilingual aggregates, not Spanish-specific scores.

When queries and documents are both Spanish; identify cross-language retrieval separately.

The 51.2-to-53.7 difference establishes neither final-answer quality nor coverage of every Spanish variety.

English retrieval

Treat Large and Large-Instruct as separate variants: Large scores higher on English MIRACL in this report. For a reranking step, use Qwen’s comparison with a common candidate pool.

When target-language evidence matters more than a multilingual average.

The Instruct label does not guarantee an improvement on every task.

Adding reranking

Start the size comparison with 0.6B. Consider 4B and 8B when their task-specific gains justify a second-stage resource budget; 8B does not lead on every metric.

When relevant documents reach the candidate pool but rank poorly; candidate count is part of the design.

Reranking cannot recover a document absent from its input pool.

Coding problems or code edits?

For an instruction-following assistant, distinguish repository editing from coding problems. StarCoder2-3B serves code completion/FIM; HumanEval does not measure chat or editing-agent quality.

Once the interface is known: editor completion, multi-file changes, or solving a programming problem.

MoE active parameters do not determine resident weight memory.

A small model for a bounded task

LFM is a small candidate for extraction and bounded workflows; MiniCPM and Granite also have reasoning and tool-use evidence. Actual weight size can differ from the rounded size in a model name.

When task scope, output format and memory constraints are defined.

A strong small-model score establishes neither lower latency nor laptop feasibility at maximum context.

Selecting a Spanish-language generator

Salamandra supplies direct evidence for several Spanish tasks. Keep it alongside general multilingual candidates; the available data does not identify a best Spanish chatbot.

When Spanish evidence must be distinguished from a general multilingual claim.

Language coverage is not evidence for every region or professional domain.

Intent and text classification

Use MassiveIntent and MTOP evidence for intent classification. An embedding model is one component of the classifier; the result also depends on the classifier and its training data.

For a defined set of labels; the published evidence supports an initial shortlist.

An embedding model or bare ParsBERT is not a ready-to-use chatbot.

SmolLM3-3B

A 3B model with two thinking modes. For short replies, choose non-thinking mode and an explicit output budget.

en: Publisher declaration
SmolLM3-3B — model card · Publisher report https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Revision / commit
a07cc9a04f16550a088caea529712d1d335b0ac1
Relevant source section
README.md: model description / architecture / intended use; matching lines 26, 33, 345
  • Model announcement counts are rounded; safetensors.total counts stored elements, not necessarily unique parameters.
SmolLM3-3B — Hub metadata · Publisher report https://huggingface.co/api/models/HuggingFaceTB/SmolLM3-3B?blobs=true ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Revision / commit
a07cc9a04f16550a088caea529712d1d335b0ac1
Relevant source section
$.sha; $.id; $.pipeline_tag; $.cardData; $.safetensors; $.siblings
  • A file inventory and tensor count do not equal runtime memory or unique parameters.
SmolLM3-3B — parameter specification · Publisher report https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Revision / commit
a07cc9a04f16550a088caea529712d1d335b0ac1
Relevant source section
README.md: parameter count / named model variant; matching lines 14, 35, 59, 90, 151, 160, 210, 216, 221, 225, 246, 261
  • Model announcement counts are rounded; safetensors.total counts stored elements, not necessarily unique parameters.
SmolLM3-3B — license declaration · Publisher report https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Revision / commit
a07cc9a04f16550a088caea529712d1d335b0ac1
Relevant source section
README.md: license / underlying-model terms; Hub $.cardData.license; matching lines 3, 31, 381, 382
SmolLM3-3B — configuration · Publisher report https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/config.json ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Revision / commit
a07cc9a04f16550a088caea529712d1d335b0ac1
Relevant source section
config.json: architectures, model_type, text_config, max_position_embeddings, quantization_config, torch_dtype/dtype
SmolLM3 release · Publisher report https://huggingface.co/blog/smollm3 ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Relevant source section
Publication date July 8 2025; SmolLM3-3B
  • Release date of the named variants, not repository creation or data extraction.
SmolLM3-3B — declared context · Publisher report https://huggingface.co/HuggingFaceTB/SmolLM3-3B/blob/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Relevant source section
Key features; Long context processing: trained at 64k; YaRN for 128k
  • Documented capacity of the named variant; separate from a test's output length or a hosting API's limit
HuggingFaceTB/SmolLM3-3B — model card · Publisher report https://huggingface.co/HuggingFaceTB/SmolLM3-3B/raw/a07cc9a04f16550a088caea529712d1d335b0ac1/README.md ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Relevant source section
Description; model details; usage
HuggingFaceTB/SmolLM3-3B — repository metadata · Publisher report https://huggingface.co/api/models/HuggingFaceTB/SmolLM3-3B?blobs=true ↗
Publisher / author
Hugging Face
Accessed
2026-09-15
Relevant source section
cardData; sha; safetensors; siblings
117 results out of 117 rows
Application
Declared model size(billion)
Recommendation basis
Role in the system
Advanced filters 3 controls
Model type
All rows
Row comparison and conditions

Comparison rule: Models may differ; task, language, dataset, benchmark version, metric and unit must be matched or disclosed.

Use × beside a heading to hide its column. Open row details for sources and conditions.

Hidden columns:
Model application guide
CompareDetails
Role in the system
Primary use
Distinguishing feature and rationale
Relevant usage condition
Recommendation basis
Get started
Official page ↗
Document retrieval Hybrid document retrieval in RAG Dense, sparse and multi-vector outputs in one model; over 100 languages, 8,192-token inputs and 1,024-dimensional dense vectors. Use FlagEmbedding for all three outputs together; GGUF or Ollama output depends on the backend. Guide's analytical recommendation
Official page ↗
Document reranking Reorder retrieved documents Reads the query and document together to score relevance; complements BGE-M3 retrieval. Retrieve a limited candidate set first. A sigmoid score is not answer correctness probability; the output is not an embedding. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction Multilingual writing and chat with long documents Aya Expanse 32B has a 131,072-token context; Persian appears among the publisher's 23 declared languages. Research release under a noncommercial license; context length does not establish uniform accuracy throughout a document. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction Multilingual writing and rewriting assistant Aya Expanse 8B combines an 8,192-token context with multilingual preference training; it is also a candidate for Persian writing evaluation. Observe this variant's noncommercial license and context cap; 32B results do not transfer to it. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming Local chat across languages Tiny Aya Global provides broad language coverage; it is distinct from the regional Earth, Fire and Water variants. Evaluate Persian examples for Persian use; regional or base variants cannot be substituted for this checkpoint without review. Guide's analytical recommendation
Official page ↗
Tool calling, Text generation, Document-grounded generation, Structured extraction, Programming Prototype a small assistant with function calling The largest SmolLM2 variant listed supports a function-calling format as well as instruction following; the two smaller variants do not inherit that feature. Primarily an English model. The host application executes functions and must validate arguments. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction Explore instruction following with a very small model The 135M-parameter SmolLM2 variant is suitable for simple text-generation prototypes and exploring very small models' limits. Keep inputs short and tasks narrow; do not expect broad knowledge or the 1.7B variant's tool calling. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction Short-text rewriting and summarization SmolLM2 360M uses SFT and DPO for instruction tuning; its model card includes a CPU execution example. English-focused; small model size does not establish speed or Persian quality. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Small assistant with selectable thinking mode SmolLM3 3B offers direct-answer and reasoning modes, a training context of 65,536 tokens and published training details. Longer context requires YaRN. The card names six native languages including German, while repository tags list eight different languages; neither list includes Persian. Requires Transformers 4.53 or later., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Prototype chat with the smallest Qwen3 Dense Qwen3 0.6B offers thinking and non-thinking modes; an official Q8_0 file is available. Limit the output budget; the family name does not establish the problem-solving ability of larger variants., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Small text assistant with reasoning control Qwen3-1.7B is among the sub-2B options listed; its message template can select direct-answer mode. In thinking mode, reasoning tokens count toward cost and output length. This record's official GGUF is Q8_0., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Text generation and analysis with dense Qwen3 Qwen3-14B is a dense model with two answer modes; official GGUF variants offer a choice of weight precision. Q4_K_M and Q8_0 packages are available; their speed and quality have not been tested by this guide., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming Tool-oriented assistant with MoE architecture Qwen3-30B-A3B uses selected experts and the Qwen-Agent pattern for connecting tools. Active parameters do not replace total weight size in memory estimates; execution requires compatible message templates and parsers., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Reasoning, Programming, Text generation, Document-grounded generation, Structured extraction, Tool calling Multi-step analysis with a larger dense Qwen3 Unlike 30B-A3B, Qwen3-32B is dense; reasoning mode can be chosen according to question complexity. Compare direct-answer and thinking modes under matched output budgets and conditions; this overview is not a quality ranking., This is a general model; the specialized Coder has different training and tool formats., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling General assistant at four billion parameters Qwen3-4B switches thinking mode within one checkpoint; official GGUF files offer several quantizations. Native context is 32,768 tokens; extending it requires context-extension settings., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Chat and text tasks with thinking control Qwen3-8B is a general dense model with 8.2B declared parameters; its message template selects direct answers or reasoning. Official Q4_K_M and Q8_0 files are available for local use; this guide has not measured their quality or runtime memory., Retrieval, document insertion and citation checks belong to the system; this relationship is not a RAG or Persian benchmark., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Programming, Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning Repository editing with a tool-using assistant Qwen3-Coder-30B-A3B-Instruct is tuned for code generation, repository editing and coding-agent workflows, with a native context of 262,144 tokens. This is a non-thinking variant. Use Coder's tool-calling format; it is distinct from general Qwen3-30B-A3B., The host application supplies file, testing and execution tools., This checkpoint produces direct answers only; it has no separate thinking mode. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Document retrieval Vector retrieval with the smallest Qwen3 Embedding The 0.6B variant produces instruction-aware embeddings with up to 1,024 dimensions and 32,768-token inputs. Add the task instruction to the query; index documents using the same model and dimension settings. Guide's analytical recommendation
Official page ↗
Document retrieval Multilingual indexing with adjustable vectors Qwen3-Embedding-4B outputs up to 2,560 dimensions and uses query instructions to adapt representations to a task. Dimension reduction requires checking retrieval quality on the intended collection; an MTEB score cannot substitute for that check. Guide's analytical recommendation
Official page ↗
Document retrieval Document retrieval with the largest listed Qwen3 Embedding Qwen3-Embedding-8B supports up to 4,096 dimensions and 32,768 tokens of context; text goes in and vectors come out. Account for indexing and vector storage separately from answer generation; this model does not write the final answer. Guide's analytical recommendation
Official page ↗
Document reranking Lighter reranking in the Qwen3 family Qwen3-Reranker 0.6B is tuned to score query–document relevance with task instructions. Requires the reranker template and yes/no scoring; do not use the chat or embedding-generation path. Guide's analytical recommendation
Official page ↗
Document reranking Instruction-aware reranking of candidate documents Qwen3-Reranker-4B is the middle option; the query, retrieval instruction and document all contribute to relevance scoring. Pass only candidate documents to this stage; more query–document pairs increase second-stage cost. Guide's analytical recommendation
Official page ↗
Document reranking Reranking with the 8B Qwen3 variant Qwen3-Reranker-8B is the largest reranker in this collection; it scores query–document relevance. To choose between this and 4B, compare retrieval quality and latency on your documents; size alone is not a ranking. Guide's analytical recommendation
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming Read images and documents with Qwen3-VL Qwen3-VL 8B Instruct combines text, image and video inputs for visual understanding and document tasks. Use the multimodal processor and message template; Thinking and Instruct checkpoints behave differently. Guide's analytical recommendation
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Multimodal prototyping with small Qwen3.5 Small Qwen3.5-2B is presented for prototyping and task-specific fine-tuning; it accepts visual inputs alongside text. Larger family variants' quality does not transfer to this 2B model., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Tool calling, Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming Multimodal agent with MoE Qwen3.5 The 35B-A3B variant uses an integrated text–image foundation and hybrid architecture, offering a candidate for multi-step workflows. Hosted Qwen3.5-Flash has different tools and default context; API specifications do not transfer to local weights., Use the multimodal settings and image limits documented for these weights., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Text and image processing at four billion parameters Qwen3.5-4B belongs to the integrated text–image generation; it is a smaller-than-9B candidate for document information extraction. Scanned documents require the multimodal path; the text context limit does not determine how many images can be processed., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Document and image assistant with dense Qwen3.5 Qwen3.5 9B combines text and images in a dense architecture. Compare against 4B with the same document data and image settings; distinguish language and vision parameter counts., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Tool calling, Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming Long workflows with a text–image agent Qwen3.8-27B is a dense text–image model with thinking control; its announcement emphasizes coding, research and multi-step work. Promised cloud-service features, such as built-in tools and default context, are not guarantees for local weights., Image dimensions and frame sampling are part of the usage conditions., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Instruction-following research with inspectable training Olmo 3 7B Instruct includes Dolma 3 and Dolci data; published training details distinguish it for reproducible research. This checkpoint is Instruct; do not attribute Olmo Think results to it. Requires Transformers 4.57 or later. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Explore reasoning distilled from R1-0528 Post-trained from Qwen3-8B-Base using DeepSeek-R1-0528 reasoning chains; distinct from standard Qwen3-8B. Full R1-0528 benchmark results do not belong to this 8B variant., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming Distilled reasoning on a Llama 70B base This checkpoint distills R1 onto a Llama base and is the largest distilled variant listed. Both the Llama base license and distillation publisher's terms matter; estimate weight size and reasoning-chain length separately., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming Explore reasoning limits in a very small distilled model The 1.5B variant distills R1 onto Qwen2.5-Math, allowing exploration of reasoning behavior transferred to a small model. Long responses can become repetitive., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming Problem solving with a mid-sized R1 distillation DeepSeek-R1-Distill-Qwen-14B uses Qwen2.5-14B and R1-generated data; its weights are not the full R1 model. Reported numbers apply to this variant and evaluation settings; avoid unmatched comparisons with direct-answer modes., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming Analysis and coding with a dense 32B distillation The Qwen-32B variant distills R1 onto a Qwen2.5 base. Declared reasoning ability is not tool calling or automatic code execution; the application must provide tools and answer validation., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming Problem solving with the 7B R1 distillation Based on Qwen2.5-Math-7B; distinct from the later R1-0528 distillation onto Qwen3. Keep this repository's template and tokenizer; comparison with the 8B variant requires a shared evaluation., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming Connect reasoning and tools in agent workflows DeepSeek-V3.2 combines DSA sparse attention with agent-oriented post-training; its chat template differs from earlier versions. Use V3.2-specific templates and execution paths; Speciale results do not transfer to this model., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Tool calling, Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming Multimodal agent for very long inputs DeepSeek-V4.1-Flash processes text and images with CED architecture; active parameter counts differ between prefill and decode. Active parameters are not a single fixed count; KV compression and a specialized execution path are part of this architecture., Use the CED-specific execution path; this relationship is not an OCR evaluation result., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation, Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming Ask about images and text with mid-sized Gemma 3 Gemma 3 12B IT is an instruction-tuned multimodal variant with 131,072 tokens of context, connecting image inputs to text answers. Requires the vision processor and Gemma 3 template; maximum context does not establish constant long-document comprehension. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming Small text assistant from the Gemma 3 family Unlike larger variants in the generation, Gemma 3 1B IT is text-only and has 32,768 tokens of context. Scanned images first require external OCR; the multimodal features of 4B and above do not apply to this variant. Guide's analytical recommendation
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming Text and image understanding with the largest listed Gemma 3 Gemma 3 27B IT is this generation's larger dense variant with visual input and multilingual coverage. Downloading weights requires accepting the Gemma license. Guide's analytical recommendation
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming Getting started with images in the Gemma 3 family Gemma 3 4B IT is the smallest multimodal model of this generation listed here; unlike 1B, it can read images alongside text. Download the 4B weights and processor together; validate extracted document numbers in the text output. Guide's analytical recommendation
Official page ↗
Reasoning, Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Multimodal reasoning with MoE Gemma 4 Gemma 4 26B-A4B is an expert model with thinking control and 262,144 tokens of context; its architecture differs from small E2B. This variant has no audio input; active parameters differ from all model weights., Requires the image processor path; E2B audio support does not transfer to this model., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Local text, image and audio processing with small Gemma 4 Gemma 4 E2B is presented for on-device execution and also supports audio input. E2B is an effective count, not all weights. Use total counts and actual files for memory calculations; audio and image paths need their matching processors., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Document-grounded generation, Structured extraction, Text generation, Programming, Tool calling Generate answers from retrieved documents Granite 3.3 2B Instruct is a small business-oriented model supporting RAG, summarization, text extraction and thinking mode. The system must retrieve documents and place them in the prompt; the model is not a search engine. Persian is not among its 12 declared languages., Validate output against the application's schema; requesting a format alone does not guarantee valid JSON. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Document retrieval Multilingual retrieval with low-dimensional vectors Multilingual E5 Small converts text to 384-dimensional vectors with a 512-token input cap, making it a candidate for short document chunks. query: and passage: prefixes are required even outside English; split long text before indexing. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling General assistant with large Llama 3.1 Llama 3.1 70B Instruct supports multilingual chat and text tasks with 131,072 tokens of context. This variant takes text input; weight use is subject to the Llama license. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Chat and text work with Llama 3.1 8B Llama 3.1 8B is tuned for general assistance and instruction following; it differs from the base model of the same size. Keep the Instruct template and tokenizer; long context alone does not guarantee document-grounded answer quality. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Local rewriting and summarization with small Llama Llama 3.2 1B Instruct is this generation's small text model, presented for narrow text tasks on local devices. Does not accept images directly; Llama 3.2 Vision models are separate products. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling On-device text assistant with Llama 3B Llama 3.2 3B Instruct is tuned for chat, rewriting and summarization, with more parameters than the 1B variant. Compare quality and latency on the intended task; neither the 1B nor the 3B text model accepts images. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Analysis and logic under tighter resource constraints Phi-4-mini-instruct emphasizes reasoning data and instruction following, with 131,072 tokens of context. Distinct from mini-reasoning and multimodal-instruct; compare mathematical or logical examples using this exact variant name. Guide's analytical recommendation
Official page ↗
Programming, Tool calling, Text generation, Document-grounded generation, Structured extraction, Image understanding Multi-file editing and repository-search agent Devstral Small 2 targets software engineering and tool-assisted code inspection and editing; this generation also accepts images. This repository's Instruct weights are FP8; use the Mistral format and dependencies for this revision, not generic chat-model settings., Agent tools and Mistral templates must match the model revision., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Image understanding, Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Multimodal assistant for edge deployment Ministral 3 3B Instruct is the family's small text–image model; the language component has 3.4B parameters and the vision encoder 0.4B. Official weights are FP8; 3B in the model name does not count every language and vision parameter. Guide's analytical recommendation
Official page ↗
Tool calling, Text generation, Document-grounded generation, Structured extraction Text assistant with Mistral function-calling format Mistral 7B Instruct v0.3 adds the third tokenizer version and function calling to this text model. Use the tool format and v3 tokenizer with this checkpoint; this variant has no image input. Guide's analytical recommendation
Official page ↗
Image understanding, Structured extraction, Text generation, Document-grounded generation, Programming, Tool calling Multilingual text and image assistant Mistral Small 3.1 24B Instruct adds image processing and long context to Small; Persian appears in the publisher's language list. Tools and JSON output need a matching template and parser; do not mix Base and Instruct results., Configure the parser and schema validation in the application layer. Guide's analytical recommendation, Publisher-documented use
Official page ↗
Document-grounded generation, Text generation, Structured extraction, Reasoning, Programming, Tool calling Document-grounded answers with a reasoning budget Nemotron Nano 9B v2 combines Mamba-2 with attention; reasoning mode and its token budget are controllable. Retrieved documents must be supplied externally; Persian is not among the six declared languages, and the backend must support the hybrid architecture., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning Adjustable-reasoning agent with the larger gpt-oss gpt-oss-120b is an MoE model with 117B total and 5.1B active parameters; expert weights are released in MXFP4. Correct execution requires harmony format; the host supplies browser or Python tools. Publisher memory claims are not site measurements., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Reasoning, Text generation, Document-grounded generation, Structured extraction, Tool calling Local reasoning with the smaller gpt-oss gpt-oss-20b has 21B total and 3.6B active parameters, targets local or specialized use and offers three reasoning levels. Requires harmony format and an MXFP4-compatible backend; do not substitute the 20B name for actual weight counts in calculations., For simple answers, disable thinking where supported or limit the output budget. Guide's analytical recommendation
Official page ↗
Tool calling, Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming Coding agent with preserved reasoning GLM-4.7-Flash is a 30B-A3B-class MoE model; its card describes preserving thinking between turns for multi-step agents. vLLM and SGLang paths depend on specific development versions., For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Text agent for coding and tool-based workflows Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Reasoning agent for long tool chains and analysis For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding Multimodal agent for coding from visual designs and document analysis For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Software development assistant for tool-based workflows For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Tool-oriented model for coding and multi-step planning For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Large direct-answer assistant with long context This checkpoint produces direct answers only; it has no separate thinking mode. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Coding agent for large, multi-file repositories This checkpoint produces direct answers only; it has no separate thinking mode. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding Multimodal assistant for analysis, code and documents For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding Small model for prototyping and task specialization For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding Multimodal MoE model for analysis and tool use For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling, Image understanding Large multimodal assistant for complex problems For simple answers, disable thinking where supported or limit the output budget., Count thinking tokens toward response time and context., Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming, Tool calling Multilingual text assistant for answers and enterprise tasks Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming English model for mathematics, logic and text generation Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Image understanding Small model for questions about images and video Requires this revision's image processor and vision components; image count, resolution and frames affect memory use. Guide's analytical recommendation
Official page ↗
Document retrieval Lightweight English embeddings for search and clustering Mean token pooling with attention mask and L2 normalization Publisher-documented use
Official page ↗
Document retrieval Multilingual embeddings with separate query and document prefixes Masked mean pooling and L2; query: for queries and passage: for documents Publisher-documented use
Official page ↗
Document retrieval Multilingual embeddings for semantic search Masked mean pooling and L2; query: for queries and passage: for documents Publisher-documented use
Official page ↗
Document retrieval Small English embedding model for document retrieval CLS pooling and normalization; retrieval instruction on queries only Publisher-documented use
Official page ↗
Document reranking English and Chinese reranker for search results Cross-encoder on query–document pairs; relevance score Publisher-documented use
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming Small coding assistant for explaining and correcting code Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming Coding assistant for generation, explanation and debugging Guide's analytical recommendation
Official page ↗
Text generation, Document-grounded generation, Structured extraction, Programming Coding assistant with more capacity for difficult problems Guide's analytical recommendation
Official page ↗
Document retrieval Small multilingual embedding model for on-device execution Over 100 languages; Matryoshka dimension reduction down to 128 Publisher-documented use
Official page ↗
Document retrieval Multilingual embedding model with task-specific adapters retrieval.query / retrieval.passage adapters; mean pooling and L2 Publisher-documented use
Official page ↗
Document reranking Multilingual reranker for longer documents Multilingual reranking with inputs up to 1,024 tokens; noncommercial license Publisher-documented use
Official page ↗
Document retrieval English embedding model for retrieval with query instructions Query instruction: Represent this sentence for searching relevant passages: Publisher-documented use
Official page ↗
Document retrieval English embedding model with long context and reducible dimensions search_query: prefix for queries and search_document: for documents; normalization and dimension reduction Publisher-documented use
Official page ↗
Structured extraction Base encoder for training classification and entity extraction Bidirectional encoder; token output or trained task head Publisher-documented use
Official page ↗
Code completion / FIM Code completion and fill-in-the-middle Base coding model with about 1.54B parameters; for FIM and code continuation, not instruction chat. Use this model's completion / FIM format; do not substitute an instruction model's chat template. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Code completion / FIM Code completion with sliding-window attention Base coding model; 16,384-token input with a 4,096-token attention window. For completion, not a chat assistant. Use this model's completion / FIM format; do not substitute an instruction model's chat template. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Programming Coding agent MoE model with 80B total and 3B active parameters, hybrid attention and direct answers; weight memory comes from the full model. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Document retrieval Multilingual retrieval with task instructions 1,024-dimensional embeddings; add a one-sentence instruction to queries and pass documents without it. Query format: Instruct: … Query: …; documents without instructions. Masked mean pooling and L2 normalization; maximum 512 tokens. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Text generation Small direct-answer assistant Instruction-tuned 4B variant with a 262,144-token text limit; this checkpoint has no thinking mode. Publisher-documented use
Official page ↗
Document retrieval Persian text embeddings Small Tooka-SBERT-V2 variant with 768-dimensional vectors; a Persian-focused candidate for retrieval and text similarity. The retrieved metadata does not specify a usage license. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Document retrieval Persian text embeddings Large Tooka-SBERT-V2 variant with 1,024-dimensional vectors; PTEB results cannot be ranked directly against other benchmarks. The retrieved metadata does not specify a usage license. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Structured extraction Training base for Persian tasks Base ParsBERT for Persian text understanding; classification and NER require task heads and training. Not a ready-made retrieval embedding model. Base weights alone are not a ready-to-use classifier or NER model., The retrieved metadata does not specify a usage license. Publisher-documented use
PYTORCH · Official ↗
Official page ↗
Text generation Iberian-language text generation An Iberian-language reference candidate with published Spanish task results. The 2B instruction model is a language-specific comparison point; these results do not establish a universal winner. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Text generation Iberian-language text generation An Iberian-language reference candidate with published Spanish task results. The 7B instruction model is a language-specific comparison point; these results do not establish a universal winner. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Reasoning Reasoning with a compact model A compact candidate for reasoning and tool workflows, with about 2.52 billion parameters reported in repository metadata. Separate the developer’s evaluations from relayed Artificial Analysis scores. This package establishes no Spanish or Persian task quality. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Structured extraction On-device information extraction A compact candidate for extraction and bounded on-device tasks. The publisher advises against programming and knowledge-intensive use. Spanish is declared, but this package has no task-level Spanish quality score. It uses a custom license. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Reasoning Reasoning with a compact model A Granite reasoning model with thinking and non-thinking modes. The 3B label is a size name; repository metadata reports about 3.66 billion parameters. Keep the native 128K context separate from the claimed extension to 512K. Publisher-documented use
SAFETENSORS · Official ↗
Official page ↗
Document retrieval Multimodal retrieval Text, image and video retrieval with 64–2048 dimensional vectors and a task context of 32K tokens. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Document retrieval Multimodal retrieval Text, image and video retrieval with 64–4096 dimensional vectors and a task context of 32K tokens. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Document reranking Multimodal reranking Scores query relevance for text, images and video; reranks retrieved candidates with up to 32K task tokens. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Document reranking Multimodal reranking Scores query relevance for text, images and video; reranks retrieved candidates with up to 32K task tokens. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Text generation, Image understanding, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Text and vision model with hybrid attention Qwen3.6-27B accepts text, images and video. Official BF16 weights are recorded; hybrid-attention runtime memory needs engine-specific measurement. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Text generation, Image understanding, Document-grounded generation, Structured extraction, Reasoning, Programming, Tool calling Text and vision model with hybrid attention Qwen3.6-35B-A3B accepts text, images and video. Official BF16 weights are recorded; hybrid-attention runtime memory needs engine-specific measurement. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Text generation Dedicated text translation Dedicated translation model with terminology and contextual prompts; not selected for general chat or document Q&A. Guide's analytical recommendation
Official page ↗
Text generation Dedicated text translation Dedicated translation model with terminology and contextual prompts; not selected for general chat or document Q&A. Guide's analytical recommendation
Official page ↗
Text generation Specialized translation Second-generation Hy-MT translator with translation instructions and terminology control. Guide's analytical recommendation
Official page ↗
Text generation Specialized translation Second-generation Hy-MT translator with translation instructions and terminology control. Guide's analytical recommendation
Official page ↗
Text generation Specialized translation Second-generation Hy-MT translator with translation instructions and terminology control. Guide's analytical recommendation
Official page ↗
Text generation Specialized translation Text and image-text translator with a documented 2K input limit and a translation-specific template. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Text generation Specialized translation Text and image-text translator with a documented 2K input limit and a translation-specific template. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Text generation Specialized translation Text and image-text translator with a documented 2K input limit and a translation-specific template. Guide's analytical recommendation
SAFETENSORS · Official ↗
Official page ↗
Text generation Specialized translation T5 translator with broad language coverage; the target-language prefix is part of its input. Guide's analytical recommendation
Official page ↗
Text generation Specialized translation Research translation baseline, noncommercial license and training inputs up to 512 tokens. Guide's analytical recommendation
PYTORCH · Official ↗
Official page ↗
Programming Coding, software agents, text and image analysis Xiaomi’s multimodal model for coding, tool use and long inputs; the RL checkpoint has downloadable weights with MIT license metadata. About 1.02T total and 42B active parameters; weight files occupy 534.1 GiB. This checkpoint does not fit on one 24 or 48 GB GPU., Artificial Analysis Intelligence Index v4.3.2: 46 for the MiMo-V2.6-Pro service; this is neither a Persian evaluation nor a local RL-checkpoint test. Guide's analytical recommendation
SAFETENSORS · Official ↗

“—” means no information is recorded, not that the model lacks the capability.

Data view

Hardware feasibility

Memory for weights, KV and runtime under the selected settings.

Calculated from specifications. Context includes history, input and output. Active requests differ from daily users.

GPU configurations · 3 Selected configuration
GPU specifications in the GPU guide →
Calculation method and scenario limits

Actual weight file size + KV cache + runtime reserve. GGUF scenarios use FP16 KV cache. The default GPU reserve is 2 GiB per card; the CPU reserve is 4 GiB. These reserves are planning assumptions.

MoE calculations include all weights. Nominal GPU capacity is used as a GiB budget; deployment planning should use the device's reported free memory. The sum of several GPUs' memory is not unified memory.

SmolLM2 FP32 scenarios use FP32 for both weights and KV cache. Scenarios above the selected file's context limit are excluded, including Aya Expanse 32B with the selected GGUF file's 8,192-token limit.

56 results out of 56 rows
Advanced filters 1 controls
Required memory(GiB)
All rows
Row comparison and conditions

Comparison rule: Hardware cells share a row's scenario. Comparing rows requires matching non-hardware conditions.

Use × beside a heading to hide its column. Open row details for sources and conditions.

Artifact feasibility by hardware configuration
CompareDetails
Model
Weight version
RTX 4090 24GB
RTX 5090 32GB
RTX 6000 Ada 48GB
Official page ↗
Q8_0 3.47 GiB 0.6 GiB 0.88 GiB
Official page ↗
Q8_0 4.58 GiB 1.71 GiB 0.88 GiB
Official page ↗
Q4_K_M 5.45 GiB 2.33 GiB 1.13 GiB
Official page ↗
Q8_0 7.11 GiB 3.99 GiB 1.13 GiB
Official page ↗
Q4_K_M 7.81 GiB 4.68 GiB 1.13 GiB
Official page ↗
Q8_0 11.24 GiB 8.11 GiB 1.13 GiB
Official page ↗
Q4_K_M 11.63 GiB 8.38 GiB 1.25 GiB
Official page ↗
Q8_0 17.87 GiB 14.62 GiB 1.25 GiB
Official page ↗
Q4_K_M 20.03 GiB 17.28 GiB 0.75 GiB
Official page ↗
Q8_0 33 GiB 30.25 GiB 0.75 GiB
Official page ↗
Q4_K_M 22.4 GiB 18.4 GiB 2 GiB
Official page ↗
Q8_0 36.43 GiB 32.43 GiB 2 GiB
Official page ↗
Q4_K_M 44.1 GiB 39.6 GiB 2.5 GiB
Official page ↗
Q8_0 74.33 GiB 69.83 GiB 2.5 GiB
Official page ↗
Q4_K_M 11.87 GiB 8.37 GiB 1.5 GiB
Official page ↗
Q8_0 18.12 GiB 14.62 GiB 1.5 GiB
Official page ↗
Q4_K_M 22.49 GiB 18.49 GiB 2 GiB
Official page ↗
Q8_0 36.43 GiB 32.43 GiB 2 GiB
Official page ↗
Q4_K_M 6.8 GiB 4.36 GiB 0.44 GiB
Official page ↗
Q8_0 9.98 GiB 7.54 GiB 0.44 GiB
Official page ↗
Q4_K_M 44.1 GiB 39.6 GiB 2.5 GiB
Official page ↗
Q8_0 74.33 GiB 69.83 GiB 2.5 GiB
Official page ↗
Q4_K_M 7.58 GiB 4.58 GiB 1 GiB
Official page ↗
Q8_0 10.95 GiB 7.95 GiB 1 GiB
Official page ↗
Q4_K_M 21.69 GiB 18.44 GiB 1.25 GiB
Official page ↗
Q8_0 35.22 GiB 31.97 GiB 1.25 GiB
Official page ↗
Q4_K_M 7.71 GiB 4.71 GiB 1 GiB
Official page ↗
Q8_0 10.95 GiB 7.95 GiB 1 GiB
Official page ↗
Q4_K_M 4.48 GiB 0.98 GiB 1.5 GiB
Official page ↗
Q4_K_M 2.27 GiB 0.1 GiB 0.18 GiB
Official page ↗
Q8_0 2.31 GiB 0.13 GiB 0.18 GiB
Official page ↗
Q8_0 2.67 GiB 0.36 GiB 0.31 GiB
Official page ↗
Q4_K_M 20.03 GiB 17.28 GiB 0.75 GiB
Official page ↗
Q8_0 33 GiB 30.25 GiB 0.75 GiB
Official page ↗
Q4_K_M 7.81 GiB 4.68 GiB 1.13 GiB
Official page ↗
Q8_0 11.24 GiB 8.11 GiB 1.13 GiB
Official page ↗
Q4_K_M 3.26 GiB 1.04 GiB 0.22 GiB
Official page ↗
Q8_0 3.98 GiB 1.76 GiB 0.22 GiB
Official page ↗
Q4_K_M 4.06 GiB 1.44 GiB 0.63 GiB
Official page ↗
Q8_0 5.13 GiB 2.51 GiB 0.63 GiB
Official page ↗
Q4_K_M 5.32 GiB 2.32 GiB 1 GiB
Official page ↗
Q8_0 6.8 GiB 3.8 GiB 1 GiB
Official page ↗
Q4_K_M 7.07 GiB 4.07 GiB 1 GiB
Official page ↗
Q8_0 10.17 GiB 7.17 GiB 1 GiB
Official page ↗
Q4_K_M 136.32 GiB 132.85 GiB 1.47 GiB
Official page ↗
Q8_0 236.24 GiB 232.77 GiB 1.47 GiB
Official page ↗
Q4_K_M 274.8 GiB 270.86 GiB 1.94 GiB
Official page ↗
Q8_0 479.24 GiB 475.3 GiB 1.94 GiB
Official page ↗
Q4_K_M 11.99 GiB 8.43 GiB 1.56 GiB
Official page ↗
Q8_0 18.07 GiB 14.51 GiB 1.56 GiB
Official page ↗
Q4_K_M 6.8 GiB 4.36 GiB 0.44 GiB
Official page ↗
Q8_0 9.98 GiB 7.54 GiB 0.44 GiB
Official page ↗
Q4_K_M 11.87 GiB 8.37 GiB 1.5 GiB
Official page ↗
Q8_0 18.12 GiB 14.62 GiB 1.5 GiB
Official page ↗
Q4_K_M 22.49 GiB 18.49 GiB 2 GiB
Official page ↗
Q8_0 36.43 GiB 32.43 GiB 2 GiB

“—” means no information is recorded, not that the model lacks the capability.

Section four · two separate views

Inference and serving software

Data view

Software comparison

Compare software for local inference and model serving.

18 results out of 18 rows
Deployment need
Execution environment
Software role
Advanced filters 36 controls
Product
Local / cloud
Tasks and modalities
Request queue
Concurrency
Batching
Admission control
Model loading / unloading
Multiple models
Cold start
Prefix caching
Speculative decoding
CPU/GPU and KV offloading
Model sharding across GPUs
Independent replicas
Streaming
Structured output
Tool calling
Reasoning control
Chat / model template
Parser
Monitoring
Metrics
Health check
Authentication
Rate limiting
How the capability is provided
Maintenance status
Release date
Review date
All rows
Row comparison and conditions

Comparison rule: Versions can be displayed together. Interface comparisons keep the backend fixed; full-stack comparisons may vary it.

Use × beside a heading to hide its column. Open row details for sources and conditions.

Hidden columns:
Versioned inference and serving software comparison
CompareDetails
What is it for?
Practical advantage
Selection condition
Get started
Ollama · v0.34.0
Inference engine / library, API server, Model manager Local setup, development and small services Model downloads, management and API setup in one tool Local and cloud execution are separate; OLLAMA_NO_CLOUD=1 disables cloud paths. Getting started
vLLM · v0.29.0
Inference engine / library, API server Language model API with concurrent requests Continuous batching and KV-cache management; OpenAI-compatible API GGUF requires the separately installed vllm-gguf-plugin; support for this format is experimental. Getting started
SGLang · v0.5.20
Inference engine / library, API server Language and multimodal model serving on one GPU or a cluster Text generation, Continuous batching, Prefix caching, Structured output Tool parsers and quantization kernels depend on the model architecture. Starting with 0.5.20, CUDA 12 packages and images are no longer published; 0.5.19 is the last release for this path. Previously published images remain available. Responses storage is off by default. Retrieval, previous_response_id and background requests require --enable-response-store; this option cannot be enabled in prefill/decode-disaggregated (PD) deployments. Tool parsers and quantization kernels depend on the model architecture. Starting with 0.5.20, CUDA 12 packages and images are no longer published; 0.5.19 is the last release for this path. Previously published images remain available. Responses storage is off by default. Retrieval, previous_response_id and background requests require --enable-response-store; this option cannot be enabled in prefill/decode-disaggregated (PD) deployments. Getting started
SGLang · v0.5.19
Inference engine / library, API server Multi-user serving and workloads with shared prefixes Batching and prefix caching; workload-specific tuning Tool parsers and quantization kernels depend on model architecture; read each benchmark row's version and settings with its result. Getting started
llama.cpp / llama-server · v0.4.1
Inference engine / library, API server GGUF on CPU, GPU and combined RAM/VRAM Control over quantization, layer partitioning and CPU execution Partial CPU execution is distinct from AirLLM's layer-wise loading. Getting started
TensorRT-LLM · v1.2.1
Inference engine / library, API server Specialized deployment on NVIDIA GPUs Model, kernel and serving optimization Support paths and settings vary by release and architecture. Getting started
Triton Inference Server · v2.72.0
API server, Deployment manager Manage model services and inference pipelines Model management, scheduling and specialist backend integration Release 2.72.0 does not include a TensorRT-LLM backend container. Getting started
Text Embeddings Inference (TEI) · v1.9.3
Inference engine / library, API server Document and query embedding service Batching and an embedding API; CPU and GPU images Embedding support does not imply reranking support for every architecture. Getting started
AirLLM · v4.0.0
Inference engine / library Run large models with limited VRAM Load and unload layers through host memory and storage The main goal is lower GPU memory use; concurrent chat requires a benchmark of that workload. Getting started
Transformers · v5.17.0
Inference engine / library, API server Research, direct model control and custom paths Direct access to model and processing logic A simple Transformers path's speed does not represent every engine for that model. Getting started
LiteLLM · v1.101.0
Model gateway, API server Manage multiple providers and backends Shared API, routing and usage management Weight execution capability comes from the backend. Getting started
Open WebUI · v0.11.3
Chat interface Chat panel connected to model backends User chat interface and runtime connection Model memory and speed belong to the connected backend. Getting started
Text Generation Inference (TGI) · v3.3.7
Inference engine / library, API server Maintenance of text-generation serving for supported models Text generation, Model sharding across GPUs (conditional), Streaming output, Continuous batching The repository has been archived and read-only since March 21, 2026. The repository has been archived and read-only since March 21, 2026. Getting started
LM Studio · 0.4.24 Build 1
Chat interface, Model manager, API server Download models, chat and start a local API Model loading and unloading, Streaming output, Structured output, Concurrent requests Model capabilities and formats depend on the selected runtime. Model capabilities and formats depend on the selected runtime. Getting started
KTransformers · v0.7.1
Inference engine / library When host memory and GPU execution are part of the deployment design. Heterogeneous CPU/GPU execution, including optimized paths for supported MoE models. Performance depends on the supported model, CPU, memory and configuration. Getting started
Sentence Transformers · v6.0.1
Inference engine / library For implementing and evaluating the retrieval layer. Embeddings for semantic retrieval and CrossEncoder scoring or reranking. Library support does not establish task or language quality. Getting started
FlagEmbedding · v1.4.2
Inference engine / library BGE family and reranker execution BGE-M3 dense, sparse and multi-vector outputs The three BGE-M3 outputs use different indexing and consumption paths; dense output alone cannot replace all three. Getting started
MLX LM · v0.31.3
Inference engine / library Local generation and quantization of compatible models on Apple silicon Text generation (conditional), Streaming output (conditional), Prefix caching (conditional), Model sharding across GPUs (conditional) A direct route for local generation and quantization on Apple silicon with compatible models. Do not double-count unified memory as RAM plus VRAM. The README’s macOS 15 requirement concerns large-model memory wiring, not every MLX LM feature. A direct route for local generation and quantization on Apple silicon with compatible models. Do not double-count unified memory as RAM plus VRAM. The README’s macOS 15 requirement concerns large-model memory wiring, not every MLX LM feature. Getting started

“—” means no information is recorded, not that the model lacks the capability.

Evaluation results

Evidence for model selection

How selection and comparison work

On Spanish MIRACL, nDCG@10 rises from 51.2 for Small to 53.7 for Large-Instruct. On English, Large scores 52.9 versus 51.5 for Large-Instruct; Persian scores are 59.0 and 59.4. A multilingual average cannot replace the target-language comparison.

Two displayed numbers are not automatically rankable. Check the model variant, benchmark version, language, generation mode and protocol. A reported score difference is neither a quality ratio nor a statistical-significance claim.

MIRACL · nDCG@10· English

ModelScoreDifference from first rowSource and settings
48 points / 100
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=1

Dataset split
development
Score scale
0–100 as printed
51.2 points / 1003.2
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=2

Dataset split
development
Score scale
0–100 as printed
52.9 points / 1004.9
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=3

Dataset split
development
Score scale
0–100 as printed
51.5 points / 1003.5
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=4

Dataset split
development
Score scale
0–100 as printed

MIRACL · Recall@100· English

ModelScoreDifference from first rowSource and settings
85.3 points / 100
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=5

Dataset split
development
Score scale
0–100 as printed
86.4 points / 1001.1
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=6

Dataset split
development
Score scale
0–100 as printed
87.6 points / 1002.3
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=7

Dataset split
development
Score scale
0–100 as printed
88.2 points / 1002.9
Source and settingsMultilingual E5 technical report, arXiv:2402.05672v1 ↗Model publisher report

Appendix / detailed MIRACL results / language=en / column=8

Dataset split
development
Score scale
0–100 as printed

Quality in published evaluations

MATH-500 · pass@1 · English · %

Select for comparisonModelScore and scaleMode / subsetReporterSource
94.5 %0..100 as printedReasoningtestDeepSeekPublisher report
Conditions and source
MATH-500 · pass@1: 94.5%Publisher report · DeepSeek· Reasoning· en
Evaluation mode
Reasoning
Model named in the report
DeepSeek-R1-Distill-Llama-70B
Benchmark / version
MATH-500
Metric and unit
pass@1 · Percent
Source document commit
b1c0b44b4369b597ad119a196caf79a9c40e141e
Language
en
Maximum output tokens
32,768
Temperature
0.6
top-p
0.95
Samples per question
64
Dataset split
test
Score scale
0..100 as printed
Access date
2026-09-15 (Gregorian)
  • Publisher score; not suitable for ranking across different sources or directly inferring cost/performance.
  • Tested weight commit and precision are unreported; the result applies only to the named variant.
  • The 64-sample protocol estimates pass@1 for sampled evaluations; it has not been extended to CodeForces rating.
DeepSeek-R1-Distill-Llama-70B — publisher evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B/blob/b1c0b44b4369b597ad119a196caf79a9c40e141e/README.md ↗
Publisher / author
DeepSeek
Accessed
2026-09-15
Revision / commit
b1c0b44b4369b597ad119a196caf79a9c40e141e
Relevant source section
README.md: Distilled Model Evaluation table; Evaluation settings; matching lines 111, 145, 155, 160, 171, 222
  • DeepSeek-R1-Distill-Llama-70B; publisher report for the named variant; tested commit is unpublished.
  • The source document commit is not treated as the tested weight commit.
  • Not a Persian result; does not transfer to a quantized artifact.
MATH-500 dataset metadata · Technical documentation https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
Publisher / author
Hugging Face H4
Accessed
2026-09-17
Relevant source section
Dataset card: Languages: English; test split: 500 rows
  • Language and split of the original MATH-500 dataset.
DeepSeek-R1: distilled model evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
Publisher / author
DeepSeek
Accessed
2026-09-17
Relevant source section
Distilled Model Evaluation table; generation settings immediately preceding the table
  • Named distilled models in one publisher table; MATH-500 pass@1 uses the reported sampling protocol. No quantized-artifact or statistical-superiority claim.
83.9 %0..100 as printedReasoningtestDeepSeekPublisher report
Conditions and source
MATH-500 · pass@1: 83.9%Publisher report · DeepSeek· Reasoning· en
Evaluation mode
Reasoning
Model named in the report
DeepSeek-R1-Distill-Qwen-1.5B
Benchmark / version
MATH-500
Metric and unit
pass@1 · Percent
Source document commit
ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562
Language
en
Maximum output tokens
32,768
Temperature
0.6
top-p
0.95
Samples per question
64
Dataset split
test
Score scale
0..100 as printed
Access date
2026-09-15 (Gregorian)
  • Publisher score; not suitable for ranking across different sources or directly inferring cost/performance.
  • Tested weight commit and precision are unreported; the result applies only to the named variant.
  • The 64-sample protocol estimates pass@1 for sampled evaluations; it has not been extended to CodeForces rating.
DeepSeek-R1-Distill-Qwen-1.5B — publisher evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B/blob/ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562/README.md ↗
Publisher / author
DeepSeek
Accessed
2026-09-15
Revision / commit
ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562
Relevant source section
README.md: Distilled Model Evaluation table; Evaluation settings; matching lines 106, 145, 155, 160, 166, 220
  • DeepSeek-R1-Distill-Qwen-1.5B; publisher report for the named variant; tested commit is unpublished.
  • The source document commit is not treated as the tested weight commit.
  • Not a Persian result; does not transfer to a quantized artifact.
MATH-500 dataset metadata · Technical documentation https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
Publisher / author
Hugging Face H4
Accessed
2026-09-17
Relevant source section
Dataset card: Languages: English; test split: 500 rows
  • Language and split of the original MATH-500 dataset.
DeepSeek-R1: distilled model evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
Publisher / author
DeepSeek
Accessed
2026-09-17
Relevant source section
Distilled Model Evaluation table; generation settings immediately preceding the table
  • Named distilled models in one publisher table; MATH-500 pass@1 uses the reported sampling protocol. No quantized-artifact or statistical-superiority claim.
93.9 %0..100 as printedReasoningtestDeepSeekPublisher report
Conditions and source
MATH-500 · pass@1: 93.9%Publisher report · DeepSeek· Reasoning· en
Evaluation mode
Reasoning
Model named in the report
DeepSeek-R1-Distill-Qwen-14B
Benchmark / version
MATH-500
Metric and unit
pass@1 · Percent
Source document commit
1df8507178afcc1bef68cd8c393f61a886323761
Language
en
Maximum output tokens
32,768
Temperature
0.6
top-p
0.95
Samples per question
64
Dataset split
test
Score scale
0..100 as printed
Access date
2026-09-15 (Gregorian)
  • Publisher score; not suitable for ranking across different sources or directly inferring cost/performance.
  • Tested weight commit and precision are unreported; the result applies only to the named variant.
  • The 64-sample protocol estimates pass@1 for sampled evaluations; it has not been extended to CodeForces rating.
DeepSeek-R1-Distill-Qwen-14B — publisher evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B/blob/1df8507178afcc1bef68cd8c393f61a886323761/README.md ↗
Publisher / author
DeepSeek
Accessed
2026-09-15
Revision / commit
1df8507178afcc1bef68cd8c393f61a886323761
Relevant source section
README.md: Distilled Model Evaluation table; Evaluation settings; matching lines 109, 145, 155, 160, 168, 220
  • DeepSeek-R1-Distill-Qwen-14B; publisher report for the named variant; tested commit is unpublished.
  • The source document commit is not treated as the tested weight commit.
  • Not a Persian result; does not transfer to a quantized artifact.
MATH-500 dataset metadata · Technical documentation https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
Publisher / author
Hugging Face H4
Accessed
2026-09-17
Relevant source section
Dataset card: Languages: English; test split: 500 rows
  • Language and split of the original MATH-500 dataset.
DeepSeek-R1: distilled model evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
Publisher / author
DeepSeek
Accessed
2026-09-17
Relevant source section
Distilled Model Evaluation table; generation settings immediately preceding the table
  • Named distilled models in one publisher table; MATH-500 pass@1 uses the reported sampling protocol. No quantized-artifact or statistical-superiority claim.
94.3 %0..100 as printedReasoningtestDeepSeekPublisher report
Conditions and source
MATH-500 · pass@1: 94.3%Publisher report · DeepSeek· Reasoning· en
Evaluation mode
Reasoning
Model named in the report
DeepSeek-R1-Distill-Qwen-32B
Benchmark / version
MATH-500
Metric and unit
pass@1 · Percent
Source document commit
711ad2ea6aa40cfca18895e8aca02ab92df1a746
Language
en
Maximum output tokens
32,768
Temperature
0.6
top-p
0.95
Samples per question
64
Dataset split
test
Score scale
0..100 as printed
Access date
2026-09-15 (Gregorian)
  • Publisher score; not suitable for ranking across different sources or directly inferring cost/performance.
  • Tested weight commit and precision are unreported; the result applies only to the named variant.
  • The 64-sample protocol estimates pass@1 for sampled evaluations; it has not been extended to CodeForces rating.
DeepSeek-R1-Distill-Qwen-32B — publisher evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/blob/711ad2ea6aa40cfca18895e8aca02ab92df1a746/README.md ↗
Publisher / author
DeepSeek
Accessed
2026-09-15
Revision / commit
711ad2ea6aa40cfca18895e8aca02ab92df1a746
Relevant source section
README.md: Distilled Model Evaluation table; Evaluation settings; matching lines 58, 110, 145, 155, 160, 169, 196, 202, 220
  • DeepSeek-R1-Distill-Qwen-32B; publisher report for the named variant; tested commit is unpublished.
  • The source document commit is not treated as the tested weight commit.
  • Not a Persian result; does not transfer to a quantized artifact.
MATH-500 dataset metadata · Technical documentation https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
Publisher / author
Hugging Face H4
Accessed
2026-09-17
Relevant source section
Dataset card: Languages: English; test split: 500 rows
  • Language and split of the original MATH-500 dataset.
DeepSeek-R1: distilled model evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
Publisher / author
DeepSeek
Accessed
2026-09-17
Relevant source section
Distilled Model Evaluation table; generation settings immediately preceding the table
  • Named distilled models in one publisher table; MATH-500 pass@1 uses the reported sampling protocol. No quantized-artifact or statistical-superiority claim.
92.8 %0..100 as printedReasoningtestDeepSeekPublisher report
Conditions and source
MATH-500 · pass@1: 92.8%Publisher report · DeepSeek· Reasoning· en
Evaluation mode
Reasoning
Model named in the report
DeepSeek-R1-Distill-Qwen-7B
Benchmark / version
MATH-500
Metric and unit
pass@1 · Percent
Source document commit
916b56a44061fd5cd7d6a8fb632557ed4f724f60
Language
en
Maximum output tokens
32,768
Temperature
0.6
top-p
0.95
Samples per question
64
Dataset split
test
Score scale
0..100 as printed
Access date
2026-09-15 (Gregorian)
  • Publisher score; not suitable for ranking across different sources or directly inferring cost/performance.
  • Tested weight commit and precision are unreported; the result applies only to the named variant.
  • The 64-sample protocol estimates pass@1 for sampled evaluations; it has not been extended to CodeForces rating.
DeepSeek-R1-Distill-Qwen-7B — publisher evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B/blob/916b56a44061fd5cd7d6a8fb632557ed4f724f60/README.md ↗
Publisher / author
DeepSeek
Accessed
2026-09-15
Revision / commit
916b56a44061fd5cd7d6a8fb632557ed4f724f60
Relevant source section
README.md: Distilled Model Evaluation table; Evaluation settings; matching lines 107, 145, 155, 160, 167, 220
  • DeepSeek-R1-Distill-Qwen-7B; publisher report for the named variant; tested commit is unpublished.
  • The source document commit is not treated as the tested weight commit.
  • Not a Persian result; does not transfer to a quantized artifact.
MATH-500 dataset metadata · Technical documentation https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ↗
Publisher / author
Hugging Face H4
Accessed
2026-09-17
Relevant source section
Dataset card: Languages: English; test split: 500 rows
  • Language and split of the original MATH-500 dataset.
DeepSeek-R1: distilled model evaluation · Publisher report https://huggingface.co/deepseek-ai/DeepSeek-R1 ↗
Publisher / author
DeepSeek
Accessed
2026-09-17
Relevant source section
Distilled Model Evaluation table; generation settings immediately preceding the table
  • Named distilled models in one publisher table; MATH-500 pass@1 uses the reported sampling protocol. No quantized-artifact or statistical-superiority claim.
Data view

Small and specialized complementary models

Compare specialized roles, input limits, vector dimensions and results for the selected language.

Test shown in the result column: MIRACL · nDCG@10 · en

This selection applies to result columns across tables. Execution conditions are available in each model’s details.

Results grouped by test ←
28 results out of 28 rows
Model type
Application
Declared model size(billion)
Advanced filters 6 controls
Evidence type
All rows
Row comparison and conditions

Comparison rule: Metric and unit must describe the same task; document/s is never implicitly converted to token/s.

Use × beside a heading to hide its column. Open row details for sources and conditions.

Hidden columns:
Small and specialized models
CompareDetails
Exact task
Output type / dimensions
Languages and selected-language evidence
Licence
Result for the selected language
Download and setup
Official page ↗
Dense, sparse and multi-vector retrieval approximately 0.569 billion 8,192 tokens 1,024-dimensional dense vector + token weights + multi-vector output multilingual en: No recorded evidence MIT 3.81 GiB / one million vectors
Official page ↗
Query–document pair reranking approximately 0.568 billion 8,192 tokens Relevance score; optional sigmoid maps to 0–1 multilingual en: Published result Apache-2.0 Does not store fixed vectors
Official page ↗
Semantic retrieval and text embeddings approximately 0.6 billion 32,768 tokens Dense vector; 1,024 default / maximum dimensions multilingual en: Published result Apache-2.0 3.81 GiB / one million vectors
Official page ↗
Semantic retrieval and text embeddings approximately 4 billion 32,768 tokens Dense vector; 2,560 default / maximum dimensions multilingual en: No recorded evidence Apache-2.0 9.54 GiB / one million vectors
Official page ↗
Semantic retrieval and text embeddings approximately 8 billion 32,768 tokens Dense vector; 4,096 default / maximum dimensions multilingual en: No recorded evidence Apache-2.0 15.26 GiB / one million vectors
Official page ↗
Query–document pair reranking approximately 0.6 billion 32,768 tokens Text-pair relevance score; no embedding output multilingual en: Published result Apache-2.0 Does not store fixed vectors
Official page ↗
Query–document pair reranking approximately 4 billion 32,768 tokens Text-pair relevance score; no embedding output multilingual en: Published result Apache-2.0 Does not store fixed vectors
Official page ↗
Query–document pair reranking approximately 8 billion 32,768 tokens Text-pair relevance score; no embedding output multilingual en: Published result Apache-2.0 Does not store fixed vectors
Official page ↗
Text retrieval and similarity 0.118 billion 512 tokens 384-dimensional vector multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result MIT 1.43 GiB / one million vectors
Official page ↗
Text retrieval and similarity 0.023 billion 256 tokens 384-dimensional vector en en: Publisher declaration Apache-2.0 1.43 GiB / one million vectors
Official page ↗
Text retrieval and similarity 0.278 billion 512 tokens 768-dimensional vector multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result MIT 2.86 GiB / one million vectors
Official page ↗
Text retrieval and similarity 0.56 billion 512 tokens 1,024-dimensional vector multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result MIT 3.81 GiB / one million vectors
Official page ↗
Text retrieval and similarity 0.033 billion 512 tokens 384-dimensional vector en en: Publisher declaration MIT 1.43 GiB / one million vectors
Official page ↗
Search result reranking 0.278 billion 512 tokens Query–document relevance score en, zh en: Publisher declaration MIT Does not store fixed vectors
Official page ↗
Text retrieval and similarity approximately 0.3 billion nominal Total parameter count unverified 2,048 tokens 768-dimensional vector multilingual en: No recorded evidence gemma Commercial use subject to licence 2.86 GiB / one million vectors
Official page ↗
Text retrieval and similarity approximately 0.57 billion nominal Total parameter count unverified 8,192 tokens 1,024-dimensional vector multilingual, af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result CC-BY-NC-4.0 Commercial use subject to licence 3.81 GiB / one million vectors
Official page ↗
Search result reranking 0.278 billion 1,024 tokens Query–document relevance score multilingual en: No recorded evidence CC-BY-NC-4.0 Commercial use subject to licence Does not store fixed vectors
Official page ↗
Text retrieval and similarity 0.335 billion 512 tokens 1,024-dimensional vector en en: Publisher declaration Apache-2.0 3.81 GiB / one million vectors
Official page ↗
Text retrieval and similarity 0.137 billion 8,192 tokens 768-dimensional vector en en: Publisher declaration Apache-2.0 2.86 GiB / one million vectors
Official page ↗
Specialization for classification and named entity recognition 0.15 billion 8,192 tokens Token representations; task head after training en en: Publisher declaration Apache-2.0 Does not store fixed vectors
Official page ↗
Text retrieval and similarity 0.56 billion 512 tokens Dense vector af, am, ar, as, az, be, bg, bn, br, bs, ca, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, hu, hy, id, is, it, ja, jv, ka, kk, km, kn, ko, ku, ky, la, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, om, or, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, so, sq, sr, su, sv, sw, ta, te, th, tl, tr, ug, uk, ur, uz, vi, xh, yi, zh en: Published result mit 3.81 GiB / one million vectors
SAFETENSORS · Official ↗
Official page ↗
Text retrieval and similarity 0.123 billion 512 tokens Dense vector fa en: No recorded evidence 2.86 GiB / one million vectors
SAFETENSORS · Official ↗
Official page ↗
Text retrieval and similarity 0.353 billion 512 tokens Dense vector fa en: No recorded evidence 3.81 GiB / one million vectors
SAFETENSORS · Official ↗
Official page ↗
Base for training classification and entity recognition 512 tokens Token representations; the task head requires training fa en: No recorded evidence Does not store fixed vectors
PYTORCH · Official ↗
Official page ↗
Text, image and video retrieval 2 billion 32,768 tokens Vector with up to 2048 dimensions en: No recorded evidence Apache-2.0 7.63 GiB / one million vectors
SAFETENSORS · Official ↗
Official page ↗
Text, image and video retrieval 8 billion 32,768 tokens Vector with up to 4096 dimensions en: No recorded evidence Apache-2.0 15.26 GiB / one million vectors
SAFETENSORS · Official ↗
Official page ↗
Reranker 2 billion 32,768 tokens en: No recorded evidence Apache-2.0 Does not store fixed vectors
SAFETENSORS · Official ↗
Official page ↗
Reranker 8 billion 32,768 tokens en: No recorded evidence Apache-2.0 Does not store fixed vectors
SAFETENSORS · Official ↗

“—” means no information is recorded, not that the model lacks the capability.

Guide articles

Model selection, memory, runtime software and quality evaluation.

Back to tables ↑
17 min

RAG, CAG, KAG, fine-tuning and instruction tuning: how do they differ, and which should you choose?

When a model gives an unsatisfactory answer, does it need more training, or simply access to the right information? Through practical analogies, this article compares RAG, CAG and KAG with fine-tuning and instruction tuning: retrieving documents, caching knowledge, reasoning over relationships and changing model behavior. A comparison table and real project scenarios help distinguish missing knowledge from unsuitable behavior before committing to training, and identify the method or combination that fits the task.

Read the standalone article ↗
13 min

Ollama, vLLM, SGLang or llama.cpp: choosing an inference engine

How scheduling, KV cache, model formats and operational controls change the choice of inference engine, with a worked memory example and a reproducible comparison method.

Read the standalone article ↗
14 min

Why the fastest GPU does not necessarily deliver the fastest response

More compute and a higher token rate do not always mean a faster response. This article examines time to first token versus completion time, memory and concurrency, the division of work between GPUs and LPUs, and communication costs—so infrastructure choices reflect the capacity to serve requests at the required quality and latency.

Read the standalone article ↗