Deployment path chooser
1. Must data stay offline?
Yes: local desktop or your own server. Download weights only from an official publisher or verified registry.
No: a managed API can reduce GPU work, but read its retention and region terms.
2. What hardware exists?
CPU / 16–32 GB RAM: target a small, quantized model through llama.cpp or Ollama.
One GPU: select a checkpoint that leaves real KV-cache and runtime headroom.
3. What must it do?
Code / agents: evaluate exact tool-call formats on your held-out tasks.
Vision / audio: verify the exact checkpoint and runtime accept the modality.
4. Can you meet the terms?
Commercial product: have counsel or procurement read the exact licence and acceptable-use policy.
Research only: still track attribution, redistribution, and jurisdiction constraints.
Quick reference: current released-weight families
“Weights available” does not imply an OSI-approved open-source licence. “Open model” and “open weight” are deployment facts; licence rights come from the exact release terms.
| Family / representative class | Weights | Licence / terms to read | Good first runtime | Hardware signal | Starting fit | Primary source |
|---|---|---|---|---|---|---|
| OpenAI gpt-oss-20b | Released | Apache 2.0 + usage policy | Ollama / llama.cpp / vLLM | Official claim: 16 GB memory | Local text reasoning / tool-use evaluation | OpenAI release |
| OpenAI gpt-oss-120b | Released | Apache 2.0 + usage policy | vLLM / hosted or self-managed GPU | Official claim: one 80 GB GPU | High-capacity controlled deployment | OpenAI model card |
| Qwen3.8-Max | Released | Apache 2.0 for listed checkpoint | vLLM / server cluster | 2.4T total; 95B active | Frontier agentic, multimodal, and long-context evaluation | Qwen release |
| Llama 4 Scout | Released | Llama 4 Community License | vLLM / managed inference | Multi-GPU class | Long-context, multimodal evaluation | Meta model page |
| Llama 3.3 70B | Released | Llama 3.3 Community License | llama.cpp / vLLM | ~48 GB at 4-bit + overhead | Large local text deployment | Meta downloads |
| DeepSeek V4.1 Flash / V4 Pro | Released | Read exact repository terms | vLLM / managed API endpoint | Flash: multi-GPU; Pro: enterprise cluster | Multimodal reasoning and agentic evaluation | DeepSeek models |
| Mistral Small 4 | Released | Apache 2.0 | vLLM / Ollama | 119B total; 6.5B active | Hybrid instruct/reasoning/coding | Mistral model card |
| Mistral Large 3 | Released | Open weights; read model card | vLLM / server cluster | Large-server class | General-purpose multimodal evaluation | Mistral models |
| Gemma 4 E4B / 12B / 26B A4B / 31B | Released | Gemma 4 licence | Gemma / Ollama / vLLM | Mobile through server classes | Multimodal deployment by footprint | Google guide |
| OLMo 2 32B | Released | Apache 2.0 | vLLM / transformers | ~24 GB+ at 4-bit, plus overhead | Open research and provenance | Ai2 OLMo |
Hardware figures are planning signals, not vendor capacity guarantees. Weights, quantization, architecture, context length, batch size, runtime, and offload materially change memory needs.
September 2026 update: DeepSeek V4.1-Flash released Sept 10, 2026 with native multimodal vision. Qwen3.8-Max (Aug 3, 2026) remains Alibaba flagship. Mistral Small 4, Large 3, Gemma 4, and Llama 4 Scout remain current across their respective deployment profiles.
The terms people blur together
Access and control
- Open weight: weights can be obtained under stated terms. It says nothing, by itself, about source code or licence freedoms.
- Open source: a licence claim about software/source; use it only where the licence supports it. Training data, weights, and code can have different terms.
- Downloadable: files are available somewhere. Verify publisher identity, revision, checksum, format, and redistribution permission.
- Local: inference runs on a device you control. A UI, updater, or retrieval connector can still expose data.
- Self-hosted: you operate serving, identity, logs, scaling, and patching—often on a cloud GPU.
- Managed open-model API: a third party operates an open-weight model. It reduces operations, not necessarily retention obligations.
Capacity vocabulary
- Parameter count: learned values; useful for storage planning, poor as a standalone quality score.
- Active parameters: computation used per token in an MoE model; it does not erase total-weight storage needs.
- Quantization: lower-precision weights such as 4-bit; it cuts memory but can change accuracy and tool reliability.
- Context window: maximum tokens in a request. A claimed maximum is not a quality promise at that length.
- KV cache: memory retaining attention state; it grows with sequence length and concurrent requests.
- Model card: upstream scope, limitations, licence, and intended-use record; minimum procurement reading.
Choose the operating shape before the model
| Shape | Privacy boundary | Operational burden | Update control | Cost visibility | Use when |
|---|---|---|---|---|---|
| Local desktop | Strongest if companion services are disabled | Low for one user; hardware-bound | You choose every file | Hardware + power | Private experiments, offline work |
| Self-hosted server | Your cloud/VPC and configured logs | High: auth, GPU, observability, patching | You pin versions and roll back | GPU-hours + operations | Stable internal workload with operators |
| Managed open-model API | Provider contract and regions | Low | Provider catalog and aliases | Per-token / request | Demand discovery before owning serving |
| Closed-model API | Provider contract and regions | Low | Limited; provider may change models | Per-token / request | Capability/support beats control need |
Runtime decision matrix
Ollama
Fast local developer workflow and simple local API. Good for a single-user proof of concept.
Poor default for production Pin artifacts, authenticate callers, and measure capacity.
Official docsllama.cpp
Portable inference and GGUF quantized models; useful for CPU, Apple silicon, and constrained devices.
Watch conversion provenance and feature parity.
Official repositoryvLLM
Throughput-oriented GPU serving with an OpenAI-compatible server. Start here for a team service after local evaluation passes.
Watch GPU compatibility, quota, and tenant isolation.
Official docsManaged route
Provider-operated endpoint for a named model. Good for operations-light demand discovery.
Watch retention, regional processing, aliases, and logs.
Compare API operating costsWorked path: one-GPU local coding assistant
- Constraint: one consumer GPU, proprietary code, code explanation plus structured JSON—not an autonomous production agent.
- Shortlist: select two released medium-class code-capable candidates whose exact licence permits the work. Use a publisher or verified registry revision.
- Budget memory: leave room beyond weights for the runtime, KV cache, and OS. A model that fits at idle can fail on a long prompt or two requests.
- Test: 30 held-out repository questions, 10 JSON-schema outputs, and 10 tool-call fixtures. Log revision, quantization, runtime, GPU, prompt/output tokens, latency, and failures.
- Gate: deploy only if it follows schema and repository-access policy at required latency; otherwise reduce scope, change model/runtime, or use hosted inference.
record: publisher/checkpoint@immutable-revision · artifact checksum · runtime version · context tokens · concurrencyLicence, distribution, and evaluation
Example: an internal support assistant may be commercial use even if it never directly charges customers. Before distribution, record the exact checkpoint URL/revision, licence version, acceptable-use terms, attribution/notice obligations, redistribution rules, geography restrictions, and whether adapters or quantizations create a derivative-distribution question. Escalate ambiguity; a model-hub tag is not permission.
- Keep the upstream licence and model card with the release record.
- Use immutable revisions/checksums; scan model files and containers through normal supply-chain controls.
- Document retention, telemetry, prompt logs, and who can download weights.
- Run hold-out prompts, structured-output fixtures, tool-call tests, latency/memory tests, and failure logging before replacing any workflow.
Advanced constraints and common mistakes
Quantization, batching, and KV cache
Quantization cuts memory and may change quality. Batching can improve aggregate throughput but add queue time. Long contexts and concurrent requests increase KV-cache pressure; test the longest expected prompt at expected concurrency before sizing hardware.
Multimodal and tool calling
Check the exact model card and runtime API: a family can contain text-only and multimodal variants, and runtimes expose different subsets. Treat tool calls as untrusted model suggestions; validate arguments and authorize actions outside the model.
Supply chain and data boundary
Weights, quantizations, packages, containers, web UIs, and retrieval connectors expand the trust boundary. Verify publisher/revision, restrict egress where appropriate, scan artifacts, keep secrets out of prompts, and assume prompt injection can reach connected tools.
Anti-patterns to reject
- Calling every downloadable model “open source.”
- Choosing from an average leaderboard while ignoring task, harness, and quantization.
- Ignoring licence conditions because a demo works.
- Selecting solely by parameter count or active parameters.
- Running untrusted model packages with production data.
- Treating a one-user demo as proof of production capacity.
Continue the decision
Frontier AI labs · Ubuntu for AI developers · AI model API pricing · AI coding agents compared.
Sources: OpenAI gpt-oss · Qwen3.8-Max · Gemma 4 · Mistral models · Llama 4.