deployment notebook

Open-weight AI models: run, license, and choose

A downloadable model is not automatically open source, private, affordable to run, or licensed for your product. Start with the deployment constraint; choose a family only after that.

Releases, licences, and endpoints are volatile. Each family row links to the publisher or its official model card.

Deployment path chooser

1. Must data stay offline?

Yes: local desktop or your own server. Download weights only from an official publisher or verified registry.

No: a managed API can reduce GPU work, but read its retention and region terms.

2. What hardware exists?

CPU / 16–32 GB RAM: target a small, quantized model through llama.cpp or Ollama.

One GPU: select a checkpoint that leaves real KV-cache and runtime headroom.

3. What must it do?

Code / agents: evaluate exact tool-call formats on your held-out tasks.

Vision / audio: verify the exact checkpoint and runtime accept the modality.

4. Can you meet the terms?

Commercial product: have counsel or procurement read the exact licence and acceptable-use policy.

Research only: still track attribution, redistribution, and jurisdiction constraints.

Decision output: offline + one consumer GPU + coding normally means “test two medium, quantized candidates locally.” It does not justify choosing the largest model that barely loads.

Quick reference: current released-weight families

“Weights available” does not imply an OSI-approved open-source licence. “Open model” and “open weight” are deployment facts; licence rights come from the exact release terms.

Family / representative classWeightsLicence / terms to readGood first runtimeHardware signalStarting fitPrimary source
OpenAI gpt-oss-20bReleasedApache 2.0 + usage policyOllama / llama.cpp / vLLMOfficial claim: 16 GB memoryLocal text reasoning / tool-use evaluationOpenAI release
OpenAI gpt-oss-120bReleasedApache 2.0 + usage policyvLLM / hosted or self-managed GPUOfficial claim: one 80 GB GPUHigh-capacity controlled deploymentOpenAI model card
Qwen3.8-MaxReleasedApache 2.0 for listed checkpointvLLM / server cluster2.4T total; 95B activeFrontier agentic, multimodal, and long-context evaluationQwen release
Llama 4 ScoutReleasedLlama 4 Community LicensevLLM / managed inferenceMulti-GPU classLong-context, multimodal evaluationMeta model page
Llama 3.3 70BReleasedLlama 3.3 Community Licensellama.cpp / vLLM~48 GB at 4-bit + overheadLarge local text deploymentMeta downloads
DeepSeek V4.1 Flash / V4 ProReleasedRead exact repository termsvLLM / managed API endpointFlash: multi-GPU; Pro: enterprise clusterMultimodal reasoning and agentic evaluationDeepSeek models
Mistral Small 4ReleasedApache 2.0vLLM / Ollama119B total; 6.5B activeHybrid instruct/reasoning/codingMistral model card
Mistral Large 3ReleasedOpen weights; read model cardvLLM / server clusterLarge-server classGeneral-purpose multimodal evaluationMistral models
Gemma 4 E4B / 12B / 26B A4B / 31BReleasedGemma 4 licenceGemma / Ollama / vLLMMobile through server classesMultimodal deployment by footprintGoogle guide
OLMo 2 32BReleasedApache 2.0vLLM / transformers~24 GB+ at 4-bit, plus overheadOpen research and provenanceAi2 OLMo

Hardware figures are planning signals, not vendor capacity guarantees. Weights, quantization, architecture, context length, batch size, runtime, and offload materially change memory needs.

September 2026 update: DeepSeek V4.1-Flash released Sept 10, 2026 with native multimodal vision. Qwen3.8-Max (Aug 3, 2026) remains Alibaba flagship. Mistral Small 4, Large 3, Gemma 4, and Llama 4 Scout remain current across their respective deployment profiles.

The terms people blur together

Access and control

  • Open weight: weights can be obtained under stated terms. It says nothing, by itself, about source code or licence freedoms.
  • Open source: a licence claim about software/source; use it only where the licence supports it. Training data, weights, and code can have different terms.
  • Downloadable: files are available somewhere. Verify publisher identity, revision, checksum, format, and redistribution permission.
  • Local: inference runs on a device you control. A UI, updater, or retrieval connector can still expose data.
  • Self-hosted: you operate serving, identity, logs, scaling, and patching—often on a cloud GPU.
  • Managed open-model API: a third party operates an open-weight model. It reduces operations, not necessarily retention obligations.

Capacity vocabulary

  • Parameter count: learned values; useful for storage planning, poor as a standalone quality score.
  • Active parameters: computation used per token in an MoE model; it does not erase total-weight storage needs.
  • Quantization: lower-precision weights such as 4-bit; it cuts memory but can change accuracy and tool reliability.
  • Context window: maximum tokens in a request. A claimed maximum is not a quality promise at that length.
  • KV cache: memory retaining attention state; it grows with sequence length and concurrent requests.
  • Model card: upstream scope, limitations, licence, and intended-use record; minimum procurement reading.

Choose the operating shape before the model

ShapePrivacy boundaryOperational burdenUpdate controlCost visibilityUse when
Local desktopStrongest if companion services are disabledLow for one user; hardware-boundYou choose every fileHardware + powerPrivate experiments, offline work
Self-hosted serverYour cloud/VPC and configured logsHigh: auth, GPU, observability, patchingYou pin versions and roll backGPU-hours + operationsStable internal workload with operators
Managed open-model APIProvider contract and regionsLowProvider catalog and aliasesPer-token / requestDemand discovery before owning serving
Closed-model APIProvider contract and regionsLowLimited; provider may change modelsPer-token / requestCapability/support beats control need

Runtime decision matrix

Ollama

Fast local developer workflow and simple local API. Good for a single-user proof of concept.

Poor default for production Pin artifacts, authenticate callers, and measure capacity.

Official docs

llama.cpp

Portable inference and GGUF quantized models; useful for CPU, Apple silicon, and constrained devices.

Watch conversion provenance and feature parity.

Official repository

vLLM

Throughput-oriented GPU serving with an OpenAI-compatible server. Start here for a team service after local evaluation passes.

Watch GPU compatibility, quota, and tenant isolation.

Official docs

Managed route

Provider-operated endpoint for a named model. Good for operations-light demand discovery.

Watch retention, regional processing, aliases, and logs.

Compare API operating costs

Worked path: one-GPU local coding assistant

  1. Constraint: one consumer GPU, proprietary code, code explanation plus structured JSON—not an autonomous production agent.
  2. Shortlist: select two released medium-class code-capable candidates whose exact licence permits the work. Use a publisher or verified registry revision.
  3. Budget memory: leave room beyond weights for the runtime, KV cache, and OS. A model that fits at idle can fail on a long prompt or two requests.
  4. Test: 30 held-out repository questions, 10 JSON-schema outputs, and 10 tool-call fixtures. Log revision, quantization, runtime, GPU, prompt/output tokens, latency, and failures.
  5. Gate: deploy only if it follows schema and repository-access policy at required latency; otherwise reduce scope, change model/runtime, or use hosted inference.
record: publisher/checkpoint@immutable-revision · artifact checksum · runtime version · context tokens · concurrency

Licence, distribution, and evaluation

Example: an internal support assistant may be commercial use even if it never directly charges customers. Before distribution, record the exact checkpoint URL/revision, licence version, acceptable-use terms, attribution/notice obligations, redistribution rules, geography restrictions, and whether adapters or quantizations create a derivative-distribution question. Escalate ambiguity; a model-hub tag is not permission.

  • Keep the upstream licence and model card with the release record.
  • Use immutable revisions/checksums; scan model files and containers through normal supply-chain controls.
  • Document retention, telemetry, prompt logs, and who can download weights.
  • Run hold-out prompts, structured-output fixtures, tool-call tests, latency/memory tests, and failure logging before replacing any workflow.

Advanced constraints and common mistakes

Quantization, batching, and KV cache

Quantization cuts memory and may change quality. Batching can improve aggregate throughput but add queue time. Long contexts and concurrent requests increase KV-cache pressure; test the longest expected prompt at expected concurrency before sizing hardware.

Multimodal and tool calling

Check the exact model card and runtime API: a family can contain text-only and multimodal variants, and runtimes expose different subsets. Treat tool calls as untrusted model suggestions; validate arguments and authorize actions outside the model.

Supply chain and data boundary

Weights, quantizations, packages, containers, web UIs, and retrieval connectors expand the trust boundary. Verify publisher/revision, restrict egress where appropriate, scan artifacts, keep secrets out of prompts, and assume prompt injection can reach connected tools.

Anti-patterns to reject

  • Calling every downloadable model “open source.”
  • Choosing from an average leaderboard while ignoring task, harness, and quantization.
  • Ignoring licence conditions because a demo works.
  • Selecting solely by parameter count or active parameters.
  • Running untrusted model packages with production data.
  • Treating a one-user demo as proof of production capacity.

Continue the decision

Frontier AI labs · Ubuntu for AI developers · AI model API pricing · AI coding agents compared.

Sources: OpenAI gpt-oss · Qwen3.8-Max · Gemma 4 · Mistral models · Llama 4.

David Veksler is a Principal AI Engineer in Denver. He leads agentic AI engineering at Antech, a Mars company, and builds AI platforms for regulated financial firms. This page was produced by a governed, multi-agent Claude Code pipeline with a git audit trail. How it's built Case studies