Skip to content
Blog
Local LLMOllamaGemma 4Quantization+1

Run Gemma 4 with Ollama: Model Tags and Memory Requirements

Run Gemma 4 locally with Ollama. Choose a model tag, budget memory for context, verify GPU placement with ollama ps, and test the native API.

16 min read
On this page

You pull a model, ask one question, and get an answer. Then you give it a repository file and the response slows to a crawl. Nothing is obviously broken: the weights fit, but context memory and runtime buffers changed the plan. A local model that answers isn't automatically a model that fits your workload.

This walkthrough chooses an artifact, checks runtime placement at an explicit context length, and saves a timing receipt. The installation and inference commands are documentation examples, not measurements from a local GPU. The Python timing example uses synthetic data.

Ollama runs a local model server and exposes an HTTP API. It binds to 127.0.0.1:11434 by default, and that endpoint doesn't require authentication.[1][2]

That endpoint is local, but local isn't a complete access-control plan. Keep the bind address local unless you have deliberately added network controls, authentication, and TLS.

Gemma 4 is Google DeepMind's Apache 2.0 open-weights family. It has five sizes: E2B, E4B, 12B, 26B A4B, and 31B.[3]

E2B and E4B use per-layer , so their total stored parameter counts exceed their effective parameter counts. The model card gives E2B and E4B 128K context limits, and 12B, 26B A4B, and 31B 256K limits.[3]

Those are model ceilings. They aren't promises that your Ollama process can allocate the full window while keeping the model on fast memory. Start with that distinction in mind, because it explains most surprising local runs.

Google introduced the unified 12B model on June 3, 2026 for machines with 16 GB of or unified memory.[4]

Its Q4_K_M Ollama artifact is 7.6 GB. That makes it a candidate for this hardware class, not a fit guarantee. If runtime memory is too tight, compare gemma4:e4b-it-qat.[5]

Ollama's published local tags currently advertise text and image input. Google's lists native audio for E2B, E4B, and 12B, so treat local audio as unsupported until the exact Ollama tag advertises it.[3][5]

In the tag names below, qat means . The model sees simulated low-precision behavior during training before its final weights are packed. That can change fit, but it isn't a substitute for measuring task quality.

Choose an exact Ollama tag

An Ollama tag names the variant and packaged artifact. It can also select or a runtime path. Before you compare quality, choose the exact artifact you intend to measure.

These published sizes were checked September 21, 2026. GB is the registry's rounded artifact size, not a measured RAM or VRAM requirement.[5]

TagPublished artifactModel limitStart here when
gemma4:e2b-it-qat4.3 GB128KTightest local memory budget
gemma4:e4b-it-qat6.1 GB128KE2B misses quality and 12B won't leave enough headroom
gemma4:12b-it-q4_K_M7.6 GB256K12B trial with room for cache and runtime
gemma4:12b-it-qat7.2 GB256KCompare QAT quality and runtime with the 12B baseline
gemma4:26b-a4b-it-qat16 GB256KWorkstation trial with much more free memory
gemma4:31b-it-qat19 GB256KDense 31B comparison with clear memory margin

gemma4:cloud and gemma4:31b-cloud call cloud models. A localhost API URL can still route to those models, so the URL alone doesn't prove local inference.[5][2] For a local-only setup, disable cloud features in Ollama's configuration or set OLLAMA_NO_CLOUD=1 in the server environment and restart it. Model downloads still require network access.[1]

The 26B A4B model is a model with 25.2B total parameters and 3.8B active parameters per . Ollama rounds that active path to 4B.[3][6]

The active count describes compute per token, while the total artifact still determines the weight-memory problem. This isn't a dense 4B model in disguise.

Diagram showing Choose candidate tag, Memory and latency pass?, no, and Reduce context or model size.
Choose candidate tag, Memory and latency pass?, no, and Reduce context or model size.

💡 Fit rule: Keep the smallest exact tag that passes your task evaluation at the context length you actually use.

Avoid short aliases in benchmark scripts. latest currently maps to E4B (9.6 GB), while 26b now maps to the 19 GB 26b-a4b-it-mtp-q4_K_M artifact. Neither name describes the full runtime choice.[5] Even a descriptive tag can be updated: record the full digest returned by /api/tags, and retain the artifact if you need an exact .[7]

Budget runtime memory

Published artifact size isn't the full runtime requirement. The process also needs memory for metadata, runtime buffers, prompt processing, the , and a safety margin.

The KV cache stores attention state from earlier tokens, so longer context and concurrent requests raise memory use. Inference: TTFT, TPS & KV Cache is the place to size that cache instead of guessing from download size.

Keep units and memory pools consistent. A 7.6 GB file is approximately 7.08 GiB because a GiB is 2³⁰ bytes. That conversion still doesn't tell you its resident allocation. Discrete GPU VRAM is separate from host RAM; unified memory is shared with the OS and other applications.

For an illustrative 16 GiB unified-memory budget, reserve 4 GiB for the OS and applications. Suppose measured resident weights use 7 GiB and runtime buffers use 1 GiB. That leaves 4 GiB for cache and any additional margin. If each independent request needs 4 GiB of cache, one fills the budget and two require 20 GiB. These are accounting assumptions, not Gemma 4 measurements; a full budget with no safety margin is already fragile.

Illustrative unified-memory budget: 4 GiB OS and applications, 7 GiB resident weights, and 1 GiB runtime leave 4 GiB out of 16 for cache; doubling independent request caches raises demand to 20 GiB
Same weights, different request capacity. The assumed cache doubles with two independent requests and no sharing. Actual cache allocation depends on the model, context, cache precision, and runtime configuration.

The final fit check happens after a real request. On a discrete , partial can hurt latency.

On Apple Silicon, CPU and GPU share unified memory, so read ollama ps placement labels next to measured throughput. Benchmark the exact machine, tag, context, and workload. Local LLM Deployment walks through the full weights-plus-cache-plus-reserve budget, including when to roll back a tag.

What QAT and MTP change

Once the model fits, a second question appears: can it generate fast enough? QAT and solve different problems.

QAT simulates low-precision behavior during training so the model can adapt before weights are compressed. Google's official QAT release covers Q4_0 checkpoints and mobile-specific formats.[8]

Ollama's default aliases and QAT tags use different quantization formats, so the published size gap mixes QAT with packaging. The gap is large for E2B (7.2 GB to 4.3 GB) and E4B (9.6 GB to 6.1 GB), but small for 12B (7.6 GB to 7.2 GB). Treat each tag as its own artifact instead of assuming one percentage.[5] Model Quantization: GPTQ, AWQ & GGUF explains how QAT differs from and why GGUF is a container rather than a bit width.

Published Ollama Gemma 4 artifact sizes: E2B Q4_K_M 7.2 GB versus QAT 4.3 GB, E4B 9.6 versus 6.1 GB, and 12B 7.6 versus 7.2 GB
Published artifact sizes, not measured peak memory or a controlled QAT ablation. The 12B difference is 0.4 GB; test quality and placement instead of assuming the E2B reduction carries over.

MTP is a speculative-decoding path. A small drafter proposes tokens and the target verifies them in a batch. Accepted drafts can raise generated-token throughput when saved target passes outweigh drafting and verification overhead. MTP doesn't shrink target weights or remove prompt-processing and KV-cache cost.[9]

Diagram showing Draft candidate tokens, Target verifies batch, Accept valid prefix, and rejection.
Draft candidate tokens, Target verifies batch, Accept valid prefix, and rejection.

Ollama's documented MTP path is runtime-specific. Ollama 0.31 added automatic MTP for Gemma 4 on Apple Silicon MLX. The published recipe is to update Ollama and re-pull gemma4:12b-mlx; no MTP flag is required.[10]

The separate gemma4:31b-coding-mtp-bf16 artifact is 64 GB. The current 26B MTP Q4_K_M artifact is 19 GB. A tag containing mtp identifies packaging, not a measured speedup on your machine.[5] Check the supported runtime and compare generation rates with prompt, context, thinking mode, and sampling held constant. Speculative Decoding explains acceptance and correction in more detail.

Install from the official package

Install Ollama from the official download page on macOS or Windows. On Linux, the official guide offers an install script and manual packages.[11][12]

The commands below download the script for inspection before execution. sh -n checks shell syntax, not trustworthiness. Running the final installation command changes the machine; use your organization's approved package process where required.

terminal
curl --fail --location --proto '=https' --tlsv1.2 \
  https://ollama.com/install.sh --output ollama-install.sh
sh -n ollama-install.sh
less ollama-install.sh
sh ollama-install.sh
ollama --version

Pull one descriptive local tag. Check free disk first, because the pull stores several gigabytes. The walkthrough uses the Q4_K_M 12B artifact throughout; substitute a smaller tag consistently if necessary.

terminal
ollama pull gemma4:12b-it-q4_K_M

Check the server and save the artifact identity before generation. This Bash snippet requires jq. A failed curl means the server or HTTP request failed; a successful request followed by a failed selection means the expected tag was not found.[7]

verify-ollama.sh
set -euo pipefail

curl --fail --silent --show-error http://127.0.0.1:11434/api/tags \
  --output ollama-tags.json
jq -e '.models[] | select(.name == "gemma4:12b-it-q4_K_M")
       | {name, digest, size, details}' ollama-tags.json

On a Linux manual install, start ollama serve in another terminal if no service is running. Desktop apps normally manage the server. Keep the bind address at 127.0.0.1 unless the deployment has a deliberate authentication, firewall, and TLS plan.[1][2]

Call the native API

The native /api/generate endpoint returns the response and runtime timing fields in one receipt. It also accepts num_ctx, so it's a useful first benchmark surface.[13][14]

Set context explicitly. Ollama's context guide describes hardware-dependent defaults, while the FAQ gives 4096 as the baseline; neither is evidence of your running allocation. The example requests 8192 tokens, far below the model's advertised ceiling.[15][1]

Disable streaming so one JSON object contains text and timings. Disable thinking for this short baseline so the output cap isn't consumed by a separate thinking response. Temperature zero reduces sampling variation but doesn't guarantee identical runs across runtimes or hardware.[13]

benchmark-gemma4.sh
set -euo pipefail

curl --fail --silent --show-error \
  http://127.0.0.1:11434/api/generate \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gemma4:12b-it-q4_K_M",
    "prompt": "List three checks for a failing unit test.",
    "stream": false,
    "think": false,
    "keep_alive": "5m",
    "options": {
      "num_ctx": 8192,
      "num_predict": 128,
      "temperature": 0
    }
  }' > gemma4-12b-run.json

jq -e '.done == true and (.response | length > 0)' gemma4-12b-run.json
jq '{model, done, done_reason, prompt_eval_count, prompt_eval_cached_count, eval_count, total_duration,
     load_duration, prompt_eval_duration, eval_duration}' gemma4-12b-run.json
ollama ps

ollama ps shows allocated CONTEXT and PROCESSOR placement. Check both after generation. A CPU/GPU split is not a full-GPU fit; on unified memory also inspect system memory pressure.[1][15]

Ollama reports durations in nanoseconds. done: true means generation ended, not that the answer was correct or complete. Inspect done_reason: a length limit can stop an answer halfway through.[14][13]

Generated-token rate is eval_count / (eval_duration / 1e9). Prompt timing needs another step: Ollama reports the total input count as prompt_eval_count, cached input as prompt_eval_cached_count, and time spent evaluating uncached tokens as prompt_eval_duration. Subtract the cached count before calculating uncached prompt tokens per second.[13]

Suppose a receipt has 1,000 input tokens, 800 cached tokens, and 0.2 seconds of prompt evaluation. The measured uncached rate is (1000−800)/0.2=1000(1000 - 800) / 0.2 = 1000 tokens/second. Dividing the full input count by that duration would report 5,000 tokens/second and mistake cache reuse for faster computation. With no uncached input, report the uncached rate as unavailable rather than dividing by zero. Keep prompt rate and decode rate separate: a model can ingest a long file slowly and still generate quickly afterward.

The evaluator below checks positive generated counts and timings, a valid cached-token count, and a nonempty response that stopped with done_reason: "stop". It checks an illustrative of at least 20 generated tokens/second and at most 5 seconds of server time. A normal stop doesn't prove that an answer fulfilled the request: a confident but wrong answer can stop normally too. These thresholds and stop checks are separate from task quality.

The synthetic has 128 generated tokens in 3.2 seconds: 40 tokens/second. Its server total is 3.42 seconds. Token counts are invented for arithmetic and aren't a tokenization of the placeholder response. Save this code as evaluate_ollama_run.py; with no arguments it evaluates that fixture. Use --receipt to evaluate your saved JSON file.

evaluate_ollama_run.py
import json
import sys
from pathlib import Path

def parse_ollama_metrics(payload: dict) -> dict[str, float | str | bool | None]:
    if not isinstance(payload, dict) or payload.get("done") is not True:
        raise ValueError("Expected a completed non-streaming response")
    for field in ("eval_count", "eval_duration", "prompt_eval_count", "total_duration"):
        if type(payload.get(field)) is not int or payload[field] <= 0:
            raise ValueError(f"Expected a positive integer: {field}")
    cached = payload.get("prompt_eval_cached_count", 0)
    if type(cached) is not int or not 0 <= cached <= payload["prompt_eval_count"]:
        raise ValueError("Cached count must be an integer between zero and prompt count")
    uncached = payload["prompt_eval_count"] - cached
    prompt_duration = payload.get("prompt_eval_duration")
    if type(prompt_duration) is not int or prompt_duration < 0:
        raise ValueError("Expected a nonnegative integer: prompt_eval_duration")
    if uncached > 0 and prompt_duration == 0:
        raise ValueError("Uncached prompt tokens require positive evaluation time")
    response = payload.get("response")
    normal_stop_with_text = (
        isinstance(response, str) and bool(response.strip())
        and payload.get("done_reason") == "stop"
    )
    decode_tps = payload["eval_count"] / (payload["eval_duration"] / 1e9)
    prompt_tps = uncached / (prompt_duration / 1e9) if uncached else None
    total_seconds = payload["total_duration"] / 1e9

    return {
        "model": payload.get("model", ""),
        "uncached_prompt_tokens_per_second": round(prompt_tps, 2) if prompt_tps is not None else None,
        "generated_tokens_per_second": round(decode_tps, 2),
        "server_total_seconds": round(total_seconds, 3),
        "normal_stop_with_text": normal_stop_with_text,
        "timing_passed": decode_tps >= 20.0 and total_seconds <= 5.0,
        "timing_and_stop_passed": normal_stop_with_text and decode_tps >= 20.0 and total_seconds <= 5.0,
    }

sample_response = {
    "model": "gemma4:12b-it-q4_K_M",
    "done": True,
    "done_reason": "stop",
    "prompt_eval_count": 16,
    "prompt_eval_cached_count": 0,
    "prompt_eval_duration": 180_000_000,
    "eval_count": 128,
    "eval_duration": 3_200_000_000,
    "total_duration": 3_420_000_000,
    "response": "1. Verify test inputs and fixtures.\n2. Check assertion diffs.\n3. Isolate environment state.",
}

# Cache reuse must not inflate the uncached processing rate.
partly_cached = {**sample_response, "prompt_eval_count": 1000,
                 "prompt_eval_cached_count": 800, "prompt_eval_duration": 200_000_000}
assert parse_ollama_metrics(partly_cached)["uncached_prompt_tokens_per_second"] == 1000.0
fully_cached = {**sample_response, "prompt_eval_cached_count": 16, "prompt_eval_duration": 0}
assert parse_ollama_metrics(fully_cached)["uncached_prompt_tokens_per_second"] is None
truncated = {**sample_response, "done_reason": "length"}
assert parse_ollama_metrics(truncated)["timing_and_stop_passed"] is False

if __name__ == "__main__":
    payload = sample_response
    if "--receipt" in sys.argv:
        receipt_path = sys.argv[sys.argv.index("--receipt") + 1]
        payload = json.loads(Path(receipt_path).read_text())
    print(json.dumps(parse_ollama_metrics(payload), indent=2))
Output
{
  "model": "gemma4:12b-it-q4_K_M",
  "uncached_prompt_tokens_per_second": 88.89,
  "generated_tokens_per_second": 40.0,
  "server_total_seconds": 3.42,
  "normal_stop_with_text": true,
  "timing_passed": true,
  "timing_and_stop_passed": true
}

Run python evaluate_ollama_run.py --receipt gemma4-12b-run.json for the saved request. Server total isn't client wall time or time to first token (TTFT). A non-streaming response can't measure user-visible TTFT; use a streaming client and timestamps for that.

Older receipts may omit prompt_eval_cached_count; the example treats the missing field as zero. That lets it read old fixtures, but doesn't establish their cache state. Label those measurements as unknown cache state unless the runtime and experiment establish that the full prompt was evaluated.

Separate cold model loading from warm requests. Preload with the same context allocation and keep-alive setting, then execute the benchmark before the model unloads:[1][13]

warmup-gemma4.sh
curl --fail --silent --show-error \
  http://127.0.0.1:11434/api/generate \
  -H 'Content-Type: application/json' \
  -d '{"model": "gemma4:12b-it-q4_K_M", "stream": false,
       "keep_alive": "5m", "options": {"num_ctx": 8192}}' > /dev/null

🎯 Measurement rule: Record cold and warm trials separately. Repeating identical prompts can reuse prompt work, so label cache state and use representative distinct prompts when measuring prefill. Hold tag, context, thinking, sampling, output cap, concurrency, and machine load constant.

Use the OpenAI-compatible API carefully

Ollama supports /v1/chat/completions and the non-stateful flavor of /v1/responses. Point the client at http://127.0.0.1:11434/v1/ and use a placeholder key such as ollama; the local server ignores that key.[16]

Check supported fields before moving a hosted integration, because this API isn't a full drop-in replacement.

OpenAI-compatible requests don't set Ollama context size per request. To pin context, create a derived model:

Modelfile
FROM gemma4:12b-it-q4_K_M
PARAMETER num_ctx 8192

Then create it and call the derived name:

terminal
set -euo pipefail

ollama create gemma4-12b-8k -f Modelfile
curl --fail --silent --show-error \
  http://127.0.0.1:11434/v1/chat/completions \
  -H 'Authorization: Bearer ollama' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gemma4-12b-8k",
    "messages": [{"role": "user", "content": "Reply with local-ready"}],
    "stream": false
  }' | jq -e '.choices[0].message.content | select(type == "string" and length > 0)'

Ollama's OpenAI compatibility docs say /v1/responses arrived in Ollama 0.13.3.[16]

It doesn't support previous_response_id or conversation state. Keep conversation state in the application when the integration depends on Responses API semantics.

Keep the template and modalities inside the runtime

Gemma 4 uses standard system, user, and assistant roles. Google recommends sampling with 1.0, 0.95, and 64 for general model behavior.[3][6]

Temperature zero in the benchmark is a test control, not Google's quality recommendation. Thinking mode changes the workload too: if enabled, preserve and inspect the separate thinking field, allow enough generation budget, and report the mode with results.[13]

Ollama applies the model when you use chat endpoints, so pass roles and content instead of assembling control tokens by hand.[6]

For native API images, use its documented base64 images field and let Ollama assemble the multimodal input.[13] The listed local tags advertise text and image only. Google's model-level audio support doesn't establish support in a particular Ollama artifact.[5]

Diagnose from the symptom

SymptomInspectNext test
curl gets connection refusedOllama serviceStart the app or ollama serve, then retry /api/tags
Pull stopsFree disk and exact errorFree space, then retry the exact tag
Model loads but the machine swapsContext and total process memoryLower num_ctx or use a smaller QAT tag
Discrete GPU shows a CPU splitollama ps PROCESSORLower context or model size, then rerun latency
First response is much slowerload_durationWarm the model, then compare warm trials
Request ends without a full answerdone_reason, response, and thinking fieldsDisable thinking for the baseline or raise the output cap
Generated tokens are sloweval_duration and eval_countCompare the exact tag and runtime with context fixed
QAT fits but answers regressTask-specific acceptance setKeep the larger artifact or choose a different model
Audio request failsOllama tag capabilitiesUse a supported runtime, or treat the workflow as unsupported

Don't change tag, prompt, context, sampling, and runtime together. One changed variable lets the run receipt answer one question.

Start with service and tag checks, then placement, then context, then timing, then task quality.

Build a reproducible local comparison

Use three prompt classes: short structured extraction, a realistic repository explanation, and fixed-length generation. The first isolates instruction following, the second introduces realistic context pressure, and the third exposes sustained decode speed.

Warm each tag once, run several measured trials, and store the raw JSON. Record:

gemma4-run-receipt.txt
machine: <CPU, GPU, and memory>
ollama_version: <ollama --version>
tag: gemma4:12b-it-q4_K_M
tag_digest: </api/tags digest>
context: 8192
processor: <ollama ps PROCESSOR>
prompt_pack: <path or commit>
sampling: temperature=0
thinking: false
output_cap: 128
concurrency: 1
warm_state: <cold model, warm model, repeated prompt>
client_wall_seconds: <measured separately>
raw_response: gemma4-12b-run.json

Latency isn't quality. Write task checks before you compare, such as valid , correct failing assertions, or a cited answer from a supplied file. A faster tag loses if it misses the required behavior. Keep the smallest tag that meets both quality and latency budgets under real context.

Share this article