Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
Picture a technician taking an inspection handheld into a utility vault with no connectivity. The product specification forbids sending facility manuals or incident text off the device. For this lesson's hypothetical terminal, the proposed targets are time to first visible answer under 250 ms, at least 18 generated tokens/s after ten minutes of use, and enough battery for an eight-hour shift. These are design requirements, not measured device results or industry standards. Actual field procedures need qualified review; fluent generated instructions alone are insufficient evidence of safety.
A lookup table can already answer "which torque specification applies to flange B-12?" A lightweight classifier can route an incident report to the right department. Generating a cited, multi-sentence repair procedure is an entirely different feature with a much stricter error budget. Just because a 4-billion-parameter checkpoint fits on the handheld's flash storage doesn't mean the feature is ready to ship.
Build on workstation memory sizing, then add battery and cooling constraints, application isolation, and a product-wide data-flow policy. Offline operation means connectivity is unavailable. An air gap is an isolation boundary; a firewall rule or a local inference call alone does not establish one.
A small language model (SLM) is small relative to its comparison set; there is no universal parameter cutoff. This lesson includes sub-billion and several-billion-parameter models. Specialization adapts behavior to a defined job through data, training, retrieval, or constraints. It does not necessarily remove parameters or erase unrelated knowledge. Edge deployment concerns execution near the user or data source, including mobile devices and larger local appliances.
Build on local memory and hardware sizing and knowledge distillation. Here we focus on what survives the journey from server-side training to physical hardware. Our executable examples demonstrate numerical routing, CPU graph export, and dispatch decisions.

Why should an edge-SLM project start from the device job instead of from the smallest downloadable model?
Answer
Parameter count alone doesn't guarantee success. The model must satisfy a specific offline job, latency target, privacy boundary, resident memory budget, power draw, and sustained thermal limit on real hardware. Those physical constraints decide the viable model size and runtime.
Start from the physical constraint, not the model size
Broad capability can be useful, but the terminal's job is narrower: extract error codes, locate approved manual passages, and explain the retrieved evidence. Start by measuring those behaviors. A specialist can outperform a larger baseline on a particular evaluation; neither its size nor a teacher relationship guarantees that outcome. Specialization is not equivalent to pruning, and local execution removes an inference round-trip only when the rest of the pipeline is local too.
Before picking weights, define each offline job and quantify its failure cost. Ask whether a deterministic lookup, a classical classifier, an extractive retrieval pipeline, or a generative SLM is required. "What model fits in 4 GB of RAM?" supplies one capacity constraint; it is not a complete engineering specification.
For an industrial inspection handheld, we define candidate operational bars across four distinct device tasks:
| Job on device | Proposed latency target | Privacy rule | Baseline to try first |
|---|---|---|---|
| Basic intent routing ("is this a fault or escalation?") | p95 < 60 ms per classification | Local only | Rules or a compact classifier |
| Field policy Q&A with citations | p95 < 250 ms time to first token (TTFT); 18 t/s at ten minutes | Local only | Local retrieval plus an instruct SLM, compared with extractive answers |
| Damaged-item photo + text decision support | p95 < 400 ms for the defined output | Local only | A vision classifier or two-stage vision/text pipeline |
| Facility policy lookup while offline | p95 < 60 ms per lookup | Local only | Indexed local documents, without generation |
Notice how no single row mandates a parameter count. An image classifier's total inference latency differs completely from a generative model's first-token latency. Time to first token (TTFT) must also measure the complete pipeline: application dispatch, local retrieval, prompt templating, tokenization, and prefill compute until the first visible answer token arrives. If a model emits hidden reasoning tokens first, track both first-generated-token latency and first-visible-token latency.
Here is a filter over invented scorecard rows. The privacy_local field is an assumed prior review result, not something this code establishes:
1jobs = [
2 {
3 "name": "intent",
4 "privacy_local": True,
5 "p95_ms": 42,
6 "max_p95_ms": 60,
7 "quality": 0.98,
8 "min_quality": 0.97,
9 },
10 {
11 "name": "policy_qa",
12 "privacy_local": True,
13 "p95_ms": 228,
14 "max_p95_ms": 250,
15 "quality": 0.93,
16 "min_quality": 0.92,
17 },
18 {
19 "name": "photo_support",
20 "privacy_local": True,
21 "p95_ms": 470,
22 "max_p95_ms": 400,
23 "quality": 0.91,
24 "min_quality": 0.92,
25 },
26]
27
28approved = [
29 job["name"]
30 for job in jobs
31 if job["privacy_local"] is True
32 and job["p95_ms"] < job["max_p95_ms"]
33 and job["quality"] >= job["min_quality"]
34]
35print(f"routes passing these fixture checks: {approved}")1routes passing these fixture checks: ['intent', 'policy_qa']Photo support misses both its latency ceiling and quality threshold. Intent routing and policy Q&A pass these preliminary checks, yet hardware qualification remains to be done.
These approved rows also separate distinct model workloads. Intent routing simply selects an execution lane. Policy Q&A must retrieve rules, ground its response in exact citations, and abstain when documentation is missing.
In the inspection handheld example, why is "maintenance policy Q&A with citations" a different model target than "basic intent routing"?
Answer
Routing selects an execution path. Policy Q&A needs correct retrieval, citations that support the claims, and abstention when evidence is insufficient. Compare extractive answers with generation; neither workload has a guaranteed latency or required parameter count before measurement.
A bandwidth worksheet, with its assumptions visible
For a dense model whose matrix weights exceed cache capacity, batch-1 decode often spends much of its time moving weights. Compute, attention, small kernel launches, format conversion, and runtime scheduling can also limit it. Embedding lookups, inactive experts, cached weights, and immediate reuse invalidate an assumption that every stored parameter is read from DRAM on every token.
Let count matrix weights used in this step. Assuming one complete packed-weight read, two floating-point operations per multiply-add, and no other traffic gives:
This model gives 1 FLOP/byte for 16-bit weights and 4 for four-bit weights. At an assumed 50 GB/s, the latter corresponds to 200 GFLOP/s of modeled matrix work. Integer TOPS, sparse operation counts, and floating-point FLOPs are different metrics; dividing this work rate by an NPU's advertised TOPS does not establish utilization or idle time. Quantization scales and unpacking also cost bytes and work.
We calculate the theoretical decoding ceiling using the weight-streaming bound:
Assume a dense model step reads 3.8 billion active matrix weights once over a hypothetical 50 GB/s bus. The raw payload is 1.9 GB at four bits or 3.8 GB at eight bits. This excludes quantization metadata, mixed-precision tensors, KV/activation traffic, and invocation overhead; it is not an artifact-size prediction for every 3.8B model:
1def weight_gb(params_billion: float, bits: int) -> float:
2 return params_billion * bits / 8
3
4def weight_streaming_ceiling(bandwidth_gbs: float, active_weight_gb: float) -> float:
5 return bandwidth_gbs / active_weight_gb
6
7params_b = 3.8
8bandwidth_gbs = 50.0 # illustrative handheld budget, not a device datasheet
9floor_tps = 18.0
10
11for bits in (8, 4):
12 artifact_gb = weight_gb(params_b, bits)
13 ceiling = weight_streaming_ceiling(bandwidth_gbs, artifact_gb)
14 print(
15 f"{bits}-bit: {artifact_gb:.1f} GB weights, "
16 f"ceiling {ceiling:.1f} t/s, "
17 f"not ruled out by this bound: {ceiling >= floor_tps}"
18 )18-bit: 3.8 GB weights, ceiling 13.2 t/s, not ruled out by this bound: False
24-bit: 1.9 GB weights, ceiling 26.3 t/s, not ruled out by this bound: TrueUnder these assumptions, the eight-bit candidate is ruled out: 50 / 3.8 is only 13.2 steps/s. Four-bit storage is not ruled out by the 26.3 steps/s estimate. It still needs a working kernel and device measurement. Changing the model's active weights, cache reuse, or decoding algorithm changes this worksheet's applicability; these numbers are not guaranteed token rates.
Total memory footprint requires its own accounting:
Flash storage, system RAM, and accelerator-accessible allocations are different budgets. A download can fit while serving allocations fail or the OS terminates the app under memory pressure. Include tokenizer and retrieval storage, multimodal encoders, runtime pools, and other applications. An individual allocation limit and a system low-memory termination are different failure mechanisms.
MobileLLM introduced dedicated architectural optimizations for sub-billion models deployed to mobile chips.[1] Use its empirical findings to select model topologies, but always measure actual energy, memory residency, and throughput on your specific target SKU.
Why does an NPU with 45 TOPS of compute still struggle to generate 20 tokens per second for a 3B parameter model?
Answer
If its active matrix weights require 1.5 GB of DRAM traffic per step and only 50 GB/s is available, that traffic alone takes 30 ms, giving about 33 steps/s before other work. This conditional bound does not explain a measured result by itself or convert integer TOPS into floating-point throughput.
Distillation: transferring dark knowledge to compact silicon
Knowledge distillation trains a student against a teacher's predictions or other selected signals. The teacher need not be a frontier language model. The student can use a different architecture, and deployment need not include the teacher's weights. Whether this improves the device task depends on the training setup and evaluation.[2]
In classification, hard labels tell the student which class won, but hide how the teacher ranked the alternatives. In language modeling, teacher logits provide a dense probability distribution over the entire vocabulary. This distribution contains "dark knowledge": the subtle semantic relationships between runner-up tokens that were plausible versus those that were absurd.
The distillation objective
When distilling token by token, we combine standard cross-entropy on ground-truth tokens with Kullback-Leibler (KL) divergence over temperature-scaled teacher and student distributions:
Here and are aligned teacher and student logits, positive is the temperature, is softmax, and balances supervision. The hard-label CE uses unscaled student logits; the soft term uses the same temperature for teacher and student.
Temperature changes how strongly logit differences affect probabilities. For a peaked distribution, increasing positive gives alternatives more mass without changing their rank:
Notice the ratio between two candidate tokens and :
For a fixed teacher, the exact gradient of the unscaled forward KL with respect to a student logit is:[2]
There are two temperature effects: the chain rule contributes , and at high temperature the probability difference is approximately proportional to . Center each logit vector by subtracting its mean. For a vocabulary of size :
Multiplying by approximately compensates for this high-temperature shrinkage. It does not guarantee equal gradients at every temperature, stable optimization, or successful training. Temperature and loss weights remain validation choices.
Check the exact derivative against a finite difference before relying on the approximation:
1import math
2
3teacher = [4.0, 2.0, 1.0]
4student = [3.2, 2.4, 1.1]
5
6def probabilities(logits, tau):
7 peak = max(logits)
8 values = [math.exp((v - peak) / tau) for v in logits]
9 return [v / sum(values) for v in values]
10
11def loss(logits, tau):
12 p, q = probabilities(teacher, tau), probabilities(logits, tau)
13 return sum(a * math.log(a / b) for a, b in zip(p, q))
14
15for tau in (1.0, 3.0, 100.0):
16 p, q = probabilities(teacher, tau), probabilities(student, tau)
17 exact = (q[0] - p[0]) / tau
18 step = 1e-4
19 plus, minus = student.copy(), student.copy()
20 plus[0] += step
21 minus[0] -= step
22 numeric = (loss(plus, tau) - loss(minus, tau)) / (2 * step)
23 assert math.isclose(numeric, exact, rel_tol=1e-6, abs_tol=1e-10)
24 print(f"tau={tau:.0f}: gradient={exact:.6f}, scaled={tau*tau*exact:.6f}")1tau=1: gradient=-0.207576, scaled=-0.207576
2tau=3: gradient=-0.029854, scaled=-0.268686
3tau=100: gradient=-0.000024, scaled=-0.235045The scaled values remain of comparable order, but are not equal. For these centered logits, the high-temperature limit is about -0.233333 for the first coordinate. This example checks a derivative, not a trained language model.
Let's test this behavior with teacher logits [4.0, 2.0, 1.0]:
1import math
2
3def softmax(values, temperature):
4 shifted = [value / temperature for value in values]
5 peak = max(shifted)
6 weights = [math.exp(value - peak) for value in shifted]
7 total = sum(weights)
8 return [weight / total for weight in weights]
9
10teacher_logits = [4.0, 2.0, 1.0]
11for temperature in (1.0, 3.0):
12 probs = softmax(teacher_logits, temperature)
13 print(f"temperature={temperature:.0f}: {[round(p, 3) for p in probs]}")1temperature=1: [0.844, 0.114, 0.042]
2temperature=3: [0.532, 0.273, 0.196]![Softmax for teacher logits [4.0, 2.0, 1.0]: probabilities change from [0.844, 0.114, 0.042] at temperature 1 to [0.532, 0.273, 0.196] at temperature 3. The ranking stays A, B, C while the B/C ratio falls from 2.72 to 1.40.](/cdn/content-image/fundamentals/slm-specialization-edge-deployment/illustrations/_generated/soft_target_temperature_dark.png?v=a3ca5df9ee1e)
At , A receives 84.4% and C 4.2%. At , A remains preferred, while B and C receive more mass. A one-hot label discards their relative preferences. If the teacher is wrong, matching those preferences can transfer that error too.
Distillation with execution-based verifiers
Direct token-level KL needs distributions over aligned token events at aligned prefixes. Matching token IDs alone is insufficient if their meanings differ. Different tokenizers need an alignment method, or sequence-level supervision using teacher text retokenized by the student. Closed APIs may expose only sampled text or a limited set of log probabilities; these are not automatically a full-vocabulary distribution.
Phi-3 combined filtered public data with synthetic examples; it is a data-curation result, not a universal requirement to use only teacher-generated text.[3] For domain specialization, include approved documents, human examples, unanswerable questions, and verified synthetic cases. Split evaluation by document, equipment family, or site before generating variants so near-duplicates cannot leak across the split.
Check each synthetic example against the contract it is supposed to teach:
- Code and tool-use verification: Run bounded tests in an environment with appropriate isolation, resource limits, and no secrets. A container alone is not a sufficient boundary for hostile code. A passing test covers its assertions, not every behavior.
- Schema validation: Check required fields, types, and allowed values. Well-formed JSON can still request an unauthorized action or contain a wrong fact.
- Retrieval grounding: Check source identity, version, applicable equipment and conditions, and whether the passage supports the claim. An existing paragraph number or matching numeric substring alone is insufficient.
Here is an invented maintenance fixture, not a real torque recommendation. Both outputs pass a naive citation-and-number check:
1manual = {
2 "M-7": {"equipment": "pump-A", "torque_nm": 12},
3 "M-8": {"equipment": "pump-B", "torque_nm": 30},
4}
5question_equipment = "pump-A"
6answers = [
7 {"citation": "M-7", "torque_nm": 12},
8 {"citation": "M-8", "torque_nm": 30},
9]
10for answer in answers:
11 source = manual[answer["citation"]]
12 naive_pass = source["torque_nm"] == answer["torque_nm"]
13 applicable_pass = naive_pass and source["equipment"] == question_equipment
14 print(f'{answer["citation"]}: naive={naive_pass}, applicable={applicable_pass}')1M-7: naive=True, applicable=True
2M-8: naive=True, applicable=FalseThe second answer quotes a real value for the wrong equipment. Even the better check above covers only this tiny structured fixture. Review ambiguous cases and measure held-out correctness and abstention; curation cannot guarantee a flawless student.
MiniLLM uses on-policy optimization of reverse KL, , with variance reduction, teacher-mixed sampling, and length normalization.[4] Reverse KL tends to favor teacher modes rather than covering every alternative. That is an objective tradeoff, not a guarantee of factual correctness or a replacement for task evaluation.
This CPU calculation combines three invented signals. In real training, freeze teacher gradients, align token prefixes, mask padding, and define reductions explicitly. Hidden-state matching also needs chosen layers and compatible dimensions, often a learned projection; unrelated hidden coordinates are not interchangeable:
1import math
2
3def log_softmax(logits, temperature=1.0):
4 if not logits or not math.isfinite(temperature) or temperature <= 0:
5 raise ValueError("nonempty logits and positive finite temperature required")
6 if not all(math.isfinite(value) for value in logits):
7 raise ValueError("finite logits required")
8 shifted = [value / temperature for value in logits]
9 if not all(math.isfinite(value) for value in shifted):
10 raise ValueError("temperature-scaled logits must remain finite")
11 peak = max(shifted)
12 centered = [value - peak for value in shifted]
13 if not all(math.isfinite(value) for value in centered):
14 raise ValueError("centered logits must remain finite")
15 normalizer = math.log(sum(math.exp(value) for value in centered))
16 return [value - normalizer for value in centered]
17
18def cross_entropy(logits, label):
19 return -log_softmax(logits)[label]
20
21def kl_divergence(target_logs, candidate_logs):
22 if len(target_logs) != len(candidate_logs):
23 raise ValueError("aligned vocabulary sizes required")
24 return sum(
25 math.exp(p_log) * (p_log - q_log)
26 for p_log, q_log in zip(target_logs, candidate_logs)
27 )
28
29teacher_logits = [4.0, 2.0, 1.0]
30student_logits = [3.2, 2.4, 1.1]
31label = 0
32temperature = 2.0
33assert cross_entropy([0.0, -1000.0], 1) == 1000.0
34assert kl_divergence(log_softmax([0.0, -1000.0]),
35 log_softmax([-1000.0, 0.0])) == 1000.0
36
37hard_loss = cross_entropy(student_logits, label)
38soft_loss = kl_divergence(
39 log_softmax(teacher_logits, temperature),
40 log_softmax(student_logits, temperature),
41) * (temperature * temperature)
42
43teacher_hidden = [0.4, -0.1, 0.2]
44student_hidden = [0.3, 0.0, 0.1]
45assert len(teacher_hidden) == len(student_hidden) > 0
46hidden_loss = sum(
47 (s - t) ** 2 for s, t in zip(student_hidden, teacher_hidden)
48) / len(teacher_hidden)
49
50alpha, beta = 0.8, 0.2
51total_loss = (1 - alpha) * hard_loss + alpha * soft_loss + beta * hidden_loss
52print(f"hard_loss={hard_loss:.4f}")
53print(f"soft_loss={soft_loss:.4f}")
54print(f"hidden_loss={hidden_loss:.4f}")
55print(f"total_loss={total_loss:.4f}")1hard_loss=0.4522
2soft_loss=0.1480
3hidden_loss=0.0100
4total_loss=0.2109Once the student finishes training, evaluate its exported and quantized artifact again. Sensitivity depends on the model, format, calibration data, kernels, and workload. There is no universal sub-2B cutoff requiring one quantization method.
These synthetic rows illustrate a quality-versus-latency decision, not measured INT8/INT4 results. quality must have a defined task metric and uncertainty before using this filter in a release:
1baseline = {"quality": 0.94, "p95_ms": 132}
2exports = [
3 {"name": "int8", "quality": 0.938, "p95_ms": 104},
4 {"name": "int4", "quality": 0.887, "p95_ms": 71},
5]
6max_quality_drop = 0.02
7latency_gate_ms = 120
8
9approved = [
10 item["name"]
11 for item in exports
12 if baseline["quality"] - item["quality"] <= max_quality_drop
13 and item["p95_ms"] <= latency_gate_ms
14]
15print(f"approved exports: {approved}")1approved exports: ['int8']In this fixture, INT4 loses 0.053 quality and fails the 0.02 tolerance; INT8 loses 0.002 and passes both listed thresholds. Neither row establishes hardware qualification. Compare paired held-out outcomes, source support, abstention, latency, and memory under the actual decode settings; perplexity alone is insufficient.
What does multiplying the KL distillation loss by tau squared compensate for, and what does it not guarantee?
Answer
The exact unscaled gradient is (p_student - p_teacher) / tau. At high temperature, the probability difference itself is approximately proportional to 1/tau, giving the familiar approximate 1/tau² shrinkage. The multiplier compensates in that regime; it does not ensure identical gradient magnitudes or stable training at every temperature.
Architecture choices that matter at small scale
A smaller parameter count doesn't guarantee efficient execution. Layer structure and attention heads change the work, memory traffic, and cache behavior.
Meta's MobileLLM research revealed key architectural principles for sub-billion models:[1]
- Depth versus width: The paper's approximately 125M/350M experiments favored deeper, thinner designs. Scaling laws do not require equal parameter allocation to depth and width, and this result is not universal across training recipes or devices.
- Embedding tying: Sharing input and output weights reduces storage. One ablation reduced 135M to 119M parameters, with a small accuracy drop; reallocating some savings to depth recovered it.
- Immediate block reuse: MobileLLM-LS repeats each block with shared weights. Cache locality can reduce repeat weight traffic, but compute and attention state still cost resources. Cache residency and zero extra DRAM traffic are not guaranteed. Its iPhone 13 FP16/MPS experiment measured 15.6 versus 16.0 ms execution, averaged over 50 iterations, not our terminal's sustained token rate.
- GQA: Fewer KV heads reduce cache payload for otherwise equal dimensions. Choose a trained architecture; replacing an existing model's attention heads is not an inference-time switch.
The same paper found its tested distillation from Llama-v2 7B slower to train and comparable or inferior in accuracy to hard-label training. Architectural improvement and successful distillation are separate hypotheses.[1]
KV cache growth can rapidly trigger out-of-memory crashes on devices with tight RAM. Let's calculate the memory savings of 4 KV heads compared to 16 KV heads across a 4,096-token context window:
1def kv_mib(layers, kv_heads, head_dim, tokens, bytes_per_value=2):
2 return 2 * layers * kv_heads * head_dim * tokens * bytes_per_value / 2**20
3
4mha = kv_mib(layers=24, kv_heads=16, head_dim=64, tokens=4096)
5gqa = kv_mib(layers=24, kv_heads=4, head_dim=64, tokens=4096)
6print(f"MHA KV cache: {mha:.1f} MiB")
7print(f"GQA KV cache: {gqa:.1f} MiB")
8print(f"cache reduction for this architecture: {mha / gqa:.1f}x")1MHA KV cache: 384.0 MiB
2GQA KV cache: 96.0 MiB
3cache reduction for this architecture: 4.0xThe hypothetical homogeneous cache shrinks from 384 to 96 MiB. This excludes alignment, pools, quantization metadata, and temporary state. It does not establish total memory fit or fourfold speed. Hybrid attention and recurrent layers require their own per-layer formulas.
These documented model examples illustrate different tradeoffs, checked September 22, 2026. They are not a ranking or an exhaustive list of current candidates:
- Phi-4-mini-instruct (3.8B): Employs GQA, shared input-output embeddings, a 128K context window, and high-density synthetic data tuning for tool calling.[5]
- Gemma 4 E2B and E4B: The card lists 2.3B/4.5B effective parameters and 5.1B/8B including embeddings. Per-layer embeddings use token lookups. Effective is not the same label as an MoE active-expert count. Include the selected artifact's stored tensors, encoder components, cache, and measured serving residency; the effective count cannot establish fit.[6]
- Qwen3-1.7B: Its card documents 28 layers, 16 query/8 KV heads, a 32,768-token context, and thinking/non-thinking chat-template modes.[7] Thinking tokens consume generation work before the visible answer. Budget and evaluate both modes under the installed runtime; an output cap can truncate the answer rather than improve it.
Why can a model advertised as "2.3B effective parameters" still cause an out-of-memory crash on a 4 GB mobile device?
Answer
The effective label excludes large lookup embeddings. E2B lists 5.1B parameters including embeddings, before considering the chosen representation and additional components. A memory budget needs artifact bytes, allocations, cache, workspaces, and device coexistence. Parameter labels alone cannot predict peak residency.
On-device runtimes: mapping graphs to silicon
Next inspect how the selected model and format execute. A runtime may use supported accelerator kernels, native CPU kernels, or both; CPU execution is not necessarily emulation.
| Runtime | Primary export format | Hardware backends | Key engineering check |
|---|---|---|---|
| ExecuTorch[8] | Backend-specific .pte plus required assets | Core ML, MPS, Qualcomm, Vulkan, XNNPACK, others | Inspect delegated partitions, CPU fallback, supported models and target versions. |
| Apple Core ML[9] | .mlpackage, compiled for execution | ANE, GPU, CPU according to supported operations and configuration | Inspect the actual compute plan; compressed weights do not imply integer NPU execution. |
| MLC LLM (Apache TVM)[10] | Converted weights, configuration, compiled model library | Platform-specific GPU backends including Metal, Vulkan, WebGPU | Ship the correct library and assets; measure the supported model on the target. |
| ONNX Runtime Mobile[11] | ONNX or optimized ORT model | CPU and platform/package-dependent execution providers | Check operator coverage and partitions in the mobile build. |
| ONNX Runtime GenAI[12] | Compatible ONNX model, generation configuration and tokenizer | Supported providers for the selected GenAI build | Verify model-builder, provider, and platform support separately from general ORT Mobile. |
| llama.cpp[13] | GGUF and any separate modality assets | CPU/GPU backends selected at build/runtime | Match model architecture, quantization type, backend kernels and placement. |
Mobile NPUs and the danger of CPU fallback
An NPU may improve energy efficiency for supported work. Advertised TOPS/W figures across formats, sparsity modes, and chips do not establish energy per generated token for your application.
Check the chosen accelerator's restrictions:
- Shape support: Shapes, ranges, and cache updates may impose restrictions. Buckets can help some backends, but padding can waste work and change state if masking is wrong.
- Numerical support: Weight storage, multiply format, and accumulator precision are separate choices. Check this backend's exact support rather than assuming all NPUs use INT4/INT8 with FP16 accumulation.
- Operator support: An unsupported operation can remain on CPU, cause export failure, or require a different lowering. Check placement and partition boundaries, not only a percentage of delegated operators.
Partial delegation can add synchronization, layout conversion, intermediate materialization, or copies. Unified physical memory does not remove all those costs, and it does not imply a mandatory copy between two private memories either. Some partitions are beneficial. Profile their frequency and critical-path time, then compare a supported fused graph and an optimized CPU path.[8][11]
Apple's preparation guide explicitly distinguishes compression methods and device strengths, including GPU-oriented four-bit weight quantization and stateful KV-cache updates. Core ML conversion is not a promise that the whole LLM runs on the Neural Engine.[9]
Testing graph export on CPU
First check capture, serialization, and declared shape constraints. torch.export records a graph that downstream tools can lower; it is not an ExecuTorch executable or an NPU compiler.[14] This CPU round-trip uses pinned PyTorch 2.14.0, the version checked for this lesson:
1from pathlib import Path
2from tempfile import TemporaryDirectory
3import torch
4from torch import nn
5torch.set_num_threads(1)
6
7class NumericRouter(nn.Module):
8 def __init__(self):
9 super().__init__()
10 self.projection = nn.Linear(3, 2)
11 with torch.no_grad():
12 self.projection.weight.copy_(torch.tensor([
13 [-1.0, -1.0, -1.0], [1.0, 1.0, 1.0],
14 ]))
15 self.projection.bias.zero_()
16
17 def forward(self, features):
18 return self.projection(features.clamp(-1.0, 1.0))
19
20model = NumericRouter().eval()
21example = torch.tensor([[-0.5, -0.2, 0.1], [0.8, 0.7, 0.6]])
22batch = torch.export.Dim("batch", min=1, max=8)
23program = torch.export.export(
24 model, (example,), dynamic_shapes=({0: batch},),
25)
26
27with TemporaryDirectory() as directory:
28 artifact = Path(directory) / "router.pt2"
29 torch.export.save(program, artifact)
30 restored = torch.export.load(artifact).module()
31 with torch.inference_mode():
32 for size in (1, 2, 5, 8):
33 inputs = torch.linspace(-2, 2, size * 3).reshape(size, 3)
34 expected, actual = model(inputs), restored(inputs)
35 torch.testing.assert_close(actual, expected)
36 assert torch.equal(actual.argmax(-1), expected.argmax(-1))
37 classes = restored(example).argmax(-1).tolist()
38 assert classes == [0, 1]
39 try:
40 restored(torch.zeros(2, 4))
41 except (AssertionError, RuntimeError):
42 pass
43 else:
44 raise AssertionError("exported graph accepted the wrong feature dimension")
45
46print(f"restored class choices: {classes}")
47print("logits match for batch sizes 1, 2, 5, 8; invalid feature shape rejected")
48print("CPU graph round-trip only: no mobile runtime or accelerator measured")1restored class choices: [0, 1]
2logits match for batch sizes 1, 2, 5, 8; invalid feature shape rejected
3CPU graph round-trip only: no mobile runtime or accelerator measuredThe example checks four permitted batches and rejects a wrong feature dimension. It uses PyTorch on CPU after reload. Backend lowering, mobile packaging, delegate precision, and device performance still need their own checks.
Why can an unoptimized NPU deployment be significantly slower than running the exact same model entirely on CPU?
Answer
Unsupported operators can add expensive partition boundaries. Synchronization, conversion, materialization, and sometimes copies may outweigh the acceleration. This depends on the memory architecture and partition schedule; measure the critical path rather than assuming every fallback is fatal.
Heat, battery, and physical edge limits
Cooling depends on the enclosure, contact with the user, ambient temperature, workload, and any active cooling. A fanless device transfers heat through conduction, convection, and radiation. There is no universal 3-to-5 W sustained limit for all handhelds. Our 3.5 W allowance is an assumed product budget, not a chip TDP or a regulatory threshold.
Dynamic voltage and frequency scaling (DVFS)
Dynamic voltage and frequency scaling (DVFS) changes operating points to manage performance and power. Thermal control can also limit available performance. Junction, battery, and chassis temperatures are different measurements; throttling thresholds and touch-temperature requirements depend on the device and applicable conditions. Neither 85°C nor 42°C is a universal rule.
Android's Thermal API exposes thermal status and headroom. Its documentation notes that mappings and support vary by device; even THERMAL_STATUS_NONE does not always establish an absence of throttling.[15] Record the actual device, OS/runtime versions, ambient conditions, charging state, workload, output length, and measurement method.

This code checks three synthetic samples. It is not a ten-minute hardware soak test, a tail-latency check, or proof of performance between samples:
1samples_tps = [23.0, 19.4, 16.8]
2minimum_sustained_tps = 18.0
3burst_tps = samples_tps[0]
4floor_tps = min(samples_tps)
5
6print(f"burst throughput: {burst_tps:.1f} t/s")
7print(f"lowest sampled throughput: {floor_tps:.1f} t/s")
8print(f"sampled throughput gate passes: {floor_tps >= minimum_sustained_tps}")1burst throughput: 23.0 t/s
2lowest sampled throughput: 16.8 t/s
3sampled throughput gate passes: FalseThe lowest supplied sample fails 18 tokens/s. A repair could involve a smaller candidate, a supported kernel, better workload scheduling, or fewer generated tokens, but each changes tradeoffs and needs measurement. Quantization may add conversion overhead; output caps can damage answer quality. Test continuous load and a realistic shift duty cycle, including warm starts and hot ambient conditions.
Battery energy accounting
Assume a 15 Wh battery; 4,000 mAh at a nominal 3.7 V corresponds to 14.8 Wh before usable-capacity and aging considerations. With constant measured power at a defined boundary and generation rate :
If the system draws 3.5W while decoding at 18 tokens per second, generating 1,000 tokens takes 55.6 seconds and consumes:
Do those watts include display, retrieval, prefill, idle time, and radios, or only incremental decode power? That boundary changes the interpretation. A lower-power candidate can consume more energy per token if it is sufficiently slower. This synthetic shift includes decode, prefill/retrieval, and remaining idle time with mutually exclusive whole-system power assumptions:
1shift_seconds = 8 * 3600
2generated_tokens = 30_000
3battery_wh = 15
4retrieval_prefill_seconds = 600
5retrieval_prefill_watts = 4
6idle_watts = 1.2
7
8for name, decode_watts, tokens_per_second in [("candidate-A", 3.5, 18), ("candidate-B", 6, 10)]:
9 decode_seconds = generated_tokens / tokens_per_second
10 idle_seconds = shift_seconds - decode_seconds - retrieval_prefill_seconds
11 assert idle_seconds >= 0
12 decode_wh = decode_watts * decode_seconds / 3600
13 total_wh = (decode_watts * decode_seconds
14 + retrieval_prefill_watts * retrieval_prefill_seconds
15 + idle_watts * idle_seconds) / 3600
16 print(f"{name}: decode={decode_wh:.2f} Wh, shift={total_wh:.2f} Wh, below nominal={total_wh < battery_wh}")1candidate-A: decode=1.62 Wh, shift=11.13 Wh, below nominal=True
2candidate-B: decode=5.00 Wh, shift=14.07 Wh, below nominal=TrueBeing below nominal capacity is not a battery qualification. Reserve, usable capacity, aging, and longer prompts can erase the margin. For varying power, integrate over time; dividing average power by average token rate requires a consistent measurement interval. Measure total energy for a complete task when comparing classifiers, extractive answers, and generators of different lengths.
What can a 10-second token generation benchmark establish, and what does it miss?
Answer
It measures the workload under the recorded initial conditions and can help debug regressions. It cannot establish an eight-hour battery budget or sustained behavior. Add representative longer runs and thermal/energy measurements; do not assume a universal warm-up time or percentage slowdown.
Hybrid edge-cloud routing with fail-closed privacy
Some products permit a cloud route for selected inputs; our restricted facility task does not. Complexity alone supplies no authority to transmit data. On an isolated terminal, an unsupported question needs local abstention or review. The optional router below illustrates a different product mode with an explicitly permitted public-data route, not an exception to the terminal's no-egress requirement.
Mathematical confidence routing
Neither self-reported confidence nor token probabilities prove an answer is correct. Some useful diagnostic features are:
- Predictive entropy: Calculate the entropy of the next-token probability distribution: Low next-token entropy means a concentrated distribution over the tokens being considered. A confidently wrong number can have low entropy; several equally correct phrasings can have high entropy. Temperature, logits filtering, vocabulary, and prefix affect this statistic.
- Top-2 softmax margin: Measure the gap between the most likely token and the runner-up: A small margin means the two token probabilities are close. It is not a statistical significance test or a universal escalation threshold.
- Evidence checks: Verify that cited passages apply and support the answer, not merely that their identifiers exist.
If you combine these into a routing score, calibrate it against held-out answer correctness for the task and recheck after model, quantization, prompt, or sampling changes. Report coverage and error rates for answered versus abstained cases. The router's 0.80 threshold is an invented policy value, not the result of such calibration.
A fail-closed transmission policy
When uncertainty triggers an escalation request, application security policies must govern whether data can actually leave the device.
A permitted hybrid product needs a transmission policy enforced outside generated text:
- Data classification: Trusted classification must cover the assembled payload: user input, retrieved documents, history, attachments, and tool results. A model or user-provided
publiclabel is not authority. Unknown or mixed sensitive inputs retain the restrictive policy. - Fail-closed rule: If a query touches
confidentialorrestricteddata, or if the device is currently offline, cloud escalation is strictly forbidden. The system must fail closed: it either routes to local human review or returns a safe local abstention. - Surrounding paths: Review telemetry, crash reports, sync, backups, tools, downloads, and fallback paths. Use minimal allowed metrics, access controls, retention, and egress restrictions. A locally running model does not establish privacy for the surrounding app.
![Diagram showing Trusted payload classification and task score, Non-boolean finite number in [0, 1]?, Supported local answer and score >= 0.80?, and local_answer.](/cdn/content-image/fundamentals/slm-specialization-edge-deployment/diagrams/_generated/content_diagram_0_dark.png?v=deac214297c9)
The following is a pure decision function over assumed trusted inputs. It does not perform redaction, classify a payload, check a permission receipt, send a request, or enforce a network boundary. Bind real authorization to the actual payload and destination immediately before transmission:
1import math
2
3def route_query(score, *, supported=False, data_class="unknown",
4 cloud_approved=False, redaction_reviewed=False, online=False):
5 valid_score = type(score) in (int, float) and 0 <= score <= 1 and math.isfinite(score)
6 if not valid_score:
7 return "human_review_local_only"
8 if supported is True and score >= 0.80:
9 return "local_answer"
10 if (data_class == "public" and cloud_approved is True
11 and redaction_reviewed is True and online is True):
12 return "cloud_allowed_after_redaction"
13 return "human_review_local_only"
14
15public_approval = dict(data_class="public", cloud_approved=True,
16 redaction_reviewed=True, online=True)
17cases = [
18 ("routine_public", route_query(0.94, supported=True), "local_answer"),
19 ("hard_public", route_query(0.41, **public_approval), "cloud_allowed_after_redaction"),
20 ("hard_private", route_query(0.38, data_class="restricted"), "human_review_local_only"),
21 ("unknown", route_query(0.41), "human_review_local_only"),
22 ("offline", route_query(0.41, **(public_approval | {"online": False})), "human_review_local_only"),
23]
24for name, actual, expected in cases:
25 assert actual == expected
26 print(f"{name}: {actual}")
27for invalid in (float("nan"), float("inf"), -0.1, 1.1, 10**1000, True, None):
28 assert route_query(invalid, supported=True, **public_approval) == "human_review_local_only"
29for permission in ("cloud_approved", "redaction_reviewed", "online"):
30 assert route_query(0.41, **(public_approval | {permission: False})) == "human_review_local_only"1routine_public: local_answer
2hard_public: cloud_allowed_after_redaction
3hard_private: human_review_local_only
4unknown: human_review_local_only
5offline: human_review_local_onlyhard_private, unknown, and offline select local review. The tests establish decisions for those supplied inputs, not protection against mislabeled content, mutable payloads, unauthorized callers, or an independent network path.
Why is a high entropy score from the model not sufficient authorization to send a request to the cloud?
Answer
Entropy describes a token distribution, not transmission authority or factual accuracy. The declared policy forbids sending restricted or unknown data. A separately authorized public route must check the complete payload and destination before any transmission.
Candidate selection and over-the-air deployment gates
With constraints defined and routing established, benchmark candidate model families against your specific hardware SKU:
| Model family | Parameters | Key architectural focus | Target operational job |
|---|---|---|---|
| MobileLLM / MobileLLM-LS (Meta)[1] | 125M / 350M research variants | Depth/tying/GQA; immediate reuse in LS | Candidate routing or extraction, with task evaluation |
| Phi-4-mini (Microsoft)[5] | 3.8B | GQA, 128K context, synthetic textbook training | Policy Q&A, structured tool calling |
| Gemma 4 E2B/E4B (Google)[6] | 2.3B / 4.5B effective; 5.1B / 8B with embeddings | Per-layer embeddings, text/image/audio inputs | Candidate multimodal assistant, with all required assets budgeted |
| Qwen3 (Qwen Team)[7] | 1.7B | Code generation, dual thinking/non-thinking modes | Multilingual field forms, structured data extraction |
Here are synthetic candidates checked against two thresholds. They are not benchmark results for the named model families above:
1candidates = [
2 {"name": "tiny_policy", "quality": 0.88, "sample_min_tps": 33},
3 {"name": "policy_student", "quality": 0.93, "sample_min_tps": 19},
4 {"name": "hot_policy", "quality": 0.95, "sample_min_tps": 11},
5]
6min_quality, min_sample_tps = 0.92, 18
7
8approved = [
9 item["name"]
10 for item in candidates
11 if item["quality"] >= min_quality
12 and item["sample_min_tps"] >= min_sample_tps
13]
14assert approved == ["policy_student"]
15print(f"candidates passing the two fixture thresholds: {approved}")1candidates passing the two fixture thresholds: ['policy_student']Only policy_student passes these two fixture thresholds. The row names do not identify actual models; a low rate does not establish thermal throttling. Quality uncertainty, memory, energy, privacy, packaging, and sustained workload checks remain outstanding.
Over-the-air (OTA) lifecycle and dual-slot rollback
An interrupted in-place update can leave the model feature unusable. Preserve the last approved complete bundle while staging a candidate. A/B slots are one design, not a universal requirement or automatic protection. Android's system-update documentation illustrates explicit active, bootable, and successful states; an application model updater needs its own activation and recovery logic.[16]
For this terminal, design the model update around four questions:
- What is authorized? Authenticate a manifest against trusted keys. Bind hashes, lengths, model configuration, tokenizer, templates, modality assets, runtime requirements, device compatibility, version, and expiry to that signed metadata. A valid signature authenticates the signed bytes under the trust policy; it does not prove compilation, quality, or safety. Plan key rotation and revocation, version checks against replay, and freshness under the offline update policy. TUF specifies separate signature, version, expiry, and artifact checks.[17]
- What survives interruption? Download to inactive storage, with bounded size and safe paths. Check all artifacts against authenticated metadata before activation. Test truncated downloads, disk exhaustion, and power loss. Atomic pointer replacement alone does not guarantee crash durability; use the platform's durable-write protocol and recovery journal. A strictly isolated terminal needs an approved offline transfer rather than online OTA.
- What is compatible? Load the candidate on each supported device/runtime family. Check tokenizer/template behavior, reference outputs within declared tolerances, citations, abstention, memory peaks, and a representative offline workload. Five smoke cases cannot establish full task quality or ten-minute thermal behavior. Different weights need not have identical logits to the previous model.
- Who decides recovery? Activate only after required checks. Start a fresh session or use an explicitly supported transition; cached states from different weights/tokenizers are generally incompatible. Track health and restore a compatible approved bundle on failure. Avoid running two full models simultaneously if peak memory cannot support it. A rollback must respect anti-replay policy and schema compatibility; it is a controlled recovery action, not permission to install any old signed version.
Why must an on-device over-the-air model update verify both an Ed25519 cryptographic signature and a SHA-256 checksum?
Answer
A digest binds file bytes only when the expected digest comes from trusted metadata. A signature authenticates that metadata under trusted keys; it does not prove the model was compiled or evaluated correctly. Verify both, plus version, freshness, compatibility, and release evidence. An attacker can replace an unsigned file and its checksum together.
Diagnosing edge deployment breakdowns
These are illustrative failure scenarios, not measured devices. Treat symptoms as observations, then collect evidence that distinguishes causes:
1. The model draws 6W and runs at 5 tokens per second
Symptom: A 1B candidate runs at 35 tokens/s on the developer laptop but only 5 on the terminal, where whole-system power is 6 W and the enclosure feels hot.
Investigate: Check delegate placement, partition timing, synchronization, formats, clocks, and competing load. Compare supported accelerator and CPU paths with the same workload. Laptop throughput is not a mobile baseline.
Change after diagnosis: Choose a supported lowering or a different candidate. Replacing activations changes the model's computation and may require retraining; GELU or SiLU support is not universal. Recheck quality and energy as well as speed.
2. INT4 quantization destroys citation accuracy
Symptom: The unquantized FP16 model cites maintenance manual paragraphs with 96% accuracy, but the INT4 quantized model hallucinates non-existent section numbers.
Investigate: Compare tokenization, retrieval, template, sampling, conversion correctness, and paired reference outputs before attributing the regression to particular layers. A text-only decoder need not have cross-attention.
Change after diagnosis: Try a supported calibrated format, greater precision for a demonstrably sensitive tensor, or a different model. AWQ/GPTQ are candidates, not guaranteed repairs, and FP16 embeddings may break the memory budget. Repeat held-out source-support and abstention evaluation on the exact artifact.
3. Application crashes after fifteen minutes of continuous conversation
Symptom: The handheld app operates smoothly during brief queries, but terminates abruptly without an error log during prolonged diagnostic sessions.
Investigate: Inspect OS termination records, allocation peaks, retained sessions, cache reservations, and crash reports. Linear cache growth is a possibility for full attention, not proof of this failure. A watchdog, unsupported operation, or leak needs a different repair.
Change after diagnosis: Bound sessions and context, reduce a measured allocation, or select a suitable trained GQA/hybrid architecture. Cache quantization requires runtime support and quality checks. Arbitrarily evicting full-attention KV changes what the model can attend to; trimming history and rebuilding a prompt is different from native sliding-window attention.
4. Throughput collapses halfway through a work shift
Symptom: The model performs well in morning testing, but field technicians complain that answers slow to an unusable crawl during afternoon inspections.
Investigate: Correlate workload and rate with thermal state, ambient temperature, charging, background jobs, and power. A decline alone cannot identify the thermal governor as the cause.
Change after diagnosis: Evaluate smaller candidates, supported kernels, shorter outputs, or workload duty-cycle changes. Size alone does not ensure power below 3 W. A generated-token cap must leave room for the required answer and needs quality evaluation.
5. Over-the-air update bricks offline functionality
Symptom: Technicians update their handhelds before heading into the field, only to find the app crashes on launch when cell connectivity drops.
Investigate: Reproduce with networking disabled. Check missing files, lazy downloads, tokenizer/template mismatches, modality assets, authorization refresh, and incompatible runtime versions. A missing-token exception is only one possibility.
Change after diagnosis: Stage and authenticate a complete compatible bundle, test first launch offline, and exercise interrupted activation and recovery. Retain an approved offline recovery path.
Core takeaways for edge engineering
Deploying language models to the edge requires balancing physical hardware limits, architectural compactness, and software safety gates:
- Write the task contract: Compare lookup, classification, extraction, and generation against the same real job. Hypothetical sizing can reject candidates; it cannot qualify them.
- Separate training signals from truth: Soft targets expose teacher preferences, not guaranteed correctness. Exact forward-KL gradients have a 1/tau factor; the familiar squared scaling is a high-temperature approximation. Verifiers cover their contracts.
- Budget the actual architecture: Stored bytes, active weight traffic, KV/recurrent state, modality components, pools, and temporary allocations are different quantities. Effective parameter labels cannot establish fit.
- Inspect the installed runtime: A CPU graph round-trip does not prove mobile lowering. Partition boundaries can help or hurt; measure placement, output quality, and critical-path time.
- Measure a shift, not only a burst: Track visible-answer latency, generation work, memory peaks, thermal behavior, and energy for the whole product under representative conditions.
- Bind release and privacy decisions to real inputs: Enforce complete-payload transmission policy and authenticated, compatible, recoverable updates. A route string, local weights, or a valid signature alone cannot establish those properties.
Follow-up architectural questions
- Assume 34 GB/s and one active-weight stream per step. A 15-step/s target permits at most about 2.27 GB for that traffic before other costs. What evidence would invalidate this assumption, and why does it not select a quantization method by itself?
- An NPU delegate reports 98% operator coverage but misses latency. Which partition boundaries occur per generated token, and how will you separate synchronization, conversion, copies, and compute?
- Derive the exact forward-KL gradient and its high-temperature approximation. What changes when teacher and student token events cannot be directly aligned?
- A validly signed update changes the tokenizer and prompt template. What cache/session transition, offline test, and recovery evidence must precede activation?
Evaluation rubric
- Differentiate parameter count from memory bandwidth and thermal dissipation constraints.
- Derive the Hinton knowledge distillation loss and explain the mathematical role of temperature scaling and gradient compensation.
- Calculate KV cache memory savings from Grouped-Query Attention.
- Diagnose and resolve mobile NPU operator fallbacks and CPU memory thrashing.
- Implement fail-closed routing and cryptographic verification for edge model artifacts.