🧰

Small Language Models & On-Device AI

When smaller and local beats bigger and cloud-hosted.

Part 1

The Multi-Tool You Carry Everywhere

why small language models trade some raw capability for something genuinely valuable in return: speed, low cost, and the ability to run almost anywhere.

Part 2

What Fits in Your Pocket

what parameter count and model size actually mean, and where the practical line between large and small language models tends to fall.

Part 3

Before There Was Anything to Carry

how language model deployment looked before small models existed as a genuinely practical, deliberately chosen option.

Part 4

The Workshop's Full Toolset

what large, frontier language models still offer that small models genuinely don't, and why that gap matters for certain tasks.

Part 5

Sharpening One Blade Really Well

how fine-tuning a small model for one narrow, well-defined task lets it genuinely match or exceed a much larger general-purpose model on that specific job.

Part 6

Fitting the Tool in Your Pocket

the core compression techniques — quantization, pruning, distillation — that pack genuine capability into a small model's compact footprint.

Part 7

The Apprentice Learning From the Master

how knowledge distillation trains a small model to mimic a larger model's behavior directly, transferring much of its capability into a far smaller footprint.

Part 8

Running Without a Trip to the Workshop

the real, practical benefits of on-device inference: lower latency, offline capability, and stronger privacy, made possible by small models.

Part 9

When the Multi-Tool Isn't Enough

the honest signals that a task has genuinely outgrown what a small model can reliably handle, and needs a larger model's full capability instead.

Part 10

The Cost of Carrying Less

the real, often dramatic cost and compute savings small language models offer, and how to weigh that savings honestly against their capability tradeoff.

Part 11

A Multi-Tool That Fits in Everyone's Pocket

how small language models are broadening genuine access to capable AI, beyond organizations that can afford significant cloud infrastructure.

Part 12

Testing the Blade Before You Trust It

how evaluating a small language model requires task-specific testing, not just comparing general benchmark scores against larger models.

Part 13

The Right Tool for the Job at Hand

a practical decision framework that pulls together every factor this series has covered, for choosing between a small and large language model on a real project.

Part 14

Two Tools Working Together

how routing between small and large models within one system lets an organization get the benefits of both, rather than committing to one size exclusively.

Part 15

Keeping the Blade Sharp Over Time

why a fine-tuned small model needs ongoing maintenance to stay reliable, as the tasks and data it handles evolve.

Part 16

Building Your Own Multi-Tool

what it takes for a team to fine-tune and deploy its own custom small model in-house, rather than relying entirely on a general-purpose vendor model.

Part 17

The Workshop's Trust in the Toolmaker

how keeping data local, on-device or within an organization's own infrastructure, offers genuine privacy and security advantages small models make practical.

Part 18

A Pocket Tool for Every Craftsperson

the practical considerations for deploying a small model consistently across many devices at scale, not just a single proof of concept.

Part 19

Keeping the Whole Toolkit in Working Order

the sustained operational discipline required to run small language models reliably in production over time, extending the practices covered in this content library's LLMOps series.

Part 20

The Full Toolkit, Ready to Carry

reassembling every piece covered across this series into the complete picture of how small language models earn their place alongside frontier models.