AI Fundamentals
A business leader's guide to the building blocks of AI — no code, no jargon, just the mental models that matter.
8 concepts 4 decision paths Diagrams
The building blocks
Click each concept to see the business analogy, what the thing actually is, and a visual explanation.
A base model is trained on trillions of tokens of text with one self-supervised objective: predict the next token. The task supplies its own answers — the target for each position is simply the token that follows, so no human labelling and no task-specific dataset are needed — and the billions of parameters are nudged by gradient descent, batch after batch, purely to get better at that prediction. Grammar, facts and reasoning patterns are what falls out of doing it well, across weeks of continuous training on thousands of accelerators.
A base model is trained on trillions of tokens of text with one self-supervised objective: predict the next token. The task supplies its own answers — the target for each position is simply the token that follows, so no human labelling and no task-specific dataset are needed — and the billions of parameters are nudged by gradient descent, batch after batch, purely to get better at that prediction. Grammar, facts and reasoning patterns are what falls out of doing it well, across weeks of continuous training on thousands of accelerators.
An embedding model converts a piece of text into a vector — a list of several hundred to several thousand numbers — positioned so that texts with similar meaning sit close together. Closeness is then arithmetic, usually cosine similarity between the two vectors. A vector database stores millions of them and answers "what is nearest to this?" by approximate nearest-neighbour search, which is why the lookup stays fast as the corpus grows.
Training continues from the pre-trained weights on a much smaller curated dataset — typically thousands of prompt-and-response pairs — at a low learning rate. The weights themselves change, so the behaviour is baked in and needs no prompt to invoke it. Push it too far and the model over-fits the narrow set and loses general ability: catastrophic forgetting. Preference-tuning methods can follow, teaching the model which of two answers is better rather than a single correct one.
Documents are split into chunks, embedded, and stored in a vector index. At query time the question is embedded too, the closest chunks are retrieved — often reranked by a second, more precise model — and pasted into the prompt as context before the model answers. The weights never change, so nothing is trained: quality comes from both halves — the retrieval pipeline (chunking, index, reranking) decides which passages the model sees, and the model decides how faithfully it uses them — and the answer can cite the chunk it came from.
The prompt is the model's entire conditioning at inference: a system instruction, any worked examples (few-shot), the retrieved context and the task itself, all as tokens in one finite context window. Nothing outside it exists to the model. Stating the reasoning steps to take, the output format to return, and one or two examples of a good answer changes the probability distribution the model samples from — no weight changes, no training run.
Generation is autoregressive: one forward pass through the network produces a probability distribution over the next token, one token is sampled from it, appended, and the whole thing runs again. It splits into two phases with different costs — prefill reads the prompt in parallel and is compute-bound, then decode emits tokens one at a time and is bound by memory bandwidth, which is why output tokens are slower and dearer than input ones. A KV cache keeps the earlier tokens' intermediate state so each step does not recompute the whole sequence.
Low-rank adaptation freezes every base weight and trains a pair of small matrices injected alongside the model's existing projections. Their product has the same shape as the weight it adjusts but a fraction of the parameters — commonly well under 1% — so training needs far less memory and produces an adapter file rather than a new model. Adapters can be merged into the weights for serving, or kept separate and switched per request on top of one shared base model in GPU memory.
Each weight is stored at lower numeric precision — 16-bit floating point down to 8- or 4-bit integers — with a scale factor per small group of weights to map the integers back to roughly their original values. Memory falls close to linearly with the bit width, and because decoding is memory-bandwidth bound, throughput improves as well. Accuracy loss stays small when the scheme is calibrated on representative data and the layers most sensitive to rounding are left at higher precision.
Decision framework
The decision ladder
Two different playbooks depending on whether you're building internal tools or shipping a product.
Internal: productivity tooling
Boosting employees, internal tools, knowledge management
- Free
- Low
- Occasional
API-based models work well — volume is low, value per query is high.
External: product integration
Embedding AI into products, pipelines, customer-facing systems
- Free
- Medium
- High
- Essential
API pricing is lethal at scale — self-hosted SLMs become the only viable path.
Frequently asked questions
What are the building blocks of AI a business leader actually needs to understand?
Eight, and none of them require code. Pre-training is the general education you buy rather than fund. Embeddings are how a model understands similarity. Prompting is the brief you write. RAG gives the model access to your data without retraining. Fine-tuning and LoRA layer domain expertise on top. Quantisation shrinks a model to fit cheaper hardware. Inference is the moment it does the work — and the line item you pay every month.
Should I use RAG or fine-tuning to give a model knowledge of my business?
Start with RAG in almost every case. RAG keeps your data out of the model weights, which makes it far easier to update, audit, and delete — drop a new document in and the knowledge is current immediately, with no retraining cost. Fine-tuning changes how a model behaves rather than what it knows, and it carries a real risk of over-specialisation. Reach for it only when the cheaper levers have genuinely failed.
Do we need to train our own model?
Almost certainly not. Pre-training a base model costs millions and takes months, and the result is a general graduate who still knows nothing about your business. You buy the graduate from someone who did it, then specialise them — with a better brief, a filing cabinet of your own documents, or a lightweight adapter. Training from scratch is a research programme, not a product decision.
What is quantisation and why does it matter commercially?
Quantisation reduces the numeric precision of a model — 32-bit down to 16, 8, or 4 — much as a RAW photo becomes a JPEG. The model gets dramatically smaller and faster with minimal quality loss. Commercially, that is the difference between an expensive multi-GPU server and a single consumer card: a 28 GB model becomes roughly 3.5 GB at 4-bit. It is the technique that makes self-hosting affordable.
Should internal tools and customer-facing products use the same AI approach?
No — the economics point in opposite directions. Internal productivity tooling is low-volume and high-value per query, so API-based models work well and the ladder rarely climbs past RAG. Product integration is the reverse: volume is high, value per query is low, and per-token API pricing becomes lethal at scale. That path runs through self-hosted small models, and usually ends at quantisation.
What is the cheapest lever to pull first?
A better brief. Prompt engineering is free, gives instant feedback, and resolves the large majority of enterprise use cases on its own — system prompts, a few worked examples, and asking the model to show its reasoning. Most complaints that "AI does not work" turn out to be vague-instruction problems, not model problems. Exhaust that before you spend anything.
Diagrams
Embed these freely — each SVG is licensed CC BY (opens in a new tab) with attribution to this page baked in.