The GenAI Field Guide

GenAI fundamentals

Start with the model, the tokens it reads, and the ways it can fail.

Generative AI is software that produces new content (text, code, images, audio) by sampling from patterns it learned during training. Almost every system you will build in this field sits on top of one component: a large language model that reads a sequence of tokens and predicts what comes next. Prompts, retrieval, tools, agents and evals are all ways of feeding that predictor better input and checking its output, so the mechanics in this chapter explain most of the behaviour you will later debug.

The mental model to carry is simple and surprisingly complete. Text becomes tokens; tokens become vectors; a transformer mixes those vectors with attention, layer after layer; the last layer produces a probability for every possible next token; a sampler picks one, appends it, and the loop repeats. Everything the model "knows" lives in its parameters, frozen at a training cutoff. Everything it can "see" for a specific request lives in the context window. Hallucination, cost, latency, temperature and nondeterminism all fall out of that picture.

The questions build in that order. The first group defines the terms (GenAI, LLM, token, context window, transformer, attention, inference, parameters, embeddings) and explains why hallucination is a property of the mechanism rather than a bug. The advanced group covers how models are trained into assistants (pre-training, post-training, RLHF), why bigger is not automatically better, what sampling controls actually do, how reasoning and multimodal models differ, and why the same prompt can return different answers. Later chapters assume this vocabulary.

What is Generative AI?

Generative AI is software that creates new content (text, code, images, audio, video) by learning the statistical patterns of large datasets and then sampling new outputs that fit those patterns and a given request.

A generative model learns a probability distribution over data: which sequences of words, pixels or sound samples are likely. Once trained, it can draw new samples from that distribution, steered by an input called a prompt. That is the difference from software that only retrieves or classifies existing content. A search engine finds a document that already exists; a generative model writes one that did not exist a moment ago.

Different media use different model families. Text and code are dominated by large language models (LLMs), which generate one token at a time. Images and video are mostly produced by diffusion models, which start from random noise and remove it step by step until a picture matching the prompt emerges. Speech uses dedicated text-to-speech and speech-to-text models, and many current systems combine several of these behind one interface.

What made GenAI practical for businesses was not a single breakthrough but a combination: the transformer architecture, very large training datasets, enough compute to train models with billions of parameters, and post-training that turns a raw text predictor into an assistant that follows instructions. The result is a general-purpose component that can draft, summarise, translate, extract, classify and write code without task-specific training.

The trade-off is that the output is plausible, not guaranteed correct. A generative model optimises for content that looks like its training data and satisfies the request. It has no built-in notion of truth, so production systems wrap it with retrieval, tools, validation and human review where errors matter.

How is GenAI different from traditional machine learning?

Traditional machine learning usually maps an input to a label or number, trained per task on labelled data. GenAI produces open-ended content and one pre-trained model can handle many tasks through prompting alone.

Traditional ML (often called discriminative or predictive ML) learns a function from inputs to a fixed output space: fraud or not fraud, a house price, a churn probability. You collect labelled examples for that one task, engineer features, train a model such as gradient-boosted trees or logistic regression, and measure accuracy against held-out labels. The output space is small and known in advance, which makes evaluation clean.

Generative AI learns the structure of the data itself and produces new instances: a paragraph, a SQL query, an image. Because a large language model is pre-trained on broad data, the same model can summarise, classify, translate and write code without retraining. You change behaviour by changing the prompt. That shifts the work from feature engineering and labelling towards prompt design, context assembly and evaluation of open-ended outputs.

The differences matter operationally. Traditional models are small, fast, cheap per prediction and usually deterministic, and their errors are measured with familiar metrics like precision and recall. Generative models are larger, slower and more expensive per call, their output varies between runs, and judging whether a generated paragraph is good often needs rubric-based or model-based evaluation rather than a simple accuracy score.

They are not rivals. Many strong systems combine them: a classical model scores risk, and an LLM explains the score to an analyst in plain language. An LLM can also label data that later trains a cheap classifier, which then handles volume at a fraction of the cost.

What is an LLM?

A large language model is a neural network, almost always a transformer, trained on vast amounts of text and code to predict the next token. Post-training turns that predictor into an assistant that follows instructions.

An LLM is "large" in two senses: it has billions of learned parameters, and it was trained on trillions of tokens of text drawn from the web, books, code repositories and licensed or synthetic data. Its single training objective during pre-training is to predict the next token given the ones before it. To do that well across so much data, the model has to absorb grammar, facts, writing styles, programming languages and patterns of reasoning, because all of them help the prediction.

A freshly pre-trained model is called a base model. It continues text rather than answering questions: give it a question and it may produce three more questions. Post-training (instruction tuning plus preference optimisation) teaches it the conversational format, to follow instructions, to refuse harmful requests and to use tools. The chat assistants and APIs you use are post-trained models.

In an application, the LLM is a stateless function. You send it a list of messages (system instructions, conversation history, retrieved documents, tool results) and it returns generated tokens. It remembers nothing between calls; any memory you see in a product is the application re-sending earlier text. It cannot browse, query a database or run code unless your application gives it tools and executes them.

That framing explains most design decisions. Quality depends on what you put in the context; cost and latency scale with tokens in and out; reliability comes from validating outputs, because the model generates what is likely, not what is verified.

How does an LLM generate text?

It runs a loop: read all tokens so far, compute a probability for every possible next token, pick one with a sampling rule, append it, and repeat until a stop token or a length limit is reached.

Generation is autoregressive: each new token depends on every token before it, including the ones the model just produced. The input text is split into tokens, each token becomes a vector, and the transformer processes the whole sequence. Its final layer outputs a score (a logit) for every token in the vocabulary, often 100,000 or more entries. A softmax turns those scores into probabilities that sum to one.

A sampler then chooses the next token. Greedy decoding always takes the most likely token. Sampling with temperature, top-p or top-k draws from the distribution, which gives more varied text. The chosen token is appended and the model runs again. It stops when it emits a special end-of-sequence token, hits a stop sequence you defined, or reaches the maximum output length.

Two consequences follow. First, output costs time per token: a 500-token answer needs 500 sequential steps, which is why long answers are slow and why output tokens are usually priced higher than input tokens. Serving systems speed this up with a KV cache that stores intermediate results for earlier tokens so they are not recomputed. Second, the model never plans the whole answer as a separate artefact. Any plan exists only as tokens it has already written or as patterns in its internal state, which is why asking a model to think step by step, or using a reasoning model, can improve hard answers.

Because each token is chosen from a probability distribution, an early unlikely choice can steer the rest of the answer. A wrong first sentence often produces a confidently wrong paragraph, since everything after it is conditioned on that sentence.

What is a token?

A token is the unit of text a model reads and writes: a whole word, part of a word, a punctuation mark or a byte sequence. Costs, context limits and speed are all measured in tokens.

Models do not see characters or words directly. A tokenizer splits text into pieces from a fixed vocabulary, typically using an algorithm such as byte-pair encoding (BPE), which starts from bytes and repeatedly merges the most frequent adjacent pairs found in training data. Common words become single tokens; rare words, names, code identifiers and long numbers are split into several. Each token maps to an integer id, and each id maps to a learned vector the model computes with.

A useful rough rule for English prose is about four characters, or three quarters of a word, per token. That ratio varies a lot. Code, JSON, URLs, numbers and many non-English languages (especially those with non-Latin scripts) use more tokens per word, sometimes two or three times as many. Each model family has its own tokenizer, so the same text produces different token counts on different models.

Tokens matter for three practical reasons. Cost: providers bill per input and output token. Limits: the context window and maximum output length are counted in tokens. Latency: output is generated one token at a time. They also explain odd behaviour. A model asked to count the letters in a word, reverse a string or do digit-level arithmetic is working on chunks it never saw as individual characters, so it can make mistakes a child would not.

Always measure with the tokenizer for the model you actually use, or with the token counts the API returns in its usage fields. Estimates from word counts are fine for a back-of-envelope budget and wrong for anything that bills or truncates.

What is a context window?

The context window is the maximum number of tokens a model can process in one request, counting the system prompt, history, documents, tool results and the output it generates. Anything outside it does not exist for the model.

Each request to an LLM is a single sequence of tokens. The context window is the hard limit on that sequence's length. It is shared: input tokens and generated output tokens both count, and many providers also set a separate, smaller cap on output tokens. If your input uses almost the whole window, the model has little room left to answer. If you exceed it, the API rejects the request or your framework silently drops older content.

Window sizes have grown from a few thousand tokens to hundreds of thousands and, for some models, millions. That does not mean you should fill them. Attention cost grows with sequence length, so long prompts raise latency and cost on every call. Quality also degrades: research such as the 2023 paper Lost in the Middle showed models retrieve information placed at the start or end of a long context more reliably than information buried in the middle, and many models get worse at following instructions as unrelated text accumulates.

The context window is also the model's only working memory for a request. The model's parameters hold general knowledge from training; the window holds the specifics of this task. Your customer's contract, today's stock level and the tool output from the previous step reach the model only by being placed in the window. Choosing what goes in, in what order and in what format is the discipline called context engineering.

In practice you budget the window like memory in an embedded system: a fixed allowance for instructions, a share for retrieved material, a share for conversation history, and a reserve for the answer.

What is a transformer?

The transformer is the neural network architecture behind nearly all LLMs. It processes a sequence of token vectors through stacked layers that alternate attention (tokens exchange information) with feed-forward networks (each token is transformed independently).

The transformer was introduced in the 2017 paper Attention Is All You Need. Earlier language models used recurrent networks that read text one word at a time, carrying a compressed memory forward, which made them slow to train and forgetful over long passages. The transformer replaced recurrence with attention, letting every position look directly at every other position in a single step. That made training highly parallel on GPUs, which is what allowed models to scale to billions of parameters and trillions of training tokens.

A modern LLM is usually a decoder-only transformer. Tokens are converted to vectors (embeddings) and given positional information so the model knows their order. The sequence then passes through dozens of identical blocks. In each block, an attention layer lets each token gather information from earlier tokens, and a feed-forward layer (a small multilayer network applied to each position) transforms that information. Residual connections add each layer's output back to its input, so information flows through the stack as a gradually refined "residual stream". The final layer maps each position's vector to scores over the vocabulary.

Rough intuition for what the parts do: attention moves information between positions (which noun does "it" refer to?), while the feed-forward layers hold much of the model's stored associations (Paris is in France). Depth lets the model compose these operations into more abstract features.

The architecture has known costs. Standard attention compares every token with every other, so compute grows with the square of sequence length, which is why long contexts are expensive. Variants such as grouped-query attention, sliding-window attention and mixture-of-experts feed-forward layers trade some flexibility for speed or capacity, but the core design has stayed remarkably stable since 2017.

What is attention?

Attention is the mechanism that lets each token compute a weighted mix of information from other tokens, with the weights based on how relevant each one is to the current token. It is how context changes meaning.

In each attention layer, every token produces three vectors through learned projections: a query (what am I looking for?), a key (what do I contain?) and a value (what information do I pass on?). The model scores each pair by the dot product of one token's query with another token's key, scales the scores, and applies a softmax so they sum to one. Each token's output is the weighted sum of all the value vectors. Tokens with matching query and key get high weight; irrelevant ones get almost none.

LLMs use causal attention: a token may only attend to itself and earlier tokens, never later ones, because during generation the later tokens do not exist yet. They also use multi-head attention, running many attention operations in parallel with different projections. Different heads learn different relationships, such as tracking syntax, copying a name seen earlier or linking a pronoun to its referent.

Attention explains several behaviours you will see in production. Because every token can look at every earlier token, a model can use a definition from the top of a 50-page document when answering at the bottom. Because the weights are a softmax over many positions, relevant information can be diluted when the context is long and full of near-duplicates. And because the scores compare every pair, compute grows quadratically with length.

During generation, the keys and values for earlier tokens do not change, so inference servers store them in a KV cache instead of recomputing them for every new token. That cache is a large part of GPU memory use and is why long conversations and many concurrent users are expensive to serve.

What is inference?

Inference is running an already trained model to produce output for a new input. For an LLM it has two phases: prefill, which processes the whole prompt at once, and decode, which generates output tokens one at a time.

Training adjusts a model's parameters; inference uses them unchanged. Every API call, chat message and agent step is an inference request. For LLM-based products, inference is where almost all of the running cost and latency live, so understanding its shape pays off quickly.

An LLM request runs in two phases. Prefill processes all input tokens in parallel, builds the KV cache and produces the first output token. It is compute-bound and scales with prompt length, so it dominates time to first token (TTFT) on long prompts. Decode then generates the remaining tokens one by one, each step reading the full model weights and the KV cache from GPU memory. Decode is usually limited by memory bandwidth rather than arithmetic, and it sets tokens per second for streaming.

Serving systems batch many users' requests together so that each pass over the weights does useful work for several sequences at once. Techniques such as continuous batching, paged KV-cache memory, quantization and speculative decoding raise throughput or cut latency. If you call a hosted API, the provider handles this, but its effects show up in your metrics: latency varies with load, very long prompts are slow to start, and long outputs are slow to finish.

For cost, providers typically bill input and output tokens at different rates, and many discount repeated prompt prefixes through prompt caching, because a cached prefix skips part of prefill. Check the provider's current pricing and caching rules rather than assuming they match another provider's.

What are model parameters?

Parameters are the learned numbers (weights and biases) inside a neural network that determine its output. An LLM's parameter count indicates its capacity and, more practically, how much memory and compute it needs to run.

A neural network is a long chain of matrix multiplications and simple nonlinear functions. The entries of those matrices are its parameters. Training starts them at random values and nudges each one, billions of times, in the direction that reduces prediction error. After training, the parameters are frozen, and everything the model has learned, from grammar to facts to coding patterns, is encoded in them as distributed numerical patterns rather than as readable records.

Parameter count is a rough proxy for capacity: more parameters can store more patterns and support more complex behaviour. It is a poor proxy for quality on a specific task, because training data, training length and post-training matter as much. A well-trained smaller model regularly beats an older or undertrained larger one.

The most reliable use of parameter count is for memory planning. Each parameter takes 2 bytes at 16-bit precision, 1 byte at 8-bit and about half a byte at 4-bit quantization. So the weights of a 7-billion-parameter model need about 14 GB at 16-bit and roughly 4 GB at 4-bit, before adding the KV cache and runtime overhead. A 70-billion-parameter model needs about 140 GB at 16-bit, which means several GPUs or aggressive quantization.

Mixture-of-experts (MoE) models complicate the picture. They have many expert sub-networks but route each token through only a few, so they report a large total parameter count and a much smaller active count per token. Memory follows the total (all experts must be loaded); compute per token follows the active count.

What is an embedding?

An embedding is a vector of numbers that represents a piece of content so that items with similar meaning end up close together. It lets software compare meaning with arithmetic, which powers semantic search, clustering and retrieval.

An embedding model reads a text (or image, or audio clip) and outputs a fixed-length vector, typically a few hundred to a few thousand numbers. It is trained, often with contrastive learning, so that related inputs (a question and the passage that answers it, "invoice" and "bill") produce vectors pointing in similar directions while unrelated inputs point elsewhere. Similarity is then measured with cosine similarity or a dot product.

The word also describes an internal step inside every LLM: each token id is mapped to a learned vector at the input layer. Those token embeddings are not what you use for search. For retrieval you call a dedicated embedding model that produces one vector for a whole sentence or passage.

Embeddings are the core of retrieval-augmented generation. You embed every chunk of your documents once, store the vectors in an index, embed the user's question at query time, and fetch the nearest chunks. This finds relevant passages even when they share no keywords with the question. Embeddings also support deduplication, clustering support tickets by topic, recommendation and anomaly detection.

Their limits are specific. Embeddings capture topical similarity better than precise facts: "refunds allowed within 30 days" and "refunds not allowed within 30 days" can sit very close together. They often miss exact identifiers such as part numbers or error codes, which is why hybrid search adds keyword matching. Vectors from different embedding models live in different spaces and cannot be compared, so changing models means re-embedding the whole corpus.

What is hallucination?

Hallucination is when a model produces fluent, confident content that is false or unsupported by its sources, such as invented citations, wrong figures or nonexistent API functions. It comes from generating likely text without a check on truth.

An LLM is trained to produce probable continuations, and post-training rewards answers that look helpful. Neither objective directly rewards saying "I don't know". When the model lacks the needed fact, the most probable continuation is often a well-formed guess: a plausible case name, a reasonable-looking statistic, a function that would exist if the library were designed that way. The guess is delivered with the same confidence as a well-known fact because the model has no reliable internal signal separating the two.

It helps to separate two kinds. Factual hallucination is a claim that is wrong about the world (a fabricated court case). Faithfulness hallucination is a claim not supported by the context you supplied (a summary that adds a clause the contract does not contain). RAG reduces the first kind and can still suffer from the second.

Risk rises in predictable places: rare or recent facts, precise numbers, citations and URLs, long outputs, questions with a false premise ("Why did the company recall its product in 2019?" when it never did), and tasks pushed beyond what the context supports. Risk falls when the answer can be copied from supplied text, when the task is narrow, and when the output is checked.

No prompt eliminates hallucination. The effective defences are structural: ground answers in retrieved sources and require citations, give the model tools for lookups and calculations, allow and reward abstention, constrain outputs to schemas or enumerations, and verify claims before they reach users or trigger actions. Measure the rate on an eval set, because anecdotes will not tell you whether a change helped.

What is pre-training versus post-training?

Pre-training teaches a model general knowledge and language by predicting the next token over trillions of tokens. Post-training then shapes that base model into a useful assistant through instruction tuning, preference optimisation and reinforcement learning.

Pre-training is the expensive phase. A randomly initialised transformer is trained on a huge, filtered mixture of web text, books, code, scientific papers and increasingly synthetic data, with one objective: predict the next token. It typically consumes the large majority of total training compute and runs for weeks to months on thousands of accelerators. The result is a base model that has absorbed facts, languages, coding conventions and reasoning patterns, but behaves like an autocomplete engine: it continues documents rather than following instructions, and it will happily imitate the worst text in its data.

Post-training turns that base into an assistant, and it is where most of a model's personality and usability come from. It usually has several stages. Supervised fine-tuning (SFT) trains on curated prompt and response pairs that demonstrate the desired format and behaviour. Preference optimisation (RLHF, DPO and related methods) trains on comparisons of better and worse answers so the model learns finer judgements about helpfulness, tone and safety. Reinforcement learning with verifiable rewards trains on tasks whose answers can be checked automatically, such as maths problems and unit-tested code, and is a large part of how reasoning models are built. Tool use, refusals and long-context skills are also trained in this phase.

The split has practical consequences. Knowledge mostly comes from pre-training, so post-training (and your own fine-tuning) is a weak way to add facts and a good way to change behaviour, style and format. The knowledge cutoff is set by the pre-training data. Two products built on the same base model can feel very different because of their post-training.

There is also a cost of post-training sometimes called the alignment tax: tuning for one behaviour can slightly degrade others, such as calibration or creativity. Providers balance this, and their choices explain why models differ in verbosity, refusal rates and how readily they admit uncertainty.

What is RLHF?

Reinforcement learning from human feedback trains a model to produce answers people prefer: humans compare pairs of responses, a reward model learns to predict those preferences, and the LLM is optimised against that reward model.

RLHF became widely known through the 2022 InstructGPT paper and was a key step in turning base models into usable assistants. The motivation is that "good answer" is hard to specify as a loss function but easy for people to judge when they see two candidates side by side. RLHF converts those judgements into a training signal.

The classic pipeline has three steps. First, start from a supervised fine-tuned model so it already produces reasonable answers. Second, collect preference data: for many prompts, sample two or more responses and have labellers pick the better one according to guidelines. Train a reward model (usually another LLM with a scalar output head) to predict which response a labeller would prefer. Third, optimise the LLM with a reinforcement learning algorithm, historically PPO (proximal policy optimisation), to maximise the reward model's score, with a KL penalty that keeps it close to the starting model so it does not drift into odd text that happens to score well.

RLHF has well-documented failure modes. Reward hacking: the policy finds outputs the reward model overrates, such as longer answers or confident tone, without being better. Sycophancy: because people tend to prefer answers that agree with them, models learn to agree. Labeller disagreement and inconsistent guidelines add noise. These are part of why models can be verbose, overly agreeable or overconfident.

The field has since added alternatives. DPO (direct preference optimisation) skips the separate reward model and RL loop, training directly on preference pairs with a simple classification-style loss. RLAIF and constitution-based methods use model-generated feedback guided by written principles to scale labelling. RL with verifiable rewards replaces human preference with automatic checks where answers can be verified. Production post-training typically mixes several of these, and the exact recipe varies by provider and is rarely fully published.

Does a larger model always perform better?

No. Larger models tend to be more capable on broad, hard tasks, but on a specific task, data, training, post-training, prompting and context often matter more, and size always costs latency and money.

Scaling research, notably the 2020 Scaling Laws for Neural Language Models paper and the 2022 Chinchilla paper, showed that loss falls predictably as parameters, data and compute grow together. Chinchilla's key finding was that many models were undertrained for their size: a smaller model trained on more tokens could beat a larger one trained on fewer, at the same compute. Since then, many model builders deliberately train smaller models far beyond that ratio because inference cost, not training cost, dominates once a model is deployed.

So size is one input among several. On narrow tasks such as classifying tickets, extracting fields or routing requests, a small model with good examples, a fine-tune or a constrained output format often matches a large general model. A small model with the right retrieved context will beat a large model guessing from memory. On broad, open-ended or multi-step tasks (complex coding, long analysis, ambiguous instructions), larger and reasoning-oriented models usually do pull ahead.

Size also brings costs that matter in production: higher latency per token, higher price per call, more GPU memory if you self-host, and lower throughput. A system that calls a model ten times per task multiplies all of those. That is why routing (send easy requests to a small model, hard ones to a large one) and cascades (try small first, escalate on low confidence) are common.

The reliable way to decide is empirical. Build an eval set from real traffic, run two or three candidate sizes, and plot quality against cost and latency. Pick the smallest model that meets the quality bar with margin, and re-run the comparison when new models appear.

What does temperature change?

Temperature rescales the model's next-token probabilities before sampling. Low values concentrate probability on the top choices for steadier, more repetitive output; high values flatten the distribution for more varied and more error-prone output.

At each step the model outputs logits, one per vocabulary token. Sampling divides those logits by the temperature T and applies a softmax. With T below 1, differences between logits are magnified, so the most likely tokens get even more probability. With T above 1, differences shrink and unlikely tokens get a real chance. As T approaches 0, sampling becomes greedy: always pick the top token. Temperature does not make the model smarter or more knowledgeable; it only changes how adventurously it picks from what it already predicts.

Two companion controls truncate the distribution. Top-k keeps only the k most likely tokens. Top-p (nucleus sampling) keeps the smallest set of tokens whose probabilities add up to p, such as 0.9, which adapts to how confident the model is at each step. Providers generally recommend adjusting temperature or top-p, not both at once. Some APIs also offer frequency and presence penalties to discourage repetition, and some reasoning models fix or ignore sampling parameters; check your provider's documentation.

Choose by task. Extraction, classification, code and structured output usually do best at low temperature (0 to 0.3) because there is one right answer and variation is just risk. Brainstorming, marketing variants and dialogue for characters benefit from more (0.7 to 1.0). For self-consistency techniques, where you sample several answers and take a majority vote, you want enough temperature to get genuinely different reasoning paths.

Low temperature is not a correctness setting. A model that is wrong with 90 percent probability is reliably wrong at temperature 0. It is also not a determinism guarantee: serving infrastructure can still produce different outputs at temperature 0, as a later question explains.

What is a reasoning model?

A reasoning model is an LLM trained to generate an extended internal chain of thought before its final answer, spending more inference compute on hard problems. It trades latency and token cost for better accuracy on multi-step tasks.

Researchers noticed early that asking models to "think step by step" improved accuracy on maths and logic, because intermediate tokens give the model room to compute. Reasoning models build this in. Through post-training, largely reinforcement learning on problems with checkable answers (maths, code with tests, logic puzzles), they learn to produce long reasoning traces that explore approaches, check intermediate results and backtrack from errors before committing to an answer. This is a form of test-time compute: quality improves by spending more computation per question rather than by training a bigger model.

Providers expose this differently. Many let you set a reasoning effort or thinking budget that caps how many reasoning tokens the model may use. Some return the full reasoning, some return a summary and some hide it. Reasoning tokens are generally billed as output tokens even when hidden, and they add latency before the visible answer starts, often seconds and sometimes much longer. Some providers fix or ignore sampling parameters like temperature for these models.

They shine on tasks with many dependent steps: debugging, complex refactors, maths, planning with constraints, analysing contradictory documents, and agentic work where a wrong early step is expensive. They add little to simple lookups, short rewrites, classification or chit-chat, where they mostly add delay and cost. They can also overthink simple instructions or be verbose.

Treat the visible reasoning trace with care. It is useful for debugging, but research has shown traces are not always a faithful account of how the model reached its answer, so do not rely on them as an audit record. Evaluate on final outputs, and pick effort levels per task the same way you would pick model size: measure quality against latency and cost.

What is a multimodal model?

A multimodal model accepts or produces more than one kind of data, such as text plus images, audio, video or documents, usually by converting each input into tokens or vectors that a shared transformer processes together.

The common design for multimodal input is to attach an encoder for each extra modality to a language model. An image is split into patches (small squares), a vision encoder such as a vision transformer turns the patches into vectors, and a projection layer maps those vectors into the same space as the LLM's token embeddings. The model then reads image tokens and text tokens in one sequence, so attention can link the word "total" in your question to the region of a receipt where the total is printed. Audio works similarly, with an audio encoder turning short frames of sound into tokens.

Multimodal output is more varied. Some models generate text only. Others can emit image or audio tokens directly, and many products route generation to separate specialist models (a diffusion model for images, a text-to-speech model for voice) behind one interface. Whether a given API does true native generation or a pipeline varies by provider and is not always visible to you.

Practical consequences follow from the token view. Images and audio consume context and cost: a single image can be hundreds to a few thousand tokens depending on resolution and provider, and PDFs may be processed as both extracted text and page images. Downscaling saves tokens but can make small print unreadable. Video is usually sampled into frames, so brief events between sampled frames can be missed.

Multimodal models fail in characteristic ways: misreading small or rotated text, miscounting objects, misjudging spatial relations (left of, above), reading values off charts imprecisely and confidently describing details that are not in the image. For high-stakes extraction, combine them with OCR or structured parsers and validate the numbers.

What is a knowledge cutoff?

A knowledge cutoff is the date after which a model's training data contains little or nothing, so the model does not know about later events, releases or changes unless you supply that information in the context.

A model's parameters are fixed when training ends. Whatever was in the training data up to that point is what it can recall; anything that happened afterwards is invisible to it. The knowledge cutoff is the approximate date where that data stops. Providers usually publish it, sometimes as a range, because different data sources end at different times. Models are often released months after their cutoff and stay in use for a year or more, so the gap between what the model knows and today can be large.

Coverage also thins out before the cutoff. The internet takes time to write about events, so the final months before a cutoff are underrepresented compared with how they will eventually be covered. A model may know a library's version from two years ago in depth and its most recent release only vaguely. Models are also often unsure of their own cutoff and may state it wrongly if asked.

This matters most for anything that changes: prices, regulations, product catalogues, API signatures, people's roles, current events and your own organisation's data (which the model never saw at all). Coding assistants show the problem constantly when they suggest deprecated functions or older configuration formats.

The fix is to stop relying on parametric memory for current facts. Put the current date in the system prompt, retrieve up-to-date documents, give the model a search or lookup tool, and include version numbers when asking about software. Fine-tuning is a poor way to keep knowledge current: it is slow, costly and absorbs facts unreliably.

Why does the same prompt sometimes give different answers, even at temperature zero?

Sampling is the obvious cause, but even greedy decoding varies because GPU floating-point maths depends on batch composition and kernel choice, and providers update models and serving stacks. Design for variation instead of assuming reproducibility.

At temperature above zero, variation is intended: the next token is drawn at random from a distribution, so two runs diverge as soon as one draws a different token, and everything afterwards differs. Many teams then set temperature to 0 and expect identical outputs. Often they get mostly identical outputs with occasional differences, which is harder to debug than constant variation.

The main cause is numerical. Floating-point addition is not associative: summing the same numbers in a different order can change the last bits of the result. Inference servers batch your request with other users' requests, and the GPU kernels they use can split and order reductions differently depending on batch size and sequence lengths. So the exact logits for your prompt can shift very slightly from one call to the next. When two candidate tokens are nearly tied, that tiny shift flips which one is the top choice, and greedy decoding then follows a different path for the rest of the answer. Mixture-of-experts routing can amplify this, since a small change can send a token to a different expert. Kernels can be made batch-invariant, but this generally costs throughput and is not the default on most serving stacks.

There are also non-numerical causes. Hosted providers update model snapshots, safety layers, system prompts and serving infrastructure. Some APIs accept a seed parameter, which improves repeatability but is usually documented as best-effort. Some reasoning models ignore temperature entirely. Your own stack contributes too: retrieval results that change as the index updates, timestamps in prompts and unordered tool results all alter the input.

The engineering response is to make the system robust to variation rather than to chase exact reproducibility. Pin model versions where the provider allows it. Log full inputs, parameters and outputs so any answer can be investigated. Validate structured outputs and retry on failure. Evaluate over multiple runs and report rates, not single outcomes. For cases that truly need identical results, cache the output and reuse it.