The GenAI Field Guide

Knowledge and reasoning

Separate learned patterns, supplied facts, and verified calculation.

Every answer an LLM gives combines two things: information it draws on and work it does with that information. The information can come from its weights, which are broad, frozen at a training cutoff and impossible to audit, or from the context you supply, which is narrow, current and citable. The work can be done in the model's tokens, which is flexible but approximate, or handed to code, which is exact. Most reliability problems in GenAI systems come from putting a job in the wrong place: asking the weights for a fact that should be retrieved, or asking tokens for a calculation that should be computed.

The mental model for this chapter is a division of labour. Let the model interpret language, choose what to do and explain results. Let retrieval and tools supply facts. Let code, solvers and rules perform anything that must be exactly right. Spend extra answer-time compute where problems are hard and checkable, and verify outputs against evidence rather than against the model's own confidence or the fluency of its reasoning.

The basic questions separate knowledge from reasoning, explain what weights actually store and why models stumble on arithmetic and recent facts, and introduce tools and symbolic reasoning as the remedies. The advanced questions cover test-time compute and inference-time scaling, self-verification and second-model verification, neuro-symbolic designs, how far to trust a reasoning trace, what happens when context contradicts memory, and how to test whether a model reasons or recalls.

What is the difference between knowledge and reasoning in an LLM?

Knowledge is the information a model can draw on, either memorised in its weights or supplied in the context. Reasoning is the process of combining that information through intermediate steps to reach a conclusion the inputs did not state directly.

The distinction matters because the two fail differently and are fixed differently. A knowledge failure means the model lacked the fact or recalled it wrongly: it cites a repealed tax rate or invents a policy clause. A reasoning failure means the facts were right but the steps were wrong: it applied the right rate to the wrong subtotal, or skipped a condition. Retrieval, tools and fresher data fix the first. Decomposition, more thinking time, code execution and verification fix the second.

Inside the model the line is blurry. There is no separate fact store and reasoning engine; both are behaviours of the same network predicting the next token. A model can appear to reason when it is actually recalling a solved example it saw in training, and it can appear to know a fact when it is actually inferring a plausible answer from related patterns. That is why a problem rephrased with new numbers can suddenly fail: the familiar version was partly memorised.

For a builder the useful split is by source, not by mental faculty. Parametric knowledge (in the weights) is broad, frozen at a training cutoff and unverifiable. Contextual knowledge (in the prompt, from retrieval or tools) is narrow, current and citable. Reasoning is the work done on top of either, and it becomes more reliable when each step is short, checkable and grounded in supplied text rather than recall.

When an answer is wrong, diagnose which part broke before changing anything. Give the model the correct facts in context and ask again. If it now succeeds, you had a knowledge problem. If it still fails with the facts in front of it, you have a reasoning problem, and adding more documents will not help.

What is stored in model weights?

Weights store billions of numeric parameters that encode statistical patterns: how words, concepts and facts tend to relate. Facts are spread across many parameters as associations, not kept as rows you can look up, list, edit or delete.

A model's weights (also called parameters) are the numbers adjusted during training so that the network predicts the next token well. Training on trillions of tokens pushes those numbers to encode whatever helps prediction: grammar, style, code idioms, common facts, and reusable procedures such as how to format a date or structure an argument.

Facts are stored distributed and associatively. Interpretability research, such as the 2021 paper Transformer Feed-Forward Layers Are Key-Value Memories, suggests that the feed-forward layers act partly like soft lookup tables: a pattern in the input activates a direction that promotes certain output tokens. A fact like a company's headquarters city is therefore not one cell. It is a tendency spread over many parameters, entangled with other facts that share those parameters.

This has practical consequences. First, recall depends on frequency: facts seen thousands of times in training are reliable, while long-tail facts seen a handful of times are recalled poorly or blended with similar ones. Second, recall is direction-sensitive: the 2023 Reversal Curse paper showed models trained on "A is B" often cannot answer "what is B?". Third, there is no index. You cannot list what a model knows, check where a claim came from, or reliably remove one fact without side effects.

Weights are also frozen at a training cutoff. Anything that changed afterwards, such as prices, laws, staff or product versions, is either missing or stale, and the model may not know which. That is the main reason production systems keep authoritative facts outside the model and feed them in at request time.

What does it mean for an LLM to “know” something?

An LLM "knows" a fact when it reproduces it reliably across phrasings and contexts. That is a behavioural, probabilistic property, not a guarantee: the same model can state a fact correctly in one prompt and invent a wrong one in the next.

In practice "knowing" is best defined by consistency under variation. If you ask the same question five different ways, in different languages and embedded in different tasks, and the answer stays correct, the knowledge is robust. If small changes in phrasing flip the answer, the model has a fragile association rather than knowledge you can depend on.

A model does not have a reliable sense of its own knowledge boundaries. It produces the most plausible continuation whether or not the underlying fact was well learned, so a correct answer and a hallucination often look identical in tone. Research on calibration shows models carry some signal about their uncertainty (token probabilities and agreement across samples correlate with accuracy), but post-training for helpfulness can weaken that signal, and the verbal confidence a model states is a poor proxy.

There is also a difference between recalling and using. A model may recite a rule correctly when asked directly and still fail to apply it inside a longer task, because the cue that triggered recall is absent. Conversely, a model given a document in its context "knows" its contents only for that request; nothing persists afterwards unless your application stores it.

For high-stakes facts such as legal citations, drug interactions or account balances, treat model output as a lead to check, not a fact. The useful question is not "does the model know this?" but "what evidence will this answer rest on, and can I verify it automatically?"

Why do models struggle with arithmetic?

Because they predict text token by token rather than run an arithmetic algorithm. Numbers are split into irregular tokens, digits are produced left to right while carries flow right to left, and nothing forces an exact result, so long or unusual calculations drift.

Several mechanisms combine. Tokenisation splits numbers in ways that ignore place value: depending on the tokenizer, 8490 might be one token, two tokens such as "84" and "90", or single digits. The model must learn arithmetic over these irregular chunks, which is much harder than over aligned digits.

Generation order works against the algorithm. Written addition and multiplication start at the rightmost digit and carry leftwards, but the model emits the leftmost digit first. To get the first digit right it must effectively compute the whole result internally in one forward pass, which works for short sums it has seen often and fails as numbers get longer.

There is no exactness guarantee. A model produces a likely-looking answer. For common calculations (12 x 12) that is reliably correct because it was memorised. For 17% of 8,490 it may produce 1,443.3 or 1,433.3 with equal confidence. Reasoning models and step-by-step prompting help a lot because they write intermediate results down, turning one hard step into many easy ones, but they still make occasional slips, and one slip in a 40-step calculation spoils the result.

The engineering answer is simple: do not ask a language model to be a calculator. Let it decide what to calculate and with which numbers, then compute in code. This also gives you a log of the exact expression, which is easier to audit than a paragraph of prose.

Why should models use tools?

Tools give a model what its weights cannot: current facts, exact computation, private data and the ability to act. The model decides which tool to call and with what arguments; deterministic code does the part that must be correct.

A model alone has three hard limits. Its knowledge is frozen at a training cutoff and contains nothing private to your organisation. Its computation is approximate, so arithmetic, date maths, sorting and counting are unreliable at scale. And it can only produce text; it cannot read a live balance, send an email or update a ticket. Tools address all three.

With tool calling (also called function calling), you describe each tool with a name, a description and a JSON Schema for its arguments. The model returns a structured request such as get_account_balance(account_id="A-1029"). Your code validates it, runs it, and returns the result to the model as new context. The model never executes anything itself; it proposes and your harness disposes.

This shifts the model's job from knowing and computing to choosing and composing, which plays to its strengths. Language models are good at understanding a request, mapping it to the right operation and explaining the result. Databases, calculators, search engines and APIs are good at being exact. Combining them gives answers that are both fluent and correct.

Tools carry costs. Each call adds latency and a model round trip. Every tool is an attack surface: tool results can contain injected instructions, and write tools can do damage if the model is manipulated. More tools also mean more chances to pick the wrong one. The discipline is to offer few, well-described tools, validate every argument, and require confirmation for actions that are hard to undo.

What is symbolic reasoning?

Symbolic reasoning manipulates explicit symbols with exact rules, as in algebra, logic, SQL or a rules engine. Each step is valid by construction and can be checked. LLMs imitate it in text; symbolic systems actually execute it.

In symbolic systems, knowledge is written as discrete symbols (variables, predicates, rules) and inference applies formal operations to them. Solving 3x + 5 = 20 by subtracting 5 and dividing by 3 is symbolic. So is a SQL query, a Prolog program, a type checker, a SAT solver or a business rules engine that says "if order total exceeds the credit limit, hold it". The defining property is that every step follows a rule exactly, so a correct rule set produces a correct answer every time, and the derivation can be audited.

LLMs are subsymbolic: they represent everything as continuous vectors and produce answers by pattern completion. They can write out symbolic steps convincingly, and on familiar problems those steps are usually right. But nothing guarantees each step obeys the rule. A model may silently drop a negative sign or apply a rule outside its conditions, and the prose will still look fluent.

The two approaches have complementary strengths. Symbolic systems are exact, transparent and cheap to run, but brittle: they need clean, structured inputs and someone to write the rules. LLMs handle messy language, ambiguity and open-ended tasks, but are approximate. Most robust production systems use the model to translate messy input into a symbolic form (an equation, a query, a structured record) and a symbolic engine to do the exact work.

A good test for which to use: if you could write the rule down and it would not change between requests, encode it in code. Ask the model only for the parts that need interpretation.

What is test-time compute?

Test-time compute is the computation a model spends while answering a request, as opposed to during training. Spending more of it, through longer reasoning, multiple attempts or search, can raise accuracy on hard problems at the cost of latency and tokens.

Traditionally, a model's capability was fixed by training: bigger model, more data, better results. Every request got roughly the same computation per output token. Test-time compute (also called inference-time compute) treats answer-time effort as an adjustable resource. A hard question can get ten times the work of an easy one.

There are two broad ways to spend it. Sequential compute lets the model think longer in one pass: it writes a chain of intermediate reasoning, checks itself, backtracks and tries again before committing to an answer. Reasoning models are trained, typically with reinforcement learning on problems with checkable answers, to use long internal reasoning productively. Parallel compute samples several independent attempts and combines them, by majority vote or by a verifier that scores each candidate. Many systems mix both.

The reason it works is that many problems are easier to check than to solve in one step. Writing out partial results gives the model more serial computation than a single forward pass allows, and lets later tokens condition on earlier work. Research such as the 2024 paper Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters found that, for some reasoning tasks, a smaller model given extra answer-time compute can match a much larger model. The gains depend heavily on problem difficulty: easy questions gain little, and the hardest may gain nothing if the model cannot produce a correct attempt at all.

The trade-off is direct. Reasoning tokens are usually billed as output tokens and add seconds or minutes of latency. Many providers expose a control (often called reasoning effort or a thinking budget) so you can choose per request. Spending maximum effort on every request is wasteful; spending none on hard multi-step tasks leaves accuracy on the table.

For product design, this means you can route by difficulty. A classifier or simple heuristic sends routine questions to a fast path and multi-step analysis, maths or code to a high-effort path, and you measure both against an eval set.

What is inference-time scaling?

Inference-time scaling is the observation, and the practice built on it, that accuracy on many tasks rises predictably as you spend more compute per answer: more samples, longer reasoning or deeper search. It is a second scaling axis alongside model size and training data.

Training-time scaling laws describe how loss falls as parameters, data and training compute grow. Inference-time scaling describes a similar curve at answer time: plot accuracy against samples per problem or reasoning tokens per answer, and on many reasoning tasks accuracy climbs, often roughly linearly in the logarithm of compute, before flattening. Doubling effort buys a fixed increment, not a doubling of quality.

The simplest form is self-consistency, from the 2022 paper of that name: sample several reasoning paths at non-zero temperature and take the majority final answer. Errors tend to scatter while correct answers converge, so voting filters noise. A stronger form is best-of-N with a verifier: generate N candidates and keep the one a checker prefers. The checker can be exact (unit tests, a type checker, a solver) or learned (a reward model that scores answers, or a process reward model that scores each step, as in the 2023 paper Let's Verify Step by Step).

The ceiling is set by the verifier. With an exact checker, coverage matters most: if any of N attempts is correct, you find it, so accuracy approaches the fraction of problems the model can ever solve. With a learned or voting-based selector, gains flatten sooner because the selector itself makes mistakes, and with large N it can be fooled by candidates that look right. Majority voting also only works when final answers are directly comparable; for free-form prose you need a judge or a different method.

Cost scales linearly with N while benefit grows roughly logarithmically, so there is always a point of diminishing returns. Find it empirically for your task: run N = 1, 2, 4, 8, 16 on an eval set, plot accuracy against cost and latency, and pick the knee. Parallel sampling can run concurrently, so latency grows less than cost, which matters for user-facing paths.

What is self-verification?

Self-verification is checking an output against requirements or evidence before returning it. It works best when the check uses something external, such as code, tests, a schema or source documents. A model rereading its own answer without new information catches far less.

The idea rests on the asymmetry between generating and checking: confirming that a total equals the sum of its line items is easier than producing a perfect invoice in one pass. A verification step turns that asymmetry into reliability by inspecting a draft and either accepting it, repairing it or rejecting it.

There are two very different kinds. Grounded verification checks the output against an independent signal: run the generated code against unit tests, validate JSON against a schema, recompute totals in code, confirm every cited passage actually appears in the source. These checks are cheap, deterministic and catch real errors. Intrinsic self-verification asks the same model to review its own answer with no new information. Research including the 2023 paper Large Language Models Cannot Self-Correct Reasoning Yet found that on reasoning tasks this often fails to improve accuracy and can make it worse, because the model shares the same blind spots that produced the error and sometimes "corrects" right answers into wrong ones.

Intrinsic review is not useless. It helps for checklist-style requirements the model can see in the output: is every requested section present, did the reply answer all three questions, is the tone appropriate. Turning requirements into an explicit checklist and asking for a pass or fail on each item works better than "check your answer". Reasoning models also perform some self-checking inside their reasoning, which partly explains their gains on maths and code.

Design verification as a layered pipeline. Put deterministic checks first because they are cheap and certain. Use model-based checks only for properties code cannot test, and feed them evidence (the source, the requirements) rather than asking for an opinion. Decide in advance what happens on failure: retry with the error message, fall back, or escalate to a human.

When should you use another model to verify an output?

Use a second model as verifier when the property cannot be checked by code, the cost of an error justifies the extra call, and you have measured that the verifier catches real failures. Expect shared blind spots, especially within one model family.

A separate verifier helps because it sees the output fresh: it is not anchored on the reasoning that produced the draft, and it can be given a narrower job ("does this summary contain any claim not supported by the source?") than the generator had. That narrowing is where most of the value comes from. Verification questions with a clear criterion and the evidence attached are far more reliable than open requests to "review this".

Independence is partial. Models trained on similar data with similar methods make correlated errors: if the generator misreads an ambiguous clause, a verifier from the same family often misreads it the same way. Model judges also have known biases, including favouring longer answers and outputs from their own family. Using a verifier from a different provider or family, giving it the source documents, and asking for specific failures rather than a score all reduce correlation, but never remove it.

Cost and latency are real. A verifier roughly doubles model spend on that path and adds a sequential call. That is worth it when errors are expensive (a contract clause extracted wrongly, a medical summary, a payout decision) and rarely worth it for low-stakes, high-volume chat. A common compromise is to verify a sample for monitoring and verify everything only on high-risk paths.

Before relying on a verifier, measure it like any classifier. Build a labelled set that includes known bad outputs and check its recall (how many real errors it flags) and precision (how many flags are real). A verifier that flags 30% of correct outputs will be ignored or will drown reviewers; one that misses half the errors gives false comfort.

What is neuro-symbolic AI?

Neuro-symbolic AI combines learned neural models with explicit symbolic components such as rules, logic, solvers or code. The neural part handles perception and messy language; the symbolic part supplies exactness, constraints and auditability.

The motivation is that each approach fails where the other succeeds. Neural networks generalise from examples and cope with ambiguity, but give no guarantees. Symbolic systems guarantee correctness relative to their rules and explain every step, but need structured input and hand-written knowledge. Combining them aims for systems that read the world like a model and decide like a program.

In current LLM systems the most common pattern is translate, then execute. The model converts natural language into a formal artefact (a SQL query, a Python expression, a JSON record, a logic formula, a constraint problem) and a symbolic engine runs it. The 2022 work on Program-Aided Language Models showed that having a model write a short program and letting an interpreter compute the result beats having the model compute in text on arithmetic word problems. Tool calling, structured output with schema validation and text-to-SQL are all everyday neuro-symbolic designs, even if nobody calls them that.

A second pattern is neural proposal, symbolic checking: the model suggests candidates and a checker accepts or rejects them. A widely reported example is AlphaGeometry (2024), in which a language model proposed auxiliary constructions and a symbolic deduction engine did the rigorous proof search. In software, a model proposes code and a compiler, type checker and test suite act as the symbolic judge. A third pattern is symbolic guardrails around neural decisions: a policy engine encodes non-negotiable rules (spending limits, eligibility criteria, regulatory constraints) and overrides or blocks model output that violates them.

The cost is integration effort. The formal layer needs a schema or rule set someone maintains, translation errors become a new failure mode (a correct solver running the wrong query still returns a wrong answer), and you need tests at the boundary. The payoff is that critical properties become testable with ordinary unit tests rather than statistical evals.

What is a reasoning trace, and can you trust it?

A reasoning trace is the intermediate text a model writes before its answer, such as a chain of thought or a reasoning model's thinking. It often improves the answer, but it is not a faithful record of how the answer was produced, so verify conclusions rather than trusting the narrative.

Traces come from two sources. Chain-of-thought prompting asks a model to write out steps before answering. Reasoning models are trained to produce long internal reasoning, often hidden from the user or shown only as a summary, depending on the provider. In both cases the trace is ordinary generated text that the model conditions on, which is why it helps: it gives the model room to compute in steps.

Helping is not the same as explaining. Research such as the 2023 paper Language Models Don't Always Say What They Think showed that when a prompt contained a biasing cue (for example, the correct option in the few-shot examples was always A), models followed the cue and then wrote plausible reasoning that never mentioned it. Later studies on reasoning models found they frequently fail to mention hints they demonstrably used. A trace can therefore be unfaithful: a coherent story that is not the actual cause of the answer. Traces can also contain correct reasoning followed by a wrong final answer, or flawed steps that happen to land on the right one.

That shapes how to use them. Traces are valuable for debugging: reading a sample of them often reveals misread instructions, missing context or a wrong assumption you can fix in the prompt. They are weak as evidence for an auditor, a customer or a reviewer. For user-facing explanations, prefer a short, separately generated justification tied to checkable facts (the rule applied, the figures used, the cited source), and verify those facts in code.

There are practical constraints too. Reasoning tokens usually count as billed output and add latency. Some providers return raw reasoning, some a summary, some nothing, and their terms may restrict how you use it. Do not build pipelines that parse a hidden trace for decisions, and do not train users to read long traces as proof.

What is a knowledge cutoff, and how do you work around it?

A knowledge cutoff is the date after which a model's training data stops, so anything newer is missing from its weights. Work around it by supplying current facts at request time through retrieval, search or tools, and by telling the model today's date.

Training data is collected up to a point, then the model is trained and released, often months later, and used for a long time after that. The knowledge cutoff is the end of that data. Events, prices, product versions, laws, APIs and people's roles that changed afterwards are absent or out of date in the weights.

The cutoff is fuzzier than a single date suggests. The last months before it are usually thinly covered, because the web had not yet written much about recent events when the data was gathered. A model may therefore know older facts well and recent-but-pre-cutoff facts poorly. Models are also often unsure of their own cutoff and of today's date, so without help they may describe an old version of a library as "the latest" or calculate ages and deadlines from the wrong year.

The workarounds are standard. Put the current date in the system prompt on every request. Use retrieval over your own documents for organisational knowledge. Use search or API tools for public facts that change, and ask the model to cite the returned source. For code assistants, give the model the installed library version and relevant documentation rather than relying on what it remembers about an API.

Fine-tuning is a poor tool for keeping facts current. It is slow and expensive to repeat, inserts facts unreliably, and still leaves you with a new cutoff. Fresh context at request time is cheaper and can be updated in minutes.

What happens when the context contradicts what the model learned?

The model has to choose between parametric knowledge in its weights and contextual knowledge in the prompt, and its choice is not reliable. It may override your document with a memorised fact, or accept a wrong document uncritically. Instruct, ground and test explicitly.

This is called a knowledge conflict. It appears constantly in RAG and tool use: your policy says the refund window is 15 days while general training data says 30 is common; your retrieved pricing page is newer than anything the model saw; a tool returns a value that contradicts a popular belief. Ideally the model follows the supplied evidence. In practice behaviour varies.

Research shows both failure directions. Studies such as the 2021 work on Entity-Based Knowledge Conflicts found models sometimes ignore a substituted fact in the passage and answer from memory, especially for very well-known facts. Later work, including the 2023 paper Adaptive Chameleon or Stubborn Sloth, found models are often highly receptive to coherent external evidence even when it is wrong, and lean towards whichever evidence agrees with their memory when sources conflict. The strength of the prior, the plausibility of the context and the instruction wording all shift the outcome.

Both directions are product risks. Over-trusting memory means a RAG system confidently answers from general knowledge when your document says something specific, which is exactly the case RAG was built for. Over-trusting context means a poisoned document, a stale page or a malicious tool result can make the model assert falsehoods, which is the mechanism behind RAG and tool poisoning attacks.

Mitigations work at several layers. Tell the model explicitly that supplied documents override general knowledge for this domain, and that it must say so when the documents do not cover the question. Require citations to the passage supporting each claim, and check them automatically. Rank and date sources so the model sees which is authoritative and current. Control what can enter the context in the first place. Finally, build eval cases specifically for conflicts: documents that state counter-intuitive but true policies, and documents that contain wrong claims the model should not repeat.

How can you tell whether a model is reasoning or recalling?

Change the surface of a problem while keeping its logic, and see whether accuracy holds. Recall breaks under new numbers, names, irrelevant details or reordered steps; reasoning that generalises survives them. Test on your own perturbed cases, not public benchmarks.

Public benchmark scores mix two abilities: solving a problem and remembering one like it. Many benchmark questions, or close paraphrases, appear on the web and can end up in training data, a problem called contamination. A high score may therefore reflect familiarity rather than a skill that transfers to your data.

The main technique is perturbation testing. Take problems the model solves and create variants that preserve the logic but change the surface: different numbers, different names and units, reordered sentences, an extra irrelevant clause, or an equivalent rephrasing. The 2024 study GSM-Symbolic did this for grade-school maths and found that accuracy dropped when only numbers changed, and dropped sharply when a plausible but irrelevant sentence was added, suggesting part of the original performance relied on pattern matching. Stronger reasoning models narrow these gaps but do not always close them.

Other signals help. Counterfactual tasks keep the reasoning identical but change a convention (arithmetic in base 9, a chess variant with swapped pieces); a model that truly applies the procedure should cope, one relying on memorised outcomes will not. Consistency checks ask the same question in logically equivalent forms and check the answers agree. Held-out private data from your own domain is the most reliable test of all, because it cannot have been in training.

The practical point is not philosophical. Whether you call it reasoning or not, you need to know if performance survives the variation your users will produce. A model that handles your test cases but fails when an invoice has one extra line item is not ready, however it scores on public leaderboards.