Connect AI to company systems with permissions, ownership, and measurable outcomes.
Enterprise AI is the work of connecting models to a company's real systems and people: document stores, ERPs, CRMs, ticketing tools and approval chains. The model is the easy part. The hard parts are making sure an assistant only sees what the asking employee could see, that an agent can only take actions it was granted, that every consequential decision can be explained months later, and that someone owns each system when it goes wrong.
The mental model is a platform with a control plane. Applications (knowledge assistants, copilots, agents) sit on top. Underneath, shared services handle identity, model access, data connectors, a registry of agents and tools, evaluation, tracing and audit. Governance decides which applications may use which pieces at what risk level, and the platform enforces those decisions in code rather than in policy documents.
The questions move from the shape of a platform and how AI should reach data, through integrations with ERP and CRM and the two most common application types, to governance. The advanced questions cover the control plane itself: agent and tool registries, authentication, audit trails, managing a fleet, environment separation, approval workflows and measuring whether any of it pays off.
A shared layer that every internal AI app uses: governed model access, identity and permissions, data connectors, a tool catalogue, evaluation, tracing and audit, cost controls and a clear owner for each app. Without it, every team rebuilds these badly.
An enterprise AI platform is the common infrastructure that sits between a company's AI applications and everything they touch: model providers, internal data, business systems and users. The reason to build one is repetition. The first AI app in a company wires up its own API key, its own logging and its own permission checks. By the fifth app, there are five sets of keys, five different ideas of what gets logged, and nobody can answer "which apps can read payroll data?" A platform answers that once.
The core pieces are stable across companies. A model gateway gives one authenticated route to approved models, with routing, rate limits, budgets and fallback. Identity ties every request to a real user or service so permissions can be enforced downstream. Data connectors reach document stores, wikis and databases while respecting source permissions. A tool registry lists the actions agents may take, with schemas and access rules. Evaluation runs before releases, and observability captures traces, cost and quality signals in production. Governance records who owns each app, what risk tier it is in and what approvals it has.
The platform should be thin where teams differ and firm where risk lives. Product teams should choose their own prompts, retrieval strategy and user experience. They should not choose whether to log, whether to check permissions, or which unapproved provider to send customer data to. A useful test: a new team should be able to ship a compliant internal assistant in days by assembling platform pieces, and it should be hard to ship one that bypasses them.
Build the platform from real demand, not a reference diagram. Most companies start with the gateway and logging (because spend and data exposure are the first problems leadership notices), then add connectors and the tool registry when the second or third app needs the same systems.
Through governed connectors that act with the requesting user's permissions, enforce them at retrieval time, and record what was read. The model should only ever see data the person asking could already open themselves.
The safe default is simple to state: the AI never has more access than the user it is serving. If an employee cannot open a folder in the document system, the assistant must not retrieve from it on their behalf. Breaking this rule turns the assistant into a permission bypass. A junior employee asks "what are the salary bands for directors?" and gets an answer drawn from an HR spreadsheet they were never meant to see.
There are two common ways to enforce this. Query-time delegation calls the source system with the user's own credentials (often an OAuth token obtained on their behalf), so the source applies its own permissions. It is always current but can be slow and depends on every source having a good search API. Indexed retrieval with access control lists copies content into a search or vector index along with each document's allowed users and groups, then filters by the user's identity at query time. It is fast and supports semantic search, but permissions are only as fresh as the last sync, so revoked access can linger.
Either way, the filter must be applied before content reaches the model, not after. Asking the model to "ignore documents the user cannot see" does not work: once text is in the context, it can leak into the answer. Filtering belongs in the retrieval layer, enforced by code that the model cannot influence.
Beyond permissions, governed access means a few more things. Connectors should cover only approved sources, with sensitive classes (health records, payroll, legal holds) excluded by default. Every retrieval should be logged with the user, the query and the document ids returned, so a later question like "who saw this contract?" can be answered. And content pulled from internal sources should still be treated as untrusted input, because a shared document can carry an indirect prompt injection just as a web page can.
Through narrow, well-defined tools on top of the ERP's own APIs, with reads used freely and writes limited to drafts or low-value actions behind validation and approval. The ERP stays the system of record and enforces its own business rules.
An ERP (enterprise resource planning system) holds a company's financial and operational records: purchase orders, invoices, inventory, the general ledger. Mistakes there have direct money and audit consequences, so the integration pattern matters more than the model. The right shape is: the model proposes, deterministic code validates, the ERP records, and a human approves anything consequential.
Expose the ERP through task-shaped tools, not raw access. A tool called draft_purchase_order(vendor_id, lines, cost_center) is safer and easier for a model to use than a generic call_erp_api(endpoint, payload) or a database connection. Narrow tools let you validate arguments (does this vendor exist, is this cost centre active, is the amount within the requester's limit) before anything is written. They also make audit readable: the log says "drafted PO for 12 laptops" rather than "POST /api/v2/obj/4471".
Use the ERP's own controls rather than duplicating them. Most ERPs already support draft or parked documents, approval workflows and segregation of duties. An AI integration should create a draft that flows into the existing approval chain, not post a final document directly. That keeps finance's controls intact and means auditors see the same process they already understand.
Read access is where most early value lies: answering "why is this invoice on hold?", reconciling mismatches between a purchase order, goods receipt and invoice (the classic three-way match), or summarising a vendor's open items. These need scoped read tools and careful handling of what data leaves the ERP, since financial records often have stricter residency and retention rules than general documents.
Read customer records the user is allowed to see to summarise, prepare and suggest, and write back only through validated, field-level tools that log the AI as the source. Most value comes from reading and drafting, not from autonomous updates.
A CRM (customer relationship management system) holds accounts, contacts, opportunities, activity history and support cases. It is rich in context that sales and support staff struggle to absorb quickly, which makes it a natural fit for summarisation and preparation: "brief me on this account before my call", "what has this customer complained about this year?", "which open deals have gone quiet for 30 days?"
Reading must respect the CRM's own sharing model. CRMs typically restrict records by owner, team, territory or role, and the assistant should query with the user's identity so those rules apply. A sales rep in one region should not get a summary that includes another region's pipeline because the integration used an administrator account.
Writing is where CRM integrations go wrong. Models are good at drafting notes and suggesting field values, and bad at knowing which updates are appropriate. An agent that freely updates stages, amounts or close dates will corrupt forecasts that leadership depends on. Safer patterns are: draft, then confirm (the rep sees the proposed activity log or field change and clicks save); field allowlists (the AI may set next_step and log activities but never touch amount or stage); and provenance marking (every AI-written record carries a source flag so reports can separate human and AI entries).
Data quality cuts both ways. CRM data is often stale or duplicated, so summaries should cite the records they used and show dates, letting the user notice that "the main contact" left two years ago. And customer data in CRMs is usually personal data, so what is sent to the model provider, and how long it is retained, falls under the same privacy rules as any other processing.
A question-answering system over a company's own documents, usually built with retrieval-augmented generation, that answers with citations and respects who may see what. Its quality depends mostly on the content and retrieval, not the model.
A knowledge assistant lets employees ask questions in plain language and get answers drawn from internal sources: policies, wikis, product documentation, past tickets, contracts. It is usually the first enterprise AI project because the need is universal (people waste time searching) and the risk is moderate (it reads, it does not act).
Under the hood it is almost always retrieval-augmented generation (RAG): the question is used to search an index of chunked documents, the most relevant passages are placed in the prompt, and the model writes an answer grounded in them. Citations link each claim to its source so the user can verify, which is essential when the answer is a leave policy or a security rule.
Most failures come from content, not models. Companies have three versions of the travel policy, an outdated one ranks highest, and the assistant confidently quotes it. Fixes are mostly editorial and operational: mark one source as authoritative per topic, record effective dates and owners, retire obsolete documents, and keep the index synced with changes. A knowledge assistant often surfaces content problems that existed long before AI.
Good assistants also know when not to answer. If retrieval finds nothing relevant, the right response is "I could not find this in the HR policies; the people team can help" rather than a plausible answer from the model's general knowledge. That behaviour has to be designed and tested, because a model's default is to be helpful. Permissions, covered in how AI should access internal data, apply here too: the assistant must retrieve only what the asking employee can see.
An assistant embedded inside the tools employees already use, such as a support console, IDE or CRM, that sees the task in front of the user and suggests drafts, summaries or next steps. The user stays in control and decides what to accept.
A copilot differs from a standalone chatbot in where it lives and what it knows. Instead of a separate window where the employee explains their situation, it sits inside the application they are working in and already has the context: the ticket being handled, the document being edited, the account being viewed. That removes the most expensive step in using AI at work, which is copying context into a prompt.
The interaction model is suggest, then accept. A support copilot drafts a reply using the ticket history and knowledge base; the agent edits and sends. A finance copilot flags which invoice lines don't match the purchase order; the analyst decides. The human remains the actor of record, which keeps accountability clear and makes errors cheaper, because a bad suggestion is ignored rather than executed.
Copilots are judged on workflow metrics, not chat quality. The useful questions are: what fraction of suggestions are accepted, how much are they edited before use, how much time per task changed, and did quality (customer satisfaction, reopen rate, error rate) hold steady. A copilot whose drafts are accepted 70% of the time but cause a rise in reopened tickets is making things worse.
There is a ceiling on value. If users accept almost every suggestion unchanged for a well-defined task, that task may be a candidate for automation with sampling review instead. If they reject most suggestions, the context or retrieval is probably wrong. Copilots also inherit the host application's permissions model, so they should act with the user's identity, not a broad service account.
The set of owners, rules and checkpoints that decide which AI systems a company runs, with what data and permissions, at what risk level, and who answers when something goes wrong. Good governance is proportionate to risk and built into the delivery process.
AI governance answers organisational questions that technology alone cannot. Who approved this assistant reading customer emails? Which model providers may receive personal data? Who is accountable if the invoice agent pays the wrong vendor? What must change before an internal pilot reaches customers? Without explicit answers, teams either ship without checks or wait months for an approval nobody knows how to give.
A working model has a few parts. An inventory of every AI system with a named business owner and technical owner. A risk classification (for example low, medium, high) based on what the system can do and to whom: an internal writing aid differs from an agent that changes customer accounts. Proportionate controls per tier, such as evals, human review, privacy assessment, security review and sign-off. And change management: a model swap, a new tool or a new data source triggers re-review for higher tiers.
External frameworks help structure this. The NIST AI Risk Management Framework, the ISO/IEC 42001 management-system standard and the EU AI Act all push toward risk-based classification, documentation and human oversight. Which obligations apply depends on jurisdiction and use case, so legal and compliance teams should map them; engineers should make sure the evidence those frameworks ask for (inventory, evals, logs, approvals) is produced automatically.
The common failure is governance as a committee that reviews slide decks. It slows everything equally and catches little, because the real risks live in permissions and data flows. Effective governance is mostly encoded: the platform will not grant a tool to an unregistered agent, a high-risk app cannot deploy without a passing eval report, and the review board spends its time on the small number of genuinely risky cases.
A system of record for every agent a company runs: its owner, purpose, risk tier, version, identity, allowed tools and data, eval status and lifecycle state. The platform reads it to grant access, so an unregistered agent cannot act.
Once a company has more than a handful of agents, basic questions get hard. Who owns the invoice agent? Which version is in production? Which agents can touch HR data? Is the agent that sent this email still supposed to be running? An agent registry answers these by keeping one record per agent, versioned like code, that the rest of the platform consults.
A useful record contains: a stable agent id and a workload identity (the credential it uses to call the gateway and tools); owners (a business owner accountable for outcomes and a technical owner on call for failures); purpose in a sentence, which reviewers use to judge whether requested permissions fit; risk tier; the model, prompt and configuration version currently deployed; the tools and data scopes it may use; budgets for tokens, cost and actions; eval status with a link to the latest report; and a lifecycle state such as draft, pilot, production, suspended or retired.
The registry only matters if it is enforced. The gateway should reject model calls from an agent id that is not registered or not in an active state. The tool layer should check the registry before honouring a call, so granting a tool means editing the registry record (through review), not editing code. Suspending an agent should be one state change that takes effect everywhere within seconds. This turns the registry from documentation into a control plane.
Treat records as code: store them in version control or a service with an audit history, require review for permission changes, and generate the human-readable catalogue from the same data. Retirement deserves attention. Agents built for a pilot often keep their credentials long after anyone uses them, and an orphaned agent with write access is a standing risk. A quarterly job that flags agents with no traffic or no valid owner pays for itself.
A managed catalogue of every action agents can take, with each tool's schema, owner, side-effect class, risk level, authorisation rules, rate limits and which agents may call it. It makes tool access reviewable, enforceable and revocable in one place.
A tool is a function an agent can call: search_tickets, issue_refund, create_jira_issue. In a single application, tools are just code. Across an enterprise, the same tool may be used by ten agents built by five teams, and the risky ones need consistent controls. A tool registry is the catalogue that holds those controls, much as an API gateway holds them for ordinary services.
Each entry records the schema (name, description, typed arguments, so every agent sees the same contract), an owner team responsible for the underlying system, a side-effect class (read, reversible write, irreversible write, external communication), a risk level, authorisation rules (which agents, acting for which users, may call it, and whether approval is required above some threshold), limits (rate, amount caps, daily totals) and a version. Protocols such as the Model Context Protocol (MCP) standardise how tools are described and discovered; the registry adds the governance that the protocol leaves to you.
Enforcement happens at call time, outside the model. When an agent emits a tool call, the runtime checks the registry: is this agent allowed this tool, does the acting user have the underlying permission, do the arguments pass validation, is the call under its limits, does it need approval? Only then is it executed. The model's prompt can list available tools, but the prompt is not the control. An agent that hallucinates a tool name or is manipulated into calling issue_refund should be stopped by the registry check, not by its instructions.
Registries also help with tool poisoning and supply-chain risk. Third-party MCP servers can change tool descriptions or behaviour after you approved them. Pinning versions, reviewing description changes, and allowing only registered servers prevents an agent from picking up a new, unreviewed tool silently. Over time the registry becomes the place where you see which actions agents actually take, which is the input for deciding what to automate further and what to lock down.
Record, for every consequential decision, the inputs and retrieved evidence, the exact model, prompt and tool versions, each tool call and its result, any human approval, and the final outcome, in tamper-evident storage, so the decision can be reconstructed and explained later.
An AI audit trail exists to answer a question asked weeks or months later: why did the system do this? "Why was this ticket escalated?", "Why was this supplier's invoice held?", "Who approved that refund?" Debug logs rarely answer these because they are sampled, short-lived and missing the versions that were live at the time. Audit records are a deliberate, structured subset designed to survive.
A useful record for each decision includes: a decision id linking all steps; the actor chain (which user, which agent, under which identity); the inputs, including references to the documents or records retrieved, not only the user's message; versions of the model route, prompt, configuration and tools, because "the model" changes over time; each tool call with arguments and results; the output or action taken; any approval with who approved, when and what they saw; and the outcome if known later (refund reversed, escalation confirmed). Model reasoning text can be stored as supporting context, but it is not a faithful explanation of why the model produced its output, so the audit case should rest on evidence and actions, not on the model's self-description.
Storage needs care. Records should be append-only and tamper-evident (write-once storage or hash chaining), with retention set by policy and legal requirements. Audit logs often contain personal data, so access must be restricted and some fields redacted or stored by reference. The tension is real: you want enough to reconstruct the decision, but not a second uncontrolled copy of every customer's records. Storing document ids and content hashes instead of full text is a common compromise when the source system keeps history.
Auditability must be tested like any other feature. Pick a past decision and try to reconstruct it from the record alone: can you see what the agent saw, which versions were live, and who approved? If the answer is "we would need to rerun it", the trail is incomplete, because model outputs are not reproducible and sources change.
Treat agents as a fleet: one registry, standard identity, shared tool and model gateways, common evals and telemetry, per-agent budgets and a central kill switch. Standardise the controls while leaving teams free to design each agent's logic.
The jump from three agents to thirty changes the problem. With three, each team knows its agent intimately. With thirty, nobody can name them all, costs appear on one shared bill, two agents may update the same CRM records with conflicting logic, and an incident starts with "which agent did this?" Fleet management is the set of practices that keeps this tractable, similar to how companies manage many microservices.
Standardise five things. Identity: every agent has its own workload identity, never a shared key, so every call is attributable and revocable. Access: tools and data are granted through the registry, never hard-coded. Telemetry: every agent emits traces in the same format (OpenTelemetry's generative AI conventions are a common choice) with the agent id, version and run id on every span, so one dashboard covers all of them. Evaluation: every agent has an eval suite that runs on every change and a minimum bar per risk tier. Budgets: per-agent limits on tokens, cost, steps per run and actions per day, enforced at the gateway rather than trusted to the agent's code.
Give operators central levers. A kill switch suspends an agent everywhere in seconds by flipping its registry state, which the gateway and tool layer honour immediately. Per-tool disable turns off one risky action across all agents while the rest keep working. Rollback returns an agent to its previous prompt and configuration version. These must be rehearsed; a kill switch that has never been pulled is a guess.
Watch for interactions between agents. Two agents that both react to new tickets can loop, each updating the ticket and triggering the other. Agents that call other agents multiply cost and blur accountability. Rules that help: each record type has one agent allowed to write it, agent-to-agent calls go through the same authorisation as user calls, and loops are capped by a hop count carried in the request. Finally, report fleet health by agent: usage, cost per task, eval score, incident count and owner status. Agents with no users and no owner should be retired.
Each environment needs its own credentials, data, tool backends and permissions, so that an agent in dev or staging physically cannot touch production systems. Configuration, prompts and models should be promoted as versioned artifacts with eval gates between environments.
For ordinary software, environment separation is routine. AI systems add two complications. First, agents take actions through tools, and a test agent wired to a real payment API can do real damage regardless of what its prompt says. Second, the behaviour depends on artifacts that are easy to change outside normal deployment: prompts, model routes, retrieval indexes, tool descriptions. If those are edited live in production, nothing you tested in staging is what is running.
Separate the blast radius. Each environment gets its own credentials and its own agent identities; staging identities must be rejected by production tools. Tool backends differ: dev uses mocks with recorded responses, staging uses sandboxes or simulators (payment providers' test modes, a staging ERP, an email sink that never delivers), and only production reaches real systems. Data differs too: dev and staging use synthetic or properly anonymised data, because copying production customer records into a staging vector index quietly multiplies your privacy exposure.
Treat prompts, model routes and tool definitions as versioned configuration promoted through environments, just like code. A change goes to dev, passes the eval suite, goes to staging, passes integration tests against simulators, and is promoted to production as the exact same artifact. Many teams then add a canary or shadow stage in production, where the new version handles a small share of traffic or runs alongside the current one without acting, because staging traffic never fully matches real users.
Some differences are deliberate. Production may use stricter budgets, larger models or different providers under contract; staging should mirror production's model routes closely enough that eval results transfer. Production logs follow retention and redaction policy; dev logs can be verbose because the data is synthetic. Finally, keep a production-equivalent eval run: a scheduled job that runs the golden set against the production configuration catches silent provider-side model updates that no deployment triggered.
Pause the agent at a defined checkpoint, persist a structured proposal, route it to the right approver by policy, show them the evidence and exact action, then execute only the approved version with an expiry, and record every decision.
An approval workflow turns an agent's intended action into a request that a person must accept before anything happens. The design questions are: which actions need approval, who approves, what they see, and what happens while waiting. Get any of them wrong and approvals become either a bottleneck everyone resents or a rubber stamp that adds no safety.
Which actions should come from policy, not the model. Rules based on side-effect class and thresholds work well: refunds over a limit, any external email to more than one recipient, any change to bank details, any deletion. The policy engine evaluates the proposed call; the model never decides whether its own action needs approval. Who approves follows the business's existing authority: the requester's manager, a finance controller above a higher limit, two people for the riskiest actions. Never route approval to the same person who requested the action when segregation of duties applies.
What the approver sees decides whether review is real. Show the exact action with concrete values ("refund 180.00 to card ending 4421 for order 7731"), the evidence the agent relied on with links, what is unusual about it, and the consequence of approving. Let the approver edit parameters within bounds or reject with a reason. A wall of model-generated justification invites skimming; a compact, structured card invites checking.
Mechanics matter for correctness. The agent run must be durable: persist state at the checkpoint and resume when the decision arrives, possibly hours later, rather than holding a process open. Bind the approval to the exact proposal with a hash, so the agent cannot execute a different action than the one approved. Give approvals an expiry, because the world changes (the order may already be refunded). Execute with an idempotency key so a retry does not double-pay. Record approver, time, what was shown and any edits in the audit trail.
Measure the workflow. If approvers accept 99.5% of requests in under ten seconds, either the threshold is too low or review has become automatic; raise the threshold or sample. If they reject often, the agent or its inputs need work. Approval is a control and a data source.
Tie each deployment to a business metric it should move, measure against a baseline or control group, count the full cost (models, platform, people reviewing output, error correction), and track quality alongside speed. Usage and satisfaction surveys alone are not evidence of value.
Enterprise AI programmes are often judged by adoption numbers: weekly active users, messages sent, "hours saved" from a survey. These show interest, not value. The useful question is whether a specific business outcome improved by more than the full cost of the system, and whether quality held while it did.
Start with a baseline for the task the AI touches, measured before launch: handle time per ticket, days to close the books, invoices processed per analyst, first-contact resolution rate. Then compare against it after launch, ideally with a control group: some teams or a random share of cases continue without the tool for a few weeks. Without a control, seasonal changes, staffing shifts and process changes get credited to the AI. A rollout in waves gives you a natural control at no extra cost.
Pair every speed metric with a quality metric. Faster ticket handling with a higher reopen rate, or faster contract review with more missed clauses, is a cost moved downstream. Useful quality signals include error rates found in sampled review, escalations, customer satisfaction and rework. For copilots, track how much suggestions are edited, not just whether they are accepted.
Count total cost, not only model spend. That includes platform and engineering time, the people who review or approve AI output, time spent correcting errors, and the eval and governance effort. A process that saves 3 minutes per case but adds 2 minutes of mandatory review has a much smaller net effect than the headline suggests. Express results as cost per completed task or per outcome, which you can compare with the pre-AI process directly.
Finally, decide in advance what result would cause you to stop or redesign. Many pilots linger because nobody set a bar. A pre-agreed threshold (for example, at least 15% reduction in handle time with no rise in reopens after eight weeks) turns the pilot into a decision.
Buy for generic tasks where a vendor's product already fits and your differentiation is low, such as general writing help or meeting notes; build where the task depends on your own processes, systems and data, or where control over actions and audit is essential.
Almost every enterprise faces this choice repeatedly: an AI feature inside an existing SaaS tool, a horizontal assistant product, or something the company builds on its own platform. The decision is less about model quality, which vendors and in-house teams access similarly, and more about fit, control and integration.
Buying fits when the task is common across companies (drafting emails, summarising meetings, general Q&A over standard office tools), when the vendor already integrates with the systems involved, and when the risk is low. You get speed and someone else's maintenance. The costs are less control over prompts and behaviour, dependence on the vendor's data handling and roadmap, and per-seat pricing that can grow faster than usage justifies. Check contracts for data retention, training use and where data is processed, since these vary by vendor.
Building fits when value comes from your specific processes and systems: an agent that understands your ERP's custom fields, a support copilot tuned to your product catalogue and escalation rules, an approval flow that mirrors your authority matrix. It also fits where you need full audit trails, fine-grained permissions or specific deployment constraints. The costs are engineering time, ongoing evaluation and maintenance, and the platform work underneath.
Most companies end up hybrid: bought tools for broad productivity, built systems for core workflows, sharing identity, data governance and, where possible, one gateway and audit approach. A useful rule is to buy the commodity and build the differentiating. Revisit the decision periodically, because vendors add features quickly and a system you built last year may now be cheaper to replace, or a bought tool may have hit the limit of what it can do for your process.
Give every agent its own workload identity and, when it acts for a person, use delegated, short-lived, narrowly scoped credentials that carry both identities, so systems enforce the user's permissions and logs show who asked and which agent acted.
Agents need credentials to call tools: the ERP, the CRM, the document store, the ticketing system. The tempting shortcut is one powerful service account per integration, shared by every agent. That makes every agent as powerful as the account, makes logs say "service-account-crm did it" with no trace of who asked, and means revoking one agent's access breaks all of them.
There are two identities to carry. The agent identity (a workload identity, such as a certificate or a token from the platform's identity service) says which registered agent is calling. The user identity says on whose behalf. For assistants that act for a person, the standard pattern is delegation: the platform exchanges the user's sign-in token for a downstream token scoped to the target system, with the agent recorded as the actor. OAuth 2.0 token exchange (RFC 8693) and the "on-behalf-of" flows offered by major identity providers implement this. The downstream system then applies the user's own permissions, so the agent can never exceed them.
For background agents with no user in the loop, such as a nightly reconciliation job, use the agent's own identity with permissions granted for that purpose only, and still avoid sharing it across agents. In both cases, keep tokens short-lived (minutes, not months) and narrowly scoped to the operations the tool needs. A tool that only reads invoices should hold a token that cannot write them.
Keep credentials out of the model's reach entirely. The tool runtime holds and attaches tokens; the model sees only tool names and results. A credential in the prompt or context can be leaked by an injection or written into logs. And when the user's access is revoked, delegated tokens should stop working at their next refresh, which is one reason short lifetimes matter.