Know what each tool is for before adding it to your stack.
The GenAI tool ecosystem grew faster than most teams could evaluate it. There are libraries for composing model calls, runtimes for stateful agents, frameworks for retrieval, multi-agent toolkits, visual automation platforms, local and high-throughput inference engines, gateways that sit in front of many providers, prompt optimizers, typed agent libraries, and platforms for tracing and evaluation. Each solves a real problem, and each adds a dependency, a mental model and an upgrade path you will maintain.
The useful mental model is a stack of layers, each with a job. Model access (provider SDKs, gateways such as LiteLLM, local servers such as Ollama, serving engines such as vLLM) gets tokens in and out. Building blocks (LangChain, LlamaIndex, Hugging Face libraries) provide integrations and retrieval. Orchestration (LangGraph, CrewAI, PydanticAI, DSPy, n8n, provider agent SDKs) decides what runs next. Observability and evaluation (Langfuse, Arize Phoenix) show what happened and whether it was good. Knowing which layer a tool occupies tells you what it can replace and what it cannot.
The Basic questions describe what each widely used tool does and when it fits. The Advanced questions cover serving, gateways, prompt optimization, typed agents and tracing in more depth, then step back to the decisions that outlast any single tool: when a framework earns its place, how to keep the option to switch, and what provider-built agent SDKs offer. Tools change; the questions you ask of them do not.
LangChain is an open-source Python and JavaScript library that gives common interfaces for chat models, prompts, retrievers, tools and output parsing, plus a large catalogue of integrations, so you can assemble LLM applications from swappable parts.
LangChain started in 2022 as a way to chain model calls with prompts, memory and tools. Its lasting value is standard interfaces. A chat model from any supported provider is called the same way, a retriever always takes a query and returns documents, and a tool is a function plus a schema the model can call. Around those interfaces sits a large set of integration packages: document loaders for file formats and SaaS systems, text splitters, embedding models, vector stores and tool wrappers.
The core abstraction is the runnable, a component with the same invoke, batch and stream methods whether it wraps a prompt template, a model or a parser. Runnables compose into pipelines, so prompt, model and parser can be joined and then streamed or batched as one unit. Recent major versions moved the library's high-level agent onto LangGraph (the graph runtime from the same team) and moved many older chain classes into a separate legacy package. Tutorials written for earlier versions often use APIs that are now deprecated, so check which version an example targets.
LangChain is useful when you need many integrations quickly: trying three vector stores, two embedding models and several providers in a prototype is mostly configuration. It is less useful when your application is a single provider call with a well-defined prompt, where the abstraction adds layers between you and the request you are debugging. The same company sells LangSmith, a tracing and evaluation platform; the library works without it, and it can send traces to other tools through callbacks or OpenTelemetry exporters.
The trade-off is control. Every wrapper decides something for you: how messages are formatted, how retries work, how tool results are inserted. When behaviour surprises you, you need to read the wrapper. Teams that succeed with LangChain tend to use a few components they understand (model interface, loaders, splitters) and write their own orchestration around them.
LangGraph is an open-source library for building stateful LLM workflows and agents as graphs: nodes do work, edges decide what runs next, and a checkpointer saves state after each step so runs can pause, resume, loop and wait for human approval.
Many agent and workflow problems are really state machines: draft, review, revise if the review fails, stop after three attempts, ask a person before sending. LangGraph makes that explicit. You define a typed state object, add nodes (functions that read the state and return updates), and connect them with edges. Conditional edges are functions that look at the state and return the name of the next node, which is how loops and branches are expressed. Its execution model is inspired by Google's Pregel system: work proceeds in steps, and each step's updates are merged into the shared state.
The feature that matters most in production is persistence. With a checkpointer configured, the state is saved after every step, keyed by a thread id. That gives you resumability after a crash, multi-turn memory for a conversation, interrupts that pause a run until a person approves and the run continues from the same point, and the ability to inspect or replay earlier states when debugging. Building that reliably yourself means designing a state store, serialisation and a resume protocol.
LangGraph comes from the LangChain team but does not require LangChain; nodes can call any SDK. It also streams intermediate updates and tokens, which helps user interfaces show progress. A commercial platform from the same company offers hosted deployment and a visual debugger, but the open-source library runs on your own infrastructure.
The cost is a new mental model and a dependency at the heart of your control flow. For a linear pipeline of three calls, plain functions are clearer. LangGraph pays off when you have loops, branches, long-running work, or human approval steps that must survive restarts.
LlamaIndex is an open-source framework, in Python and TypeScript, focused on connecting LLMs to your data: loading and parsing documents, building indexes, retrieving relevant chunks and synthesising answers, with event-driven workflows and agents built on top.
LlamaIndex began as a library for indexing documents for LLMs, and retrieval-augmented generation (RAG, answering from retrieved documents) is still its centre of gravity. The pipeline it models has clear stages. Readers (data connectors) pull content from files, databases and SaaS tools. Node parsers split documents into nodes, its name for chunks, keeping metadata and relationships such as which node came before. Indexes store nodes, most commonly as embeddings in a vector store. Retrievers fetch candidate nodes for a query, optional postprocessors rerank or filter them, and a response synthesiser builds the prompt and produces an answer. A query engine bundles retriever and synthesiser behind one call.
The value is that each stage is a replaceable part with sensible defaults, and the framework includes the retrieval techniques that are tedious to build: hierarchical and sentence-window chunking, metadata filters, hybrid search, recursive retrieval over document summaries, and evaluation helpers. Beyond RAG it offers Workflows, an event-driven way to compose steps, and agent abstractions that can use query engines as tools. The company also runs a commercial cloud service, including document parsing for complex PDFs; the open-source library does not require it.
Choose LlamaIndex when the hard part of your application is the data: messy documents, many sources, retrieval quality that needs iteration. If your retrieval is a single SQL query or a small, clean corpus, the framework is more machinery than you need.
Defaults are a starting point, not a design. The default chunk size, top-k and prompt templates were chosen to work on average, and your documents are not average. Measure retrieval quality on your own questions before and after each change.
CrewAI is an open-source Python framework for multi-agent systems in which you describe agents by role, goal and backstory, assign them tasks, and run them as a crew with a sequential or manager-led process, plus Flows for more deterministic orchestration.
CrewAI models a team. An agent has a role ("market researcher"), a goal, a backstory that shapes its persona, an LLM and a set of tools. A task has a description, an expected output and an assigned agent, and can take other tasks' outputs as context. A crew groups agents and tasks and runs them under a process: sequential, where tasks run in order and each output feeds the next, or hierarchical, where a manager agent decides which agent handles what and checks results. The framework turns those definitions into prompts and runs the tool-calling loops for each agent.
The appeal is that a role-based description is quick to write and easy for non-specialists to read. For tasks that genuinely decompose into separate roles with distinct tools, such as a researcher with web search and a writer with a style guide, it gets a working prototype running in an afternoon. CrewAI also offers Flows, which wire steps together with explicit, event-driven control and state, so you can mix fixed logic with crews where autonomy helps. The company sells a hosted platform; the library runs on its own.
The trade-off is that much of the behaviour lives in generated prompts and agent-to-agent hand-offs you did not write. A hierarchical crew can spend many model calls deciding who should do what, and errors compound as outputs pass between agents. Multi-agent designs are not automatically better than one well-prompted agent with good tools; they are better when roles need different context, tools or permissions.
Treat role descriptions as prompts: test them, version them, and read the traces to see what each agent actually received.
n8n is a workflow automation platform with a visual node editor, hundreds of app integrations, triggers such as webhooks and schedules, and AI nodes for model calls and agents. It can be self-hosted under its source-available licence or used as a hosted service.
An n8n workflow is a set of nodes joined on a canvas. A trigger node starts it: an incoming webhook, a schedule, a new form submission, a message in a chat tool. Action nodes call services such as a CRM, a spreadsheet, email or a database, and logic nodes branch, merge, loop and wait. Data passes between nodes as JSON items, and a Code node lets you drop into JavaScript or Python when the built-in nodes run out. Each run is stored as an execution you can open to see the input and output of every node.
Its AI features include nodes for calling chat models, embedding text, working with vector stores, and an agent node that lets a model pick from tools you connect to it. Those AI nodes are built on LangChain's JavaScript library under the hood. This makes it practical to put a model step in the middle of an existing business process, such as classifying an inbound email, summarising it, and creating a ticket in the right queue, without writing a service.
n8n is distributed under a fair-code, source-available licence rather than a standard open-source one: you can self-host it for internal use, but there are restrictions on reselling it as a service, so check the current licence terms for your case. Self-hosting gives you data control, and it can scale out with a queue mode that spreads executions across workers.
The trade-off is that visual workflows are harder to review, test and version than code. Logic spread across thirty nodes is easy to build and hard to change safely. Teams often use n8n for integration-heavy glue and keep complex AI logic in a service the workflow calls.
Ollama is an open-source tool that downloads, runs and serves open-weight models on your own machine with one command, exposing a local HTTP API (including an OpenAI-compatible endpoint) so applications can call a local model like a hosted one.
Running an open-weight model yourself normally means choosing an inference engine, finding a compatible file, picking a quantisation and wiring up an API. Ollama packages those steps. ollama pull fetches a model from its library, usually already quantised (stored at lower numeric precision, such as 4-bit, so it fits in less memory). ollama run starts a chat in the terminal, and a background server answers HTTP requests on a local port. It runs on macOS, Linux and Windows, uses the GPU when one is available, and falls back to CPU, much more slowly, when it is not.
Under the hood Ollama uses an inference engine descended from llama.cpp and the GGUF model file format from that project. A Modelfile lets you define a variant: a base model plus a system prompt, parameters such as temperature and context length, or an adapter. Because it offers an OpenAI-compatible endpoint, many tools and SDKs work against it by changing the base URL.
It is designed for a single user or a small team: local development, offline demos, privacy-sensitive experiments, and testing whether an open model is good enough for a task. It is not designed as a high-throughput multi-user server. Engines built for serving, such as vLLM, batch many concurrent requests far more efficiently on data-centre GPUs.
Two settings cause most surprises. The context window Ollama allocates by default is often smaller than the model supports, and input beyond it is silently truncated, so long prompts appear to be ignored. And the quantisation level changes quality: a 4-bit model is not the same as the full-precision model whose benchmarks you read.
Hugging Face runs the Hub, the main public registry of open models, datasets and demo apps, and maintains widely used open-source libraries for loading, training and fine-tuning models, plus paid services for hosted inference and compute.
The Hub is a set of Git-based repositories for models, datasets and Spaces (small hosted demo apps). A model repository holds weights, configuration, tokenizer files and a model card describing training data, intended use, limitations and licence. Dataset repositories work the same way with dataset cards. Many organisations publish their open-weight models there first, and some models are gated, meaning you must accept a licence or request access before downloading.
The libraries are the second pillar. transformers loads and runs thousands of model architectures behind a common interface. datasets loads and processes large datasets efficiently. tokenizers provides fast tokenisation. peft implements parameter-efficient fine-tuning such as LoRA, trl covers supervised fine-tuning and preference optimisation, and accelerate handles multi-GPU and mixed-precision training. The safetensors format, which stores weights without executable code, came from this ecosystem and is now the common default.
On the commercial side, Hugging Face offers dedicated inference endpoints, a routing layer to third-party inference providers, paid compute for Spaces, and enterprise features such as private organisations and access controls. Availability and terms change, so check current offerings.
Treat the Hub like any public package registry. Anyone can upload, so trust comes from the publisher, the model card and your own evaluation, not from download counts. Two specific risks: older pickle-based weight files can execute code when loaded, and some models require a trust_remote_code option that runs Python from the repository. Prefer safetensors and review remote code before enabling it.
vLLM is an open-source inference engine for serving language models at high throughput. PagedAttention manages KV-cache memory in blocks, continuous batching keeps the GPU busy, and an OpenAI-compatible server lets applications call self-hosted models through a familiar API.
The bottleneck in serving an LLM to many users is usually memory, not arithmetic. Each active request holds a KV cache, the stored attention keys and values for every token so far, and it grows with every generated token. Naive servers reserve one contiguous slab per request sized for the maximum length, wasting most of it. vLLM's PagedAttention, introduced in the 2023 paper Efficient Memory Management for Large Language Model Serving with PagedAttention, stores the cache in fixed-size blocks allocated on demand, like virtual memory pages in an operating system. Less waste means more concurrent requests fit on the same GPU, and blocks can be shared between requests with a common prefix.
The second mechanism is continuous batching. Instead of waiting for a whole batch to finish, the scheduler adds new requests and removes finished ones at every decoding step, so the GPU never idles while one long answer completes. On top of that vLLM supports automatic prefix caching (reusing the cache for repeated system prompts), tensor and pipeline parallelism across GPUs, many quantisation formats, speculative decoding, structured output constrained to a JSON schema or grammar, and serving several LoRA adapters on one base model.
Operationally, vllm serve starts an HTTP server with OpenAI-compatible chat and completion endpoints, so application code written for a hosted API often works with a changed base URL. By default it pre-allocates most of the GPU's memory for weights and cache at start-up, which is efficient but means you should not share that GPU with other processes without lowering the fraction.
vLLM is one of several serving engines; others take different approaches to scheduling, kernels and hardware support, and the right choice depends on your models, hardware and traffic. Whatever you choose, the trade-offs are the same: maximum context length reduces how many requests fit, batching raises throughput but can raise per-request latency, and quantisation saves memory at some quality cost. Load-test with your real prompt and output lengths.
LiteLLM is an open-source Python library and proxy server that puts many model providers behind one OpenAI-style interface, translating requests and responses, and adds gateway features such as fallbacks, load balancing, per-key budgets, rate limits and logging.
Providers differ in request format, message roles, tool-calling schemas, error types and streaming chunk shapes. LiteLLM translates between a single OpenAI-style request and each provider's native API, then normalises the response back. You name the model with a provider prefix and the library handles the rest, including mapping provider-specific errors onto a common set of exception types so retry logic works the same everywhere. It also tracks token usage and estimated cost per call.
It comes in two forms. The SDK is imported into your Python code. The proxy is a standalone server (an AI gateway) that any language can call as if it were one OpenAI-compatible endpoint. The proxy reads a configuration file listing model deployments under friendly aliases, and adds the controls a platform team wants centrally: virtual API keys per team or application, spend budgets, rate limits, routing and load balancing across deployments of the same model, fallbacks to another model when one fails or times out, caching, and callbacks that send logs and traces to observability tools.
The value is decoupling. Application code asks for chat-default; the platform team decides which provider and model that alias maps to, can shift traffic during an outage, and can see spend per team without collecting provider keys from every service. That is the same role commercial AI gateways play, and some teams build a thin version themselves.
The limits come from translation. A common interface covers the common features; provider-specific capabilities (particular caching controls, reasoning settings, new tool types, file handling) may arrive later, behave slightly differently, or need pass-through parameters. A fallback model also behaves differently from the primary, so a silent fallback can change output quality and format. And the proxy is now a component on your critical path that needs monitoring, redundancy and upgrades like any other service.
DSPy is an open-source Python framework, from Stanford NLP, that replaces hand-written prompts with declared input and output signatures composed into modules, then uses optimizers to search for instructions and few-shot examples that maximise a metric you define on your own data.
Hand-tuned prompts are brittle: they are tuned to one model, break when a step changes, and encode guesses about what works. DSPy treats the prompt as a parameter to learn. You declare signatures, the inputs and outputs of a step, such as a question and context going in and an answer coming out, with optional descriptions. You compose modules that implement a strategy for a signature, such as a direct prediction, step-by-step reasoning before answering, or a tool-using agent loop. A program is ordinary Python that wires modules together, for example a retrieval step feeding an answering step.
The distinctive part is the optimizer (earlier versions called them teleprompters). You supply a training set of examples and a metric, a Python function that scores an output. The optimizer runs your program, collects traces where the metric is high, and uses them to choose few-shot demonstrations, rewrite instructions, or both. Some optimizers search over candidate instructions with a model proposing variants; others can fine-tune model weights from successful traces. The result is a compiled program whose prompts were chosen by measurement rather than intuition, and you can re-compile when you change models.
The quality of the result is bounded by the metric and data. A metric that rewards the wrong thing will be optimised precisely, which is the familiar problem of overfitting to the eval. Optimization also costs many model calls, so budget it, and hold out a test set the optimizer never sees.
DSPy fits when you have a multi-step pipeline, a measurable definition of good, and at least tens to a few hundred labelled examples. It fits less well when quality is mostly subjective, when you have no labelled data, or when the team needs to read and hand-edit every prompt. Many teams use it to discover good prompts, then inspect and keep what it found.
PydanticAI is an open-source Python agent framework from the Pydantic team that makes agent inputs, dependencies, tools and outputs typed: outputs are validated against Pydantic models, failures are fed back for retry, and dependencies are injected so agents are testable.
Most agent bugs at integration boundaries are type bugs: a missing field, a string where a number was expected, a date in the wrong format. PydanticAI applies Pydantic, Python's widely used data validation library, to every boundary. You declare an agent with a model, instructions and an output type, usually a Pydantic model. The framework turns that type into a JSON schema the model must follow, validates the response, and if validation fails it sends the error back to the model and asks again, up to a retry limit. Your code receives a typed object, not a string to parse.
Tools are ordinary Python functions registered on the agent; their type hints and docstrings become the tool schema, and arguments are validated before your function runs. Dependency injection passes run-time resources such as a database connection, the current user or an HTTP client to tools and instructions through a typed context object, instead of globals. That makes tests straightforward: inject a fake database and a test model, and the agent runs offline and deterministically.
The framework is model-agnostic across many providers, supports streaming of partially validated structured output, and integrates with OpenTelemetry-based tracing, including the Pydantic team's own observability product, though it does not require it. For multi-step control flow it offers a typed graph library, and agents can call other agents as tools.
The trade-off is scope. PydanticAI is deliberately close to plain Python, so it gives you less pre-built orchestration and fewer integrations than larger frameworks. And validation proves shape, not truth: a perfectly typed invoice with the wrong total still passes. Add business-rule validators, and evaluate content separately.
Langfuse is an open-source LLM engineering platform, self-hostable or hosted, that records traces of model calls, tool calls and retrieval, tracks cost and latency, manages versioned prompts, and runs evaluations on production traces and test datasets.
A request through an AI feature can involve retrieval, several model calls and tool executions. Langfuse captures that as a trace: a tree of nested observations, where each step is a span and model calls are recorded as generations with the model name, parameters, input and output messages, token counts, cost and latency. Traces can carry a user id, a session id to group a conversation, tags and metadata. You instrument through its SDKs, decorators, integrations with common frameworks, or OpenTelemetry, which its current SDKs are built on.
Around traces it adds the loop teams need after launch. Prompt management stores prompts with versions and labels such as production, so an application fetches the labelled version at run time and a trace records which version produced each answer. Scores attach quality signals to traces: user feedback, code checks, model-graded evaluations run automatically on a sample of traffic, and human labels from annotation queues. Datasets and experiments let you run a prompt or application version over a fixed set of cases and compare results, which connects offline evals to production data.
Being open source and self-hostable matters when traces contain sensitive data: prompts and outputs often include personal or confidential information, and self-hosting keeps them inside your infrastructure. The trade-off is operating it; the self-hosted deployment uses several data stores and needs capacity planning as trace volume grows. A managed cloud option exists, and some advanced features may differ between editions, so check current terms.
Whatever tool you choose, the hard work is the same: deciding what to log, redacting what must not be stored, sampling at high volume, and defining scores that reflect real quality.
Arize Phoenix is an open-source observability and evaluation tool from Arize AI that collects OpenTelemetry traces of LLM applications, visualises retrieval and tool steps, and runs evaluations and experiments over traces and datasets, locally or self-hosted.
Phoenix is built on OpenTelemetry, the vendor-neutral standard for traces, using a set of semantic conventions for AI called OpenInference. These conventions define how to record an LLM call, a retrieval, a reranking step, an embedding or a tool call as span attributes. Auto-instrumentation packages exist for many provider SDKs and frameworks, so adding a few lines at start-up produces traces without changing application code. Because the data is standard OpenTelemetry, the same spans can also be sent to other backends.
Phoenix can run as a local app during development, including inside a notebook, or as a self-hosted service for a team. Its interface shows each trace as a tree, with retrieved documents and their scores visible next to the generation that used them. That makes retrieval problems easy to see: the right document ranked eighth, a chunk cut mid-sentence, or an empty result that the model answered anyway.
On the evaluation side it provides model-graded evaluators for common questions, such as whether retrieved documents are relevant, whether an answer is grounded in them, and whether a response is toxic, and lets you write your own. You can run evals over collected traces, store the results as annotations on spans, build datasets from interesting traces, and run experiments that compare application versions on those datasets. A prompt playground lets you replay a captured call with a changed prompt or model.
Arize also sells a commercial platform for larger-scale production monitoring; Phoenix itself is free to run, under its own licence, so check the terms if you plan to redistribute it. As with any built-in evaluator, calibrate its judgements against human labels on your own data before trusting the numbers.
Use a framework when it removes complexity you would otherwise build and maintain, such as persistence, integrations or tracing, without hiding the prompts, control flow and errors you must understand. For a few direct model calls, the provider SDK and plain code are usually clearer.
A framework is a bet that its authors' abstractions match your problem. When they do, you get tested code for hard parts: durable state and resume, dozens of data connectors, streaming, retries, tracing hooks. When they do not, you fight the abstraction, debug through layers you did not write, and inherit upgrade churn. The GenAI ecosystem moves quickly, and major frameworks have reorganised their APIs more than once, so the cost of a dependency includes future migrations.
A useful test is to list what your application actually needs and what you would have to build without the framework. A single structured-output call with validation needs a provider SDK and a schema library. A RAG system over messy multi-source data benefits from ingestion and retrieval components. A long-running agent with human approvals benefits from a runtime with checkpoints. Model access across many teams benefits from a gateway. Observability benefits from a tracing tool almost from day one, because you cannot fix what you cannot see.
Keep the parts that define your product's behaviour visible: the exact prompt, the tool definitions, the control flow and the error handling. Many teams settle on a mixed approach, using framework components for integrations and infrastructure while writing orchestration as plain code. Others adopt a framework's runtime fully and wrap it behind their own interface so it can be replaced.
Before adopting one, check basic health signals: how actively it is maintained, how breaking changes have been handled, whether you can see the raw requests it sends, whether it supports the provider features you rely on, and whether its licence fits how you ship. Then build a small spike of your hardest case, not the tutorial case, and measure how much code and confusion it saved.
Put your own thin interfaces at the boundaries that change: model calls, retrieval, tracing and orchestration. Keep prompts, schemas and eval sets as plain files you own, and use open standards such as OpenTelemetry and OpenAI-compatible APIs where they exist.
Lock-in in this ecosystem rarely comes from a contract. It comes from code: framework-specific classes spread through business logic, prompts stored only inside a vendor's platform, traces in a proprietary format, and features that only one provider offers. Each makes the next change more expensive. Because models, providers and frameworks change faster than most application code, the ability to swap a part is worth designing for.
The main technique is a narrow internal interface at each boundary. Application code calls your generate() or retrieve() function with your own request and response types; one adapter module translates to the provider SDK or framework. Swapping a provider then means writing one adapter and running the eval set, not editing every call site. Keep these interfaces small and shaped around your needs; a wrapper that mirrors every provider option has merely moved the lock-in.
The second technique is to own the assets that carry quality: prompts, output schemas, tool definitions and, above all, the eval dataset. Store them in your repository in plain formats. An eval set is what lets you prove that a replacement model or framework is as good as the current one; without it, every migration is a leap of faith.
Third, prefer open standards where they are mature. OpenTelemetry-based tracing means your instrumentation survives a change of observability backend. OpenAI-compatible endpoints are offered by many providers and self-hosted servers, which makes routing and fallback easier, although each implementation supports a different subset of features. The Model Context Protocol (MCP) gives a common way to expose tools to different agent hosts.
Accept some lock-in deliberately. Provider-specific features such as particular caching or reasoning controls may be worth using, as long as the dependency is isolated in the adapter and you know what you would lose by switching.
Many model providers publish agent SDKs that wrap their APIs with a tool-calling loop, tool definitions, hand-offs, guardrails and tracing. They are thin and track new model features quickly, but are usually tuned for that provider, unlike independent frameworks built for many.
Calling a model with tools requires a loop: send the request, check whether the model asked for a tool, run it, send the result back, and repeat until the model gives a final answer or a limit is reached. Provider agent SDKs package that loop along with the pieces around it: declaring tools from functions, connecting to MCP servers, limits on turns, hand-offs between specialised agents, input and output checks, streaming, and built-in tracing. Some also expose hosted tools the provider runs, such as web search or code execution, and some mirror the runtime the provider uses in its own agent products.
Their strengths follow from who builds them. They support new model capabilities, such as new tool types, reasoning settings or caching controls, early and correctly, because the same company ships the model. They tend to be small, with few abstractions between you and the API. Several can call other providers' models too, through compatible endpoints or adapters, but features are typically designed first for the home provider.
Independent frameworks make the opposite trade. They aim for provider neutrality and a broad integration catalogue, and add higher-level structure such as graph runtimes, data connectors or role-based crews. They may lag behind a brand-new provider feature.
The choice depends on what you expect to change. If you are committed to one provider and want the closest fit to its features, its SDK is a natural start. If you need to route across several providers, or want orchestration features the SDK lacks, an independent framework or your own thin loop may fit better. In either case, keep tools and prompts in your own code so the SDK can be replaced.