The GenAI Field Guide

AI-native software development

Use AI across the development cycle while keeping human judgment on the consequential parts.

AI-native software development is what happens when models take part in every stage of building software: shaping requirements, comparing designs, prototyping interfaces, writing code and tests, reviewing changes and supporting incidents. The interesting change is not that a model can write code. It is that drafting anything (code, tests, specs, analyses) becomes cheap, so the expensive parts of the job move to deciding what to build, specifying it precisely and proving that what was produced is right.

The mental model is a pipeline whose bottleneck has moved. Generation is fast; verification, review and understanding are not. Teams that gain from AI invest in the parts that make output easy to accept or reject: written specs, repository instructions for agents, fast and meaningful tests, gates that fail loudly, risk-based review and staged releases. Teams that only add generation get more code, more review load and, after a while, more defects.

The Basic questions walk through the lifecycle stage by stage, showing where a model helps and where it needs a person: requirements, architecture, UI and UX, tests, code review and incidents. The Advanced questions are about the operating model: which work must stay under human review, which skills grow in value, how a senior engineer's role changes, how to verify AI-assisted code, what to automate first, how to measure whether any of it helps, and the new risks generated code brings.

What is AI-native software development?

AI-native software development is a way of building software where models take part in every stage, from requirements to operations, while the team redesigns its process around fast generation and rigorous verification rather than bolting a chat window onto the old workflow.

Most teams start with AI-assisted development: autocomplete in the editor and a chat tab for questions. The process is unchanged; individuals type a little faster. AI-native development is a different operating model. The team assumes that drafting code, tests, documents and analyses is cheap, and reorganises around the things that are still expensive: deciding what to build, writing a precise specification, and proving that what was produced is correct.

The mechanism that makes this work is the shift of the bottleneck. When a coding agent can turn a clear spec into a plausible implementation in minutes, the slowest steps become ambiguity in the requirement and the human time needed to review the result. AI-native teams therefore invest in written specs, repository instruction files such as AGENTS.md, fast and trustworthy test suites, and automated gates, because those are what let a model's output be accepted or rejected quickly.

In practice this touches every stage. Requirements are drafted and stress-tested by a model before anyone estimates them. Architecture options are compared and challenged in writing. Agents implement scoped tasks on branches and open merge requests. Model reviewers comment before a human does. During incidents, an assistant pulls together deploys, logs and metrics into a first set of hypotheses. People remain accountable for every decision that ships.

The trade-off is real. Generation speed rises sharply, but so does the volume of code a team must understand and maintain, and review becomes the new constraint. Teams that adopt agents without strengthening tests and review often ship faster for a few weeks and then slow down under defects and code nobody fully understands. AI-native is less about how much code the model writes and more about how reliably the team can tell good output from bad.

How can AI help with requirements?

AI helps requirements by drafting stories from rough notes, then attacking the draft: finding missing rules, contradictions, undefined terms and edge cases, and turning each into a question for a person who owns the answer.

Requirement defects are the most expensive kind because they survive into code, tests and documentation before anyone notices. A model is useful here less as an author and more as a tireless reader. Given a draft spec, it can list every place where a term is undefined, two statements conflict, a state has no exit, or a numeric boundary is missing. It does this by pattern-matching against the many specifications and bug reports it saw in training, which is exactly the experience a junior analyst lacks.

The best workflow keeps the model from inventing answers. Ask for questions, not decisions: "refunds within 30 days" should produce "30 calendar days or business days, measured from order or delivery, in which time zone?" rather than the model quietly picking one. Each question then goes to the product owner, and the resolved answer is written back into the spec. Models are also good at converting a resolved spec into acceptance criteria in a given-when-then form, and at generating boundary cases that later become tests.

Context matters. A model that sees only the new feature will miss conflicts with existing behaviour. Feed it the related specs, the data model, the relevant policy text and examples of past incidents in the area, and it will find gaps that matter to your system rather than generic ones.

The limit is that a model cannot know your business. It does not know the regulator's position, the sales promise made to a key customer, or why a rule exists. Treat its output as a checklist of questions, weigh them, and drop the ones that do not apply. A list of 40 generic questions wastes more time than it saves; ask it to rank by impact and cap the list.

How can AI help with architecture?

AI helps architecture by mapping the existing system from code, laying out options with their trade-offs, and arguing against a proposed design. It is a fast sparring partner, but it cannot see your load, team or operating history unless you give them to it.

Architecture decisions depend on two kinds of knowledge: general patterns (queues, caches, consistency models, failure modes) and specific facts about your system (traffic, data sizes, team skills, what broke last year). Models are strong on the first and blind to the second. Used well, they bring the breadth of a well-read architect to a team; used badly, they produce confident textbook answers that ignore your constraints.

Three uses pay off reliably. First, mapping: a coding agent can read a repository and produce a dependency map, a list of services and the calls between them, or a summary of where a given table is written. This is tedious work humans skip, and it grounds every later discussion. Second, option generation: given the problem and constraints, a model can list three to five designs with their failure modes, cost drivers and operational burden, which prevents the team from anchoring on the first idea. Third, challenge: ask the model to argue against the chosen design, list its weakest assumptions, and propose the cheapest experiment that would prove each one wrong.

Give the model the numbers. "Should we use a queue?" gets a generic answer. "Peak 400 orders per minute, the payment provider times out at 30 seconds about once a day, and we have two engineers on call" gets a specific one. Write the outcome as an architecture decision record (ADR), a short document capturing context, decision and consequences, so the reasoning survives the conversation.

Verify any factual claim about a library, cloud service limit or protocol against current documentation. Models mix up versions and state outdated limits with full confidence, and those details often decide the design.

How can AI help with UI and UX?

AI helps UI and UX by generating several layout and copy alternatives quickly, turning sketches into working prototypes, and checking interfaces for accessibility and consistency. Choosing what users actually need still requires research with real users.

Design benefits from breadth early: the first idea is rarely the best. Models make breadth cheap. From a short brief, a model can produce three or four distinct layouts as working HTML or component code, write alternative microcopy for buttons and error states, and fill screens with realistic sample data so a prototype feels real in a usability session. What used to take a designer a day of mock-ups becomes a morning of choosing and refining.

Prototypes are where the gain is largest. A clickable prototype built from your own component library can go in front of five users the same week, and the cost of throwing it away is near zero. That changes behaviour: teams test more ideas instead of defending the one they invested in.

Models also help with checks. They can review a screen for missing labels, unclear error messages, inconsistent terminology and contrast problems, and they can read a component diff and flag a focus trap or a missing keyboard handler. Pair this with a deterministic accessibility scanner: the scanner is reliable for rule violations such as contrast ratios and missing alt text, while the model is better at judgement calls such as whether an error message tells the user what to do next.

The limits are taste and evidence. Models tend toward the average of what they have seen, which produces generic, interchangeable screens unless you give them a clear design system and direction. And no model knows whether your users understand a flow; only observing users does. Treat generated designs as hypotheses to test, never as validated answers.

How can AI write tests?

AI writes useful tests when it works from the specified behaviour and boundaries rather than from the implementation. It can enumerate cases, write the code and fill in fixtures quickly, but each test must be checked to fail when the behaviour breaks.

A model can write a hundred tests in a minute, which makes test count meaningless as a quality signal. The question is whether each test would catch a real defect. The most common failure is the mirror test: the model reads the implementation and writes assertions that restate what the code does, bugs included. Such tests pass today and will keep passing when the code is wrong in the same way.

The fix is to change the input. Give the model the spec or acceptance criteria, not just the function body, and ask it to enumerate cases first: the happy path, each boundary (zero, one, maximum, just over), each error path, and each permission. Review that list, because it is far quicker to check twenty one-line case names than twenty test bodies. Then ask it to implement the tests. Where possible, write the tests before the implementation, so the code is shaped to pass behaviour the team agreed on.

Check that the tests can fail. A simple discipline is to break the code on purpose (invert a condition, remove a check) and confirm a test goes red. Mutation testing tools automate this by making many small changes to the code and reporting which ones no test notices; it is the most direct measure of whether generated tests have teeth. Watch for weak assertions too, such as only checking that a response is not empty, and for over-mocking that tests the mocks rather than the system.

Models are especially good at the tedious parts: building fixtures, generating realistic test data, converting a bug report into a failing regression test, and adding property-based tests that state an invariant ("decoding an encoded value returns the original") and let a library search for counterexamples.

How can AI review code?

AI reviews code by reading a diff with its surrounding context and flagging likely bugs, missing checks, risky patterns and unclear code for a human reviewer. It is a strong first pass, not an approver: it misses intent and design problems and produces false positives.

A model reviewer reads the diff, the files it touches and ideally the related tests and spec, then comments on lines that look wrong. It is good at the things tired humans skip: an unchecked return value, a query built by string concatenation, a missing await, an off-by-one in pagination, a new endpoint without the authorization decorator its siblings have, or a renamed field still used elsewhere. It never gets bored on the fortieth file of a large change.

Its weaknesses are just as predictable. It cannot know whether the change does what the product owner wanted, whether the design fits the system's direction, or whether a tricky pattern is intentional. It produces false positives, and if more than a few of its comments are noise, people stop reading all of them. It also tends to review only what it can see; a bug caused by a caller in another file is missed unless the review tool retrieves that context.

Good setups make the reviewer specific. Give it a checklist derived from your own incident history (authorization on every handler, money as integer minor units, no logging of personal data), ask it to cite the exact line and a concrete failure scenario for each finding, and rank findings by severity. A second pass that tries to refute each high-severity finding, by a separate model call or a person, cuts noise further. Track the acceptance rate of its comments, and tune the prompt when it drops.

Keep humans accountable for approval. The model narrows where a reviewer looks; the reviewer decides whether the change is right.

How can AI help with incidents?

AI helps incidents by gathering the evidence a responder needs, such as recent deploys, error changes, metrics and similar past incidents, into one summary with ranked hypotheses. People still decide on and carry out mitigation, especially anything that changes production.

The first twenty minutes of an incident are mostly searching: what changed, when did the errors start, which services are affected, has this happened before. That search is exactly what a model with read access to your tools does well. It can pull deploys and config changes in the window, compare error signatures before and after, read the top stack traces, and find past postmortems that mention the same symptom, then present them as a short timeline with hypotheses ranked by evidence.

The value is correlation across sources that live in different tabs. A latency spike that begins four minutes after a deploy touching a database query, alongside a jump in slow-query logs for that table, is a strong lead. A model can state that link in one paragraph and cite the evidence for it, which speeds up the human's judgement rather than replacing it.

Keep the boundary clear. The assistant should have read-only access by default. Mitigations such as rollback, feature-flag changes or scaling should be proposed with the exact command and the evidence, then run by a person or by a pre-approved runbook action with its own safeguards. A wrong hypothesis acted on automatically during an outage can turn one incident into two.

Models also help after the incident. They can draft the timeline from chat and alert history, summarise the contributing factors and propose action items, which removes the chore that makes postmortems late. A person should edit the result for accuracy and blame-free framing. Beware hallucinated causality: a model will produce a plausible story even when the evidence is thin, so require each claim to cite a log line, metric or change.

Which work still needs human review?

Humans must review work where a mistake is costly, hard to reverse or a matter of values: product and policy choices, security and permission boundaries, money and data movement, irreversible operations, and the evidence that a change is correct.

The rule of thumb is to scale review with the blast radius of a mistake, not with how hard the work was to produce. A model can write a database migration as easily as a CSS tweak, but a wrong migration can destroy data while a wrong colour costs a redeploy. As generation gets cheaper, the effort it took stops being a signal of risk, so teams need an explicit way to classify risk instead.

Four categories need a person in every case. Decisions about what to build and for whom, because they encode values and trade-offs no model is accountable for. Security boundaries: authentication, authorization, secret handling, input validation at trust boundaries, and anything that changes who can see what. Irreversible or high-cost operations: schema migrations, data deletion, payment flows, infrastructure changes, public communications. Evidence of correctness: someone must look at what the tests actually check and judge whether that proves the change works, because an agent can make a weak test pass.

Low-risk work can be reviewed more lightly: formatting, documentation drafts, test scaffolding, internal tooling with easy rollback, dependency bumps covered by a strong suite. Lighter does not mean none; it means a quick skim plus automated gates instead of line-by-line reading.

Make this mechanical rather than a matter of mood. Code owners on sensitive paths (auth, billing, migrations, infrastructure) force a named reviewer. Labels or path rules can route a change to deeper review automatically. Track which approvals were real: a reviewer who approves a 2,000-line agent-written change in three minutes did not review it, and the process should make that visible.

Human review has a capacity limit. If agents generate more than reviewers can genuinely read, the answer is smaller changes and stronger automated checks, not faster approvals.

What becomes more valuable for developers?

As writing code gets cheaper, value shifts to the work around it: framing the problem precisely, understanding the domain, designing systems, judging and verifying output, and communicating decisions. Reading code critically becomes as important as writing it.

Economics explains the shift. When one input to a process becomes cheap, the scarce inputs around it gain value. Typing code is now cheap. What remains scarce is knowing which problem is worth solving, describing it so precisely that a model cannot misread it, and recognising whether the result is right. Those skills were always part of the job; they were just bundled with a lot of typing.

Problem framing and specification come first. A vague task produces plausible but wrong code at high speed. A developer who can write a short spec with explicit constraints, examples and non-goals gets usable output on the first pass. Domain knowledge is the second: knowing that a refund cannot exceed the captured amount, or that a dosage field needs units, is what turns generic code into correct code, and models do not have your domain rules.

Architecture and system thinking matter more because agents optimise locally. An agent will happily add a fourth caching layer or a second way to do the same thing. Someone has to keep the system coherent, decide where boundaries go and say no to complexity.

Verification and judgement may be the most valuable skill of all. That means reading a diff critically, knowing which test would prove real user value, spotting a weak assertion, debugging a failure the agent could not fix, and knowing when to throw away generated work and start again. Debugging in particular stays human-heavy, because it requires forming and testing hypotheses about a specific system.

There is a risk for early-career developers. The tasks that used to teach fundamentals are the ones agents now do. Teams should deliberately keep learning paths: have juniors write some code by hand, explain generated code in review, and debug without assistance at times. A developer who never built the skill cannot verify the model's work.

What is a senior engineer's role in an AI-native team?

A senior engineer sets the boundaries agents and people work within, owns the verification system, makes the trade-offs that span components, reviews the risky changes and grows the team's judgement. The role moves from producing code to shaping the system that produces it.

On an AI-native team, the senior engineer's leverage comes less from their own output and more from the environment everyone else, people and agents, works in. An agent with a clear instruction file, a fast test suite and well-defined module boundaries produces good work; the same agent in a messy repository with flaky tests produces a stream of plausible defects. Building that environment is senior work.

Concretely, the role has five parts. Set boundaries: decide module ownership, which areas agents may change freely, which need extra review, and which tools and permissions agents get. Own verification: keep the test suite fast and meaningful, make sure CI gates fail loudly (including when they checked nothing), and define the rollout criteria for critical services. Resolve trade-offs: cross-cutting decisions on consistency, cost, latency and complexity, recorded as ADRs. Review what matters: spend scarce attention on the changes in the critical tier rather than spreading it thin. Grow capability: teach others to write specs, read diffs and debug, and turn recurring review comments into checks or instructions so they never need saying again.

That last point is a multiplier. Every time a senior engineer catches the same mistake twice, it should become a lint rule, a test, a line in AGENTS.md or a reusable skill. Over time the team's standards are encoded in the harness rather than in one person's head, which is how quality scales with agent volume.

The trap is becoming a full-time reviewer of agent output. If seniors spend all day approving pull requests, the system has a throughput problem that more reviewing will not solve. The fix is upstream: smaller tasks, better specs and stronger automated checks.

How should AI-assisted code be verified?

Verify AI-assisted code in layers: automated gates (types, lint, tests, security scans) that must pass and must have checked something, a human review of the diff and the tests, a run of the real user workflow, and monitoring after a staged release.

AI-assisted code fails differently from human code. It is usually syntactically clean and well formatted, which makes it look trustworthy, but it can call functions that do not exist, use an API in a deprecated way, handle the happy path while ignoring an edge case the spec mentioned, or quietly change behaviour outside the task. Verification has to target those failure modes, not just style.

Start with automated gates, run locally by the agent and again in CI: type checks, linting, the unit and integration suites, static security analysis, secret scanning and dependency checks. Each gate must fail on empty input and print how much it examined; a test runner that matched zero tests after a path change reports success while proving nothing.

Then a human review that reads two things: the diff, looking for scope creep and changes outside the task, and the tests, asking whether they would fail if the behaviour were wrong. Check that the agent did not weaken or delete an existing test to get green. Keep changes small; review quality collapses beyond a few hundred lines.

Next, exercise the real workflow. Run the feature end to end in a realistic environment: click through it, call the API with production-like data, check the database state after. Many integration bugs (wrong config, missing migration, broken permission in the real role) are invisible to unit tests.

Finally, release progressively and watch. Feature flags and canary releases limit the blast radius, and dashboards plus error tracking tell you within minutes whether the change behaves in production. Define the rollback trigger before release, not during an incident. Each layer catches what the previous one missed; skipping any of them shifts the cost to users.

What should be automated first?

Automate first the work that is frequent, tedious, checkable by a machine and cheap to get wrong. Changelog drafts, test scaffolding, dependency updates with a strong suite and documentation upkeep are good starts; payments, auth and data migrations are not.

A good first candidate has four properties. It is frequent, so the time saved adds up. It is tedious, so people are glad to hand it over and will not resist. Its output is checkable, ideally by a machine: a test passes, a build succeeds, a format validates. And its failure cost is low and reversible: a bad draft is edited or discarded with no harm done. Score candidates on these four and start with the highest.

Typical winners: drafting changelogs and release notes from merged pull requests, writing test scaffolding and fixtures, small dependency upgrades that a strong test suite can verify, updating documentation when code changes, converting bug reports into failing regression tests, triaging and labelling issues, and mechanical refactors such as renames or API migrations across many files where the compiler checks the result.

Poor first candidates are the opposite: rare, high-judgement or hard to verify. Payment logic, authentication, data migrations and production infrastructure changes have a high failure cost and need deep review, so automation saves little and risks a lot. Tasks without a clear check, such as "improve the architecture", produce output nobody can accept or reject efficiently.

Start in suggestion mode: the automation produces a draft or a merge request, and a person accepts it. Measure the acceptance rate and how much editing each output needs. When a task is accepted with little change for several weeks, raise its autonomy, for example by auto-merging when CI passes for a narrow class of changes. Keep a way to switch it off, and review a sample of auto-merged changes regularly.

How do you measure whether AI tools actually help a team?

Measure delivery outcomes, not activity: lead time, change failure rate, time to restore, review time and rework, compared before and after adoption or across a controlled rollout. Self-reported speed and lines generated are unreliable signals.

The easy metrics mislead. Lines of AI-generated code, suggestion acceptance rates and survey answers about feeling faster all measure activity or perception. Perception is a weak guide: a 2025 randomised study by the research group METR found experienced open-source developers took longer to complete tasks with AI tools on their own mature repositories, while believing they had been faster. Other studies on different tasks and populations have found speed-ups. The honest conclusion is that the effect depends on the task, the codebase and the developer, so you have to measure it in your own setting.

Measure what the business cares about: how fast a change gets from idea to production and how often it breaks things. The four DORA metrics (deployment frequency, lead time for changes, change failure rate and time to restore service) are a well-established starting point. Add measures specific to AI-assisted work: review time per merge request, rework rate (code changed or reverted within a few weeks of merging), escaped defects traced to generated code, and the acceptance rate of agent-opened merge requests.

Design the comparison so it can show a negative result. Compare similar teams or similar work before and after, or roll tools out to some teams first. Control for task mix, because agents may be pointed at the easy tickets. Look at distributions rather than averages: if median lead time drops but the slowest 10 percent of changes get slower because review queues grow, the bottleneck has moved, not gone.

Watch for costs that appear late. Code volume growing faster than the team's ability to understand it shows up months later as slower onboarding, harder debugging and rising defect rates. Track codebase size and churn alongside throughput, and ask in retrospectives which generated code people struggled to change.

What new risks does AI-generated code introduce?

AI-generated code adds risks beyond ordinary bugs: hallucinated or look-alike dependencies, outdated or insecure patterns, licence questions, leaked secrets, prompt injection against coding agents, and a growing body of code nobody on the team fully understands.

Hallucinated dependencies are the most concrete new risk. Models sometimes import packages that do not exist, with plausible names. Researchers have shown these names recur across runs, so an attacker can register them on a public package index with malicious code, a technique sometimes called slopsquatting. Typosquatted names that differ by one character are a related risk. The defence is mechanical: every new dependency must already be in a lockfile or an allowlist, or be approved by a person who checked its source, age and maintainers.

Insecure and outdated patterns come from training data. Models reproduce what was common, including code written before a vulnerability was understood: string-built SQL, weak hashing for passwords, disabled certificate checks in examples, deprecated crypto APIs. Static analysis and secret scanning should run on every change regardless of who wrote it. Licence risk arises when a model reproduces a recognisable block of licensed code; some tools offer filters that block suggestions matching public code, and your legal team may have a policy on their use.

Coding agents add an attack surface of their own. An agent that reads issues, web pages or dependency READMEs can meet text written to instruct it, an indirect prompt injection, and act on it with the permissions it holds. Run agents in a sandbox with scoped credentials, no production secrets and restricted network access, and require review before anything they produce is merged.

The slowest risk is comprehension debt: code that works but that no one on the team has really read. It costs nothing until an incident at 2 a.m. requires someone to understand it. Small changes, mandatory explanations in review and periodic walkthroughs of generated modules keep the team able to own what it ships.