Quick answer

An agentic development workflow is a process split into stages, where an AI agent decides its own next steps from what it observes, while a fixed and repeatable structure wraps that reasoning and keeps the overall process predictable.

Five stages carry it. An idea becomes a written spec, and the agent runs it inside a governed harness built from shared rules and Agent Skills. Verification goes to someone other than whoever built the thing, shipping carries evidence, and production feeds back into the spec that follows. Without that structure, output quality varies with whoever prompted the agent.

When a single engineer spends months driving real productivity through a coding agent, teammates naturally start borrowing pieces of the configuration. This organic adoption eventually creates a push for unified standards across the team. Searching for an agentic development workflow, however, usually yields two distinct paradigms. On one side lies the detailed documentation of a solo developer’s local environment. On the other, platform vendors outline six architectural pillars that require a dedicated internal engineering team to build and maintain. Neither model fits a mid-sized product group of twelve engineers.

Team size is only part of the divide. DORA’s 2025 State of AI-assisted software development report shows that artificial intelligence functions chiefly as an amplifier. The report points out that underlying organizational systems yield far higher returns than specific tool choices, magnifying whatever systemic traits already exist.

A five-stage AI software development lifecycle fills this gap directly within your existing repository. Teams can adopt this structured approach without forming a platform group or renovating infrastructure.

Two Extremes, and the Missing Middle

The clearest account of the solo end comes from Tim Deschryver, an open-source enthusiast who sits on the core teams of NgRx and Angular Testing Library. His AGENTS.md is deliberately short, and he goes back to prune it before it can rot. Agent Skills come from a registry, which lets one set of instructions serve more than a single project.

Larger work goes through OpenSpec, where spec-driven development turns out a proposal and a design, plus a tasks file to work from. In his framing, the AI rides as a domestique and the developer leads the team. One person built everything, and that same person checks it.

Port describes the far end. Its workflow orchestrator wires triggers and actions into one flow, with AI nodes and conditions deciding what happens in between.

Six pillars sit underneath, built once by a platform team. A Context Lake holds engineering context, with reusable actions and scorecards above it and the orchestrator tying them together. Access controls decide who runs what, and engineering intelligence reports on all of it.

Two Extreme Methods and the Missing Middle One
Two Extreme Methods and the Missing Middle One

Neither account is wrong for its reader. The gap opens between them: a growth-stage team has no platform function to staff and has outgrown one developer’s machine. It needs an agentic SDLC (software development lifecycle), where agents carry much of the implementation under rules the team wrote down.

The One Idea Worth Borrowing From Both: Deterministic Structure, Non-Deterministic Reasoning

Port articulates this distinction more clearly than anyone else writing about it. Some steps in a pipeline are deterministic. A trigger fires, a scorecard passes or fails, a gate opens once its conditions are met. Run them a hundred times, and they behave the same way every time.

The agent’s reasoning stays non-deterministic. It picks an implementation approach and reads intent out of a task description. Where it goes next depends on what it observed a moment earlier.

The One Idea Worth Borrowing From Both
The One Idea Worth Borrowing From Both

The fixed half is what you can write down, which makes this portable to a team of twelve. Those steps go into the repository, get reviewed like any other change, and apply the same way to whoever is at the keyboard.

Stage #1. Baseline & Frame: Turning an Idea Into a Spec

This is where you define what is ‘correct’, when changing the answer is still cheap. Scope goes down with acceptance criteria next to it, and so do the edge cases someone would otherwise discover during review. The spec also identifies which calls can be undone later, and which calls require sign-off before implementation starts. Reversibility is a measure of how much argument a decision deserves.

Both ends already run some form of spec-driven development. OpenSpec hands a single developer a proposal, a design, and a tasks file. Enterprise guidance has an agent draft the spec for a human to review before delegation: same discipline, different tools.

At SPD Technology, spec-first delivery adds a decision protocol that sorts the questions a spec raises by what it would cost to get them wrong. A reversible call gets a line in an architecture decision record (ADR); an expensive one earns the ceremony.

Our article on spec-first development breaks down the full framework behind this stage.

Stage #2. Design & Build: Executing Inside a Shared Harness

The harness is everything wrapped around the model that shapes its behavior. Deschryver splits the context inside it into two. AGENTS.md is a static context, a rules file loaded on every task. Agent Skills are dynamic context, loaded only when a task calls for them, which keeps the context window clear of what most tasks never need.

At solo scale, one laptop holds all of it. Team scale turns AGENTS.md and the Skills sitting beside it into a shared harness that lives in the repository and gets versioned the way code does. The other thing a shared harness buys you is that tool access becomes a decision somebody signed off on.

Practitioners worth listening to split over where that access should come from. Deschryver avoids Model Context Protocol (MCP) servers on context-bloat grounds, preferring Agent Skills and command-line tools he can inspect. Port treats MCP as core infrastructure for governed developer self-service.

Both positions hold. The question inside your team is whether each piece of tool access was deliberate and reviewed. Unreviewed access is a real risk, a well-governed MCP layer is a legitimate tool, and guardrails belong on either arrangement.

We go deeper into configuring and governing that layer in our agent harness engineering guide.

Stage #3. Assure: Verifying the Result Independently

Each account goes quiet at the same point. Deschryver reviews his own agent’s output carefully, and the judgment that directed the work is the only judgment checking it. Port puts a human checkpoint at each workflow stage, without separating a summary confirmation from verification of the change.

The principle underneath both gaps: the author of a change, human or AI, should never be its only judge. Human-in-the-loop stops meaning much once the human reads a document the agent wrote about itself.

Skipping independent verification has a measurable price. DORA’s research describes a verification tax: time saved writing code gets re-spent auditing it, and it reports that 30% of developers have little to no trust in AI-generated code.

Tests and evaluation are quite distinct. Tests confirm specific behaviors on demand, independent verification covers what no assertion can express, and neither replaces the other. What a review catches depends on which artifact the reviewer opens first.

Serhii Leleko:AI & ML Engineer at SPD Technology

Serhii Leleko

AI & ML Engineer at SPD Technology

“A reviewer who reads the agent’s summary of a change is reviewing a description written by the thing that made the change. The only review that catches a spec the agent quietly reinterpreted is one that starts from the spec itself and works toward the diff, in that order — which is why we treat the direction a reviewer reads in as part of the process rather than a matter of preference.”

​

In practice, the reviewer opens the spec first, writes down what acceptance looks like, and only then reads the diff against that list. Work the agent added past the spec shows up as surplus, and anything it quietly dropped shows up as a gap.

The second reader does not have to be your most senior engineer. Anyone working forward from the spec catches reinterpretation reliably, and a separate review profile with its own rules can take the mechanical part of the pass before a person starts reading.

Read our guide on independent evaluation and verification to go deeper into the framework.

Stage #4. Ship: Evidence, Not Just a Merged PR

What happens after the merge gets treated as an afterthought in both accounts. Attributable catalog changes are what Port logs agent actions as, and while the practice is genuinely useful, what it leaves behind still falls short of anything an investor, or someone running compliance review, could act on. Open that record six months later, and a new engineer finds only a log.

We call the result an evidence package. Assembled while the change ships, it records what changed, what was tested, what was scanned, who reviewed it, and how to undo it. That evidence package doubles as the audit trail a regulated customer eventually asks for, and a verified build arrives with one attached to its pull request (PR).

Calibrated autonomy is what we call the second decision here: how much of this stage runs unattended and how much waits for a signature, scaled to the blast radius of what is shipping.

Response time to a pull-request-ready fix fell from over 60 minutes to under 30 minutes on Incident Pilot, an AI incident-management system for a US fintech and SaaS platform. Incidents there had required 24/7 on-call engineering coverage across time zones. The agentic workflow we built triages and investigates on its own, then stops at a pull-request-ready fix for a human to take over. Autonomous resolution succeeds in up to 70% of cases, and round-the-clock staffing is no longer necessary.

Speed and evidence together get the full treatment in measuring AI development velocity without hiding rework.

Stage #5. Repeat: Production Feedback Into the Next Spec

Every stage before this one runs on assumptions written down before the code existed. Production checks those assumptions. This stage routes what production returns back into the next spec, so a failure that reached users becomes an acceptance criterion the following change has to satisfy.

Most mature workflows have a feedback loop somewhere. It tends to run on whoever was on call that week, and the return path rarely has an owner or a defined place to land. The loop closes each cycle differently, which is why the same class of failure keeps coming back.

Two artifacts absorb what comes back. Specs take the behavioral gaps, so an edge case missed on one pass becomes an acceptance criterion for the next. The harness takes the pattern-level ones, where a mistake the agent made twice becomes a rule in AGENTS.md, or a Skill that loads for that kind of task.

Limiting post-incident feedback to the responding engineer doesn’t scale very well across the organization. Cadence-based reviews added to recent evidence packages make informal observations a predictable phase in the life cycle.

The Full Workflow at a Glance

The Full Agentic Development Workflow at a Glance
The Full Agentic Development Workflow

The whole agentic development workflow fits on one page, which is the point of writing it down. Each row below names what a stage does, what it borrows, and what this version adds. Production feedback closes the loop at the bottom, then re-enters at the top.

Stage
What Happens
What's Borrowed
What This Workflow Adds

#1. Baseline & Frame

Idea becomes a written spec — scope, acceptance criteria, reversibility of key decisions

Spec-driven planning (both solo and enterprise sources)

Ceremony scaled to how expensive a wrong decision would be, not applied uniformly

#2. Design & Build

The agent executes inside a shared, versioned harness (rules + Skills)

AGENTS.md and Skills as context-management primitives (solo-developer source)

The harness is a team asset, reviewed and governed, not one developer’s local setup

#3. Assure

The result is verified independently, not by its own author

Human checkpoints at key steps (enterprise platform source)

The reviewer works independently from the original spec, not from a summary of the change

#4. Ship

The change ships with a package of evidence, not just a merged PR

Attributable action logging (enterprise platform source)

Evidence is assembled as a normal part of shipping, usable by a non-engineer reviewer

#5. Repeat

Production feedback and new failures feed back into the next Baseline & Frame stage

The general feedback-loop principle (present in some form in most mature workflows)

The loop is explicit and closes the same way every cycle, not ad hoc

What Happens

Idea becomes a written spec — scope, acceptance criteria, reversibility of key decisions

The agent executes inside a shared, versioned harness (rules + Skills)

The result is verified independently, not by its own author

The change ships with a package of evidence, not just a merged PR

Production feedback and new failures feed back into the next Baseline & Frame stage

What's Borrowed

Spec-driven planning (both solo and enterprise sources)

AGENTS.md and Skills as context-management primitives (solo-developer source)

Human checkpoints at key steps (enterprise platform source)

Attributable action logging (enterprise platform source)

The general feedback-loop principle (present in some form in most mature workflows)

What This Workflow Adds

Ceremony scaled to how expensive a wrong decision would be, not applied uniformly

The harness is a team asset, reviewed and governed, not one developer’s local setup

The reviewer works independently from the original spec, not from a summary of the change

Evidence is assembled as a normal part of shipping, usable by a non-engineer reviewer

The loop is explicit and closes the same way every cycle, not ad hoc

When You Need More (or Less) Than This

The wrong size hurts in both directions. A one-line copy change needs none of this, and pushing it through five stages produces the ceremony mismatch that spec-first practice exists to prevent.

There is tangible value in proportionality. InfoQ, reporting on DORA’s ROI of AI-Assisted Software Development report, notes Stanford’s Software Engineering Productivity research finding gains of 35 to 40% on simple greenfield tasks against often 10% or less on complex legacy code, which explains why the same process weight lands so differently on a new feature and a ten-year-old billing module.

Teams genuinely reaching platform scale should build the infrastructure they keep working around, since dozens of engineers and many concurrent workflows outrun a documented process. For a team already building informally and ready for governed structure, moving from vibe coding to production-grade delivery is the same transition on an existing codebase. Five stages fit a specific and common stage of company growth, and they make a poor universal law.

Practical Workflow Readiness Checklist

An agentic development workflow is easier to assess than to design, and six practices separate one that holds up from a set of good individual habits. Read the checklist as a diagnostic on the team you have today.

✓
Practice
Stage

⃣

Features of real size get a written spec before implementation, not just a prompt

Baseline & Frame

⃣

AGENTS.md and Skills are versioned in the repository, not living in one developer’s local setup

Design & Build

⃣

Someone other than a change’s author (human or AI) verifies it before it ships

Assure

⃣

Tests confirm specific behaviors, separate from broader review of quality and approach

Assure

⃣

Releases carry a record of what changed, what was tested, and who reviewed it

Ship

⃣

Production feedback routinely becomes input to the next round of specs, not just informal awareness

Repeat

✓

⃣

⃣

⃣

⃣

⃣

⃣

Stage

Baseline & Frame

Design & Build

Assure

Assure

Ship

Repeat

A team checking the first two rows and a few below them plans and executes well without verifying much, and that gap is usually the fastest and cheapest one to close. The more comprehensive AI production-ready checklist covers the model side.

Our Expertise

Most of what this article describes comes out of one engagement. A long-standing SPD Technology’s client runs agent-assisted delivery on its own engineering platform at production scale, with our engineers working inside that platform’s repositories. The five stages describe how it operates day-to-day.

Pipelines in these environments run through mandatory review gates before any pull request gets created, and sensitive code gets a stricter look than the rest. Automated deployment actions trigger only after explicit confirmation gates, while production incident tools operate strictly in read-only mode.

  • Specs and decision records are versioned in the client’s own repository from day one.
  • The shared harness holds the rules and the Skills, with tool access reviewed as part of it.
  • Independent review is structural. The process that produced a change never signs it off.
  • Evidence emerges from the delivery process as a byproduct.

Key Takeaways

  • The best current guidance sits at two extremes: a personal setup for one developer and shared infrastructure for a platform team, leaving the product team between them improvising.
  • Separating fixed structure from an agent’s variable reasoning needs a documented, versioned process and no platform team. Spec-driven development is where both extremes already agree.
  • AGENTS.md for always-loaded context and Agent Skills for on-demand context is a genuinely good split that has to become a shared repository asset once a second engineer depends on it.
  • Credible practitioners disagree on whether MCP servers are worth the context-bloat risk, and governance of that access decides the answer.
  • A developer reviewing their own AI’s output is careful but not independent, since the judgment that directed the work is the only check on it. Independent verification has to start elsewhere.
  • Recording every agent action as an attributable change is useful, and still falls short of an evidence package a non-engineer can read.
  • A small change does not need the full weight of five stages, and a team reaching real platform scale should invest in that infrastructure.

In short: five stages wrap unpredictable reasoning in structure a product team can write down this quarter, the middle ground the solo write-ups and platform pitches step over.

FAQ

  • What is an agentic development workflow?

    The agentic development workflow decomposes software delivery into separate phases. It helps the AI models continuously observe the process and infer what to do next. It starts with written specifications that drive execution in a governed harness. Then the work goes through an independent validation review before releases gather evidence packages, and the loop closes with production feedback. This framework is commonly referred to as an agentic software development lifecycle in industry teams.