• AI coding agents
  • Agentic coding
  • Developer workflow

AI Coding Agents: A Workflow That Holds Up

The short version

Treat a coding agent like a fast colleague who needs a clear ticket and a way to check its own work. Write the task with acceptance criteria, ask for a plan before any code, give the agent commands that prove the change works, keep every diff small, and review it as the owner. Speed is the agent’s job. Judgment stays with you.

Where the time goes now

Coding agents write and change code quickly, and they follow instructions well. That shifts where the effort goes in a project. Less of it is typing. More of it lands in two places: saying exactly what you want, and checking that you got it.

Both are old engineering skills. Writing a clear ticket and reviewing a change were always part of the job, and they matter more now because the cost of producing a change has dropped while the cost of understanding it has not. The workflow below does not depend on any vendor. It works whether your agent lives in a terminal, an editor or a pull request bot.

Write the task like a ticket

An agent fills every gap you leave with its own best guess. A vague request returns a confident answer to a question you did not quite ask. A useful task has four parts: the goal, the constraints, how to tell it is done, and what is out of scope.

Prompt
Goal
Add a "resend verification email" button to the account page.

Constraints
Reuse the existing sendVerificationEmail service. Do not add dependencies.
Limit resends to 3 per hour per user.

Done when
* A signed in user with an unverified email sees the button.
* The fourth request within an hour returns a 429 and shows a clear message.
* Tests cover the limit and the already verified case.

Out of scope
Email template changes, the signup flow, anything under /billing.

The last section matters more than it looks. Without it, an agent will helpfully tidy things you never mentioned, and you end up reviewing a larger change than the one you wanted.

Ask for a plan before any code

For anything beyond a small edit, ask the agent to read the code and propose a plan, and do not let it change files until you have read that plan. Most agents have a planning or read only mode for exactly this. A plan is cheap to read and cheap to correct. A wrong plan caught here saves a wrong implementation and the review that would have followed.

When I read a plan I check three things. Does it touch only the files I expect? Does it reuse code that already exists instead of writing a parallel copy? Does it say how the change will be tested? If a plan proposes a new dependency or a new pattern, I ask why before it writes a line.

Make correctness something the agent can run

The biggest single improvement to agent output is giving it a way to check itself. If tests, type checks and linting run with one command, the agent can change code, run the command, read the failures and fix them without you in the loop.

JSON
{
  "scripts": {
    "check": "eslint . && tsc --noEmit && vitest run"
  }
}

Tell the agent in plain words to run that command before it says the task is finished. Better still, start from a failing test that describes the behaviour you want and let the agent make it pass. A test you wrote is a specification. A test the agent wrote for its own code can quietly confirm its own assumptions, so read those more closely than the code itself.

Keep every change small

Review quality falls quickly as a diff grows. A change you can read in a few minutes gets a real review. A change spread over thirty files gets a skim. Ask for one concern at a time, commit at checkpoints, and treat a revert as an ordinary tool instead of a failure.

You can enforce this instead of hoping for it. This script fails the build when a change is bigger than you agreed to review. Run it in CI against the branch you are merging.

JavaScript
import { execFileSync } from 'node:child_process'

const base = process.env.BASE_REF ?? 'origin/main'
const maxFiles = 15
const maxLines = 400

const stat = execFileSync('git', ['diff', '--numstat', base + '...HEAD'], { encoding: 'utf8' })
const rows = stat.trim().split('\n').filter(Boolean).map((line) => line.split('\t'))
const files = rows.length
const lines = rows.reduce((sum, [added, removed]) => sum + (Number(added) || 0) + (Number(removed) || 0), 0)

if (files > maxFiles || lines > maxLines) {
  console.error('Diff budget exceeded: ' + files + ' files, ' + lines + ' changed lines. Split this change.')
  process.exit(1)
}
console.log('Diff within budget: ' + files + ' files, ' + lines + ' lines')

Fifteen files and four hundred lines are a starting point, not a rule. Pick limits your team can genuinely review, and exclude lockfiles and generated files once they start causing false alarms.

Teach the repository with an instructions file

An agent begins each session knowing nothing about your project. Most tools read a plain text instructions file from the root of the repository. AGENTS.md has become a common shared name that several tools understand, and some tools also read a file of their own, such as CLAUDE.md. If your team uses more than one tool, keeping the shared rules in one place avoids them drifting apart.

What goes in it is short and practical: the commands to build and test, the conventions that are not obvious from reading the code, the folders that are off limits, and the mistakes that have happened before. Start small. Add a line when you see the same mistake twice, and where you can, add a test that fails if it returns, so the rule is enforced and not merely written down.

Prompt
# AGENTS.md

## Commands
* Run everything with: npm run check
* Never run the seed script against a shared database.

## Conventions
* Money is stored as integer minor units. Never use floats for amounts.
* New endpoints validate input with the shared schemas in src/schemas.

## Off limits
* Do not edit src/generated or any migration that has already run.

Review it as the owner

When a pull request arrives, the reflex is to open the changed files. Start one step earlier. Read the task, decide what a correct change would look like, and then compare. Reviewing against intent catches the most expensive problem of all, a change that is neat and well tested and solves the wrong thing.

  • Scope. Did it do what the task asked and nothing more?
  • Edge cases. Empty input, a repeated click, a missing record, a slow network.
  • Security. Who is allowed to call this? Is every input validated? Are secrets kept out of the code and the logs?
  • Dependencies. Did it add a package, and is that justified?
  • Performance. Queries inside loops, lists with no limit, whole files loaded into memory.
  • Tests. Would they fail if the feature were broken?

This kind of review rests on fundamentals. If you are earlier in your career, the guide on vibe coding versus learning to code covers the basics that make it possible.

Know when to take the keyboard back

Some situations are better handled by hand, or at least with you driving closely.

  • The same error has come back three times in a row.
  • The requirement is unclear and the agent is guessing between interpretations.
  • The code is security sensitive: authentication, permissions or payments.
  • The action cannot be undone, such as a data migration or a delete.

Apply the same least privilege you would give a new teammate. Work on a branch, keep production credentials out of reach, and restrict destructive commands so a bad guess costs you a revert and not a recovery.

The loop in one place

  • Write the ticket with a goal, constraints, a definition of done and what is out of scope.
  • Ask for a plan and read it before any file changes.
  • Have the agent run the check command and fix what fails.
  • Keep the diff small enough to review properly.
  • Review against intent, then against the checklist.
  • Record any repeated mistake in the instructions file, ideally with a test.

Expect the first few tasks to feel slower while you write tickets and set up checks. The payoff is that the output starts arriving in a state you can trust and review quickly, and that is where the time saving from an agent actually comes from.

Keep reading