Skip to content
Ishaan Reddy

Blog · architecture

Why Agentic Coding Needs CI/CD More Than Human Coding Does

2026-08-127 minagentic codingCI/CDtestingautomationDevOps

Agentic coding makes a weak CI pipeline expensive very quickly.

A developer can ask an agent to change an API, update every caller, rewrite the tests, and prepare the commit in one run. Producing the change may take minutes. Verifying that the change preserves every assumption in the repository still takes as long as it did before.

That mismatch is why CI/CD matters more once agents start writing code. Human-written code needs the same checks, but a human naturally limits the rate of change. We pause to read files, wait for commands, notice an odd abstraction, and carry the reason for a decision from one edit to the next. An agent can cross those boundaries much faster and produce code that looks coherent even when its first assumption was wrong.

Faster generation is useful only if failure can travel back through the system just as fast.

The agent needs an external judge

An agent can test its own work. It can run the compiler, inspect a failure, make another edit, and try again. That loop is valuable, but it still depends on which commands the agent chose to run and what the local machine happened to contain.

CI moves the acceptance criteria out of the agent's prompt and into the repository. Every change gets the same commands, secrets policy, operating-system images, and clean starting environment. “Works on my machine” stops being an outcome because the machine that decides whether the change passes is disposable and repeatable.

The pipeline also prevents a bad assumption from surviving several rounds of agent work. If an API change quietly breaks a second package, the integration test fails before the agent builds more code on top of it. The logs become the next input to the coding loop.

This diagram shows the decisions, not the exact order of jobs. A real pipeline will run many of these checks in parallel. A green result proves only that the change satisfied the checks the team wrote down, which lets the reviewer spend more time on behavior, architecture, and missing requirements.

Give each CI stage one question to answer

A pipeline that runs one test command gives an agent a narrow definition of “done.” Each stage should reject a different kind of mistake.

  1. Can the repository build? Compilation and type checks catch missing imports, invalid signatures, broken generated code, and references to files that exist locally but were never committed. Run these first because later jobs cannot tell you much about code that does not build.

  2. Does the behavior still work? Unit tests check small pieces of logic. Integration tests check the assumptions between those pieces: the API response still matches the client, the migration still matches the model, and the command-line interface still passes the right values into the library. Agents are good at updating code consistently within the files they found. Integration tests catch the boundary they did not find.

  3. Does the change follow repository rules? Linters and formatters remove mechanical review work. They also catch unused values, unsafe patterns, and configuration mistakes when the chosen tools support those rules. A reviewer should spend attention on the change, not on whitespace that a command can settle.

  4. Did the dependency surface change safely? Dependency and security scans can flag known vulnerabilities, unexpected lockfile changes, and packages that violate a project policy. They make dependency changes visible before deployment while deeper trust questions remain part of review.

  5. Does it work outside the agent's environment? A platform matrix runs the same build and tests on every operating system or runtime the project claims to support. This is where CI catches code that is valid on the agent's Linux runner and invalid on Windows.

These checks also improve the agent's next attempt. “CI failed” is vague. “The Windows build cannot resolve this Unix-only import” gives the agent a location, a constraint, and a clear success condition.

The Rust failure you want before release day

Consider a network-monitoring application written in Rust and shipped on Windows, macOS, and Linux. A change to interface discovery compiles on Linux, passes the unit tests, and looks reasonable in review. The implementation calls an API that does not exist on Windows.

Without a build matrix, someone has to remember to test Windows manually. That often happens near a release, when the original change is no longer fresh and unrelated commits have landed around it. The debugging job now includes finding which change introduced the failure.

A CI matrix runs the Windows, macOS, and Linux builds on the pull request. The Windows job fails on the commit that introduced the unsupported call. The agent can read that log, add the correct platform-specific implementation or conditional compilation, and rerun the same matrix.

The pipeline only has to enforce the project's existing promise that all three targets build.

Agentic coding increases the value of this check because the agent can make cross-cutting changes faster than a developer can manually open three operating systems. The matrix scales verification with generation instead of asking a reviewer to absorb the difference.

CD comes after evidence

Continuous deployment gets safer when it is the last step of the same pipeline. The commit that passed the build, tests, scans, and review becomes the commit that the deployment job releases. Nobody rebuilds it differently on a laptop or copies files by hand over SSH.

That repeatability matters when a release goes wrong. The team can redeploy a known-good artifact or revert to a previous commit through the same path. Rollback still needs engineering, especially when a release changes stored data, but the application deployment itself should not depend on remembering a sequence of terminal commands during an incident.

Giving an agent direct production access does not remove this order. The agent proposes a change, CI gathers evidence, a human or policy approves it, and the deployment job uses a restricted production credential. The pipeline defines what the agent may ship and records what happened.

A minimum pipeline for agent-written code

Start with checks that can reject a change without human judgment:

  1. Build or typecheck the entire repository from a clean checkout.
  2. Run unit and integration tests.
  3. Run linting and formatting checks without silently rewriting the CI workspace.
  4. Scan dependency and lockfile changes.
  5. Build and test every supported platform or runtime version.
  6. Create a preview or staging deployment when the application supports one.
  7. Require review before the production deployment job runs.

The exact tools matter less than whether the commands are deterministic, required, and visible to the agent. A warning that everyone ignores is not a gate. A test that runs only on a developer's laptop cannot protect the shared branch.

CI/CD exposes bad code while the change is still small, the reasoning is still available, and the agent can respond to the failure. Before an agent is allowed to produce more code, the repository needs a reliable way to reject what it already produced.