Skip to main content
AI Engineering Hiring

OpenAI's agentic CI/CD: what it means for DevOps hiring

Calin Muresan•
#agentic CI/CD#DevOps hiring#SRE#platform engineering#AI agents#Cloud & DevOps

OpenAI’s agentic CI/CD: what it means for DevOps hiring

Some of OpenAI’s build-test-deploy systems saw roughly 10x more load in about six months, growth that might take most companies two or three years, as Codex agents took on implementing, testing, and opening pull requests. For anyone hiring senior DevOps, SRE, or platform engineers, the takeaway isn’t “AI makes engineers faster.” In our view, the scarce skill has moved from keeping pipelines running to designing the risk tiers, review gates, and observability that let agents use those pipelines without a human approving every change.

Senior DevOps screens used to test whether a candidate could keep CI green. At OpenAI, agents now watch CI until it’s green. What still needs a human is setting the rules for which changes an agent may approve.

On September 15, 2026, Gergely Orosz published Inside OpenAI’s agentic software factory in The Pragmatic Engineer. Venkat Venkataramani, OpenAI’s VP of Engineering for Applied Infra, told him the company is seeing “roughly a 10x increase in load on some systems” in about six months. At most companies, he said, that growth might take two or three years.

This analysis covers what OpenAI described, why it’s a platform engineering problem rather than a productivity story, which skills it makes scarce, and six screening questions to use on your next senior DevOps or SRE hire. We’re former software engineers who run Cloud & DevOps searches, so we read this the way your platform lead would. We worked from the free portion of the article; later sections are paywalled.

Key takeaways

  • OpenAI reports roughly 10x load growth on some build-test-deploy systems in about six months, while Codex agents implement, test, open PRs, and watch CI until it passes.
  • It handles the volume with risk tiers: specialist agent reviewers, stricter paths for high-risk changes, and opt-in auto-approval for low-risk PRs in some areas of the codebase.
  • Humans still own incidents. OpenAI’s Sevbot agent gathers context and proposes mitigations, but engineers stay on call.
  • Industry-wide, AI is lifting throughput while hurting stability. DORA’s 2025 report links higher AI adoption to more delivery throughput and lower delivery stability.
  • For senior DevOps and SRE hiring, screen for risk classification, per-change observability, and incident boundaries for agents. Pipeline maintenance alone is no longer the senior bar.

What happened inside OpenAI’s software factory?

The Pragmatic Engineer described a pipeline where OpenAI’s Codex agents handle much of the mechanical delivery loop (implementing, testing, opening PRs, watching CI), and humans set the rules those agents run under. The details below come from the article’s free sections, published September 15, 2026.

Key facts:

  • Agents run the build-test loop. Codex implements a change, builds and runs tests, fixes broken tests, opens a pull request, and monitors CI until it’s green.
  • Load is rising fast. PRs per engineer are growing “like a hockey stick,” and every part of the build-test-deploy pipeline is seeing more load. Venkataramani put it at roughly 10x on some systems in about six months.
  • Review is split by domain. Multiple agent reviewers, each configured as a specialist in an area such as cloud infrastructure or security, review changes.
  • Changes are classified by risk. High-risk changes can get more AI reviews or require a human review after the agents finish. Areas of the codebase can opt in to an agent that auto-approves low-risk PRs, which removes human acceptance as a bottleneck.
  • Deploys bring their own monitoring. When an agent deploys a change, it builds its own monitoring dashboard to watch it.
  • Incidents get an agent assistant, not an agent owner. Sevbot, OpenAI’s incident response agent built on Codex, collects context, works out possible mitigations, answers questions in Slack, and acts when an engineer tells it to. Engineers remain on call. The stated goal is for it to mitigate routine outages autonomously.

Codex use also spread beyond engineering: non-engineering teams at OpenAI went from roughly 0% in early 2025 to 90% by April 2026, according to the same article.

A risk-tiered agentic delivery pipeline, as described at OpenAI A Codex agent writes code, runs tests, and opens a pull request. Specialist agent reviewers check it. A risk classification step sends low-risk changes in opted-in areas to agent auto-approval, and high-risk changes to extra AI review or a human review. Both paths deploy, and the deploying agent builds its own monitoring dashboard. If an incident follows, the Sevbot agent proposes mitigations and an on-call engineer decides. Codex agent writes code, runs tests, opens PR Specialist agent reviewers cloud infra, security, other domains Risk classification Low risk, opted-in area agent auto-approves High risk more AI review or human review Deploy agent builds its own monitoring dashboard Incident: Sevbot proposes mitigations; on-call human decides
Simplified from the free sections of The Pragmatic Engineer's reporting (Sep 15, 2026). Low-risk changes outside opted-in areas follow standard review. OpenAI's internal thresholds and tooling details aren't public.

Why is 10x load in six months a platform problem, not a productivity story?

When agents multiply the number of changes, the constraint moves from writing code to verifying, merging, deploying, and watching it, and those are platform engineering jobs. Every extra PR is another CI run, another review, another deploy, and another thing that can page someone at 3 a.m.

OpenAI is the extreme case, but the direction is industry-wide. GitHub’s Octoverse 2025 report, covering September 2024 to August 2025, counted 43.2 million pull requests merged per month on average (+23% year over year), nearly 1 billion commits (+25.1%), and 11.5 billion GitHub Actions minutes in public projects (+35%).

Year-over-year growth in GitHub delivery activity, September 2024 to August 2025 Horizontal bar chart. Pull requests merged per month grew 23%. Commits pushed grew 25.1%. GitHub Actions minutes in public projects grew 35%. GitHub delivery activity up 23–35% in a year Pull requests merged per month +23% Commits pushed +25.1% GitHub Actions minutes (public projects) +35% 0% 10% 20% 30% 40%
Source: GitHub Octoverse 2025 (published Oct 28, 2025), year-over-year change. For scale, roughly 10x the load on some of OpenAI's systems means about +900%, in about six months.

In our view, most platform teams plan capacity around growth closer to GitHub’s 23–35% a year than to OpenAI’s 10x in six months. The metrics differ, but the gap in scale is the point. A team that scales CI capacity once a year, reviews every merge by hand, and builds dashboards per service rather than per change doesn’t bend under that. It breaks. That’s why we read this as a hiring signal, not an anecdote about fast engineers.

Throughput is outrunning stability

Google’s 2025 DORA report, based on nearly 5,000 technology professionals and published September 23, 2025, found that 90% of respondents use AI at work. It linked higher AI adoption to more software delivery throughput, but also to lower delivery stability. In DORA’s words, “that acceleration can expose weaknesses downstream.”

The same report found that 90% of organizations have adopted at least one internal platform, and that platform quality correlates directly with an organization’s ability to get value from AI. The platform team decides whether AI speed turns into shipped value or into incidents.

Confidence is running ahead of control

Leaders report high confidence in AI-written code and, in the same survey, more production issues from it. That gap is where senior DevOps and SRE judgment earns its pay.

CloudBees’ 2026 State of Code Abundance report, a vendor survey of 200+ enterprise technology leaders published May 19, 2026, found that 92% are confident AI-generated code is production-ready. In the same survey, 81% reported more production issues tied to AI-generated code. Developers are more skeptical: in the 2025 Stack Overflow Developer Survey, 45.7% said they distrust the accuracy of AI tools and 32.7% said they trust it.

Confidence in AI code versus reported problems and distrust, across three surveys Lollipop chart. 92% of leaders are confident AI code is production-ready (CloudBees). 81% of leaders report more production issues from AI code (CloudBees). 45.7% of developers distrust AI tool accuracy (Stack Overflow). 32.7% of developers trust AI tool accuracy (Stack Overflow). 30% of DORA respondents report little or no trust in AI-generated code. Confident, but reporting more production issues Leaders confident AI code is production-ready (CloudBees) 92% Leaders seeing more production issues from AI code (CloudBees) 81% Developers who distrust AI tool accuracy (Stack Overflow) 45.7% Developers who trust AI tool accuracy (Stack Overflow) 32.7% Little or no trust in AI-generated code (DORA) 30% 0% 25% 50% 75% 100%
Three surveys, three populations; compare within a survey, not across. Sources: CloudBees 2026 State of Code Abundance (200+ enterprise leaders, May 2026); Stack Overflow Developer Survey 2025; DORA 2025 (nearly 5,000 respondents).

OpenAI’s answer to that gap isn’t more trust or less. It’s routing: decide which changes an agent may approve, which need more review, and which need a person. Someone has to design those rules, measure whether they hold, and change them when they don’t. That’s the senior DevOps and SRE job now.

Which DevOps and SRE skills just became scarce?

The scarce skills are risk classification, per-change observability, and setting safe boundaries for agents in delivery and incidents. Pipeline maintenance still matters, but at OpenAI, agents already do much of it. The table below is our reading of what OpenAI’s setup implies for the senior bar.

Area Old senior bar New senior bar
CI/CD Keeps pipelines green and fast Designs pipelines that scale with agent PR volume, including merge queues, test selection, and CI cost per change
Code review Enforces human review on every merge Defines risk tiers: what can auto-approve, what gets extra AI review, what must reach a human
Observability Dashboards and alerts per service Signals per change, so a bad deploy is caught and tied to its PR within minutes
Incidents Runs on-call and writes postmortems Sets what an incident agent may propose, run, and never touch, with an audit trail
Metrics Deploy frequency and uptime Throughput and stability together, plus the escape rate from each risk tier

Risk tiering is a design skill, not a policy document

OpenAI lets areas of the codebase opt in to auto-approval for low-risk PRs. That one sentence hides the hard work: defining “low risk” in terms a machine can check, such as changed paths, blast radius, test coverage, config versus code, and reversibility. It also means proving the classifier is right, by tracking how often auto-approved changes cause rollbacks. A senior candidate should be able to design that and say how they’d know it was failing.

Observability has to follow the change

When agents ship many small changes, service-level dashboards stop answering the only question that matters in an incident: which change did this? OpenAI’s deploying agent builds its own monitoring dashboard to watch the rollout. The human skill is deciding what those dashboards must contain and wiring deploy events, feature flags, and SLOs together so a regression points back to one PR.

Agents in incidents need hard boundaries

Sevbot proposes mitigations and acts when an engineer tells it to. The engineer stays on call. Moving Sevbot from proposing mitigations to handling routine outages on its own, which OpenAI describes as the goal, is a reliability engineering problem: which actions are reversible, which runbooks are safe to automate, and what the kill switch is. That’s classic SRE thinking, applied to a new kind of operator. If you’re weighing DevOps vs SRE profiles for this work, this is the part that leans SRE.

We argue in AI-native engineering that the scarce skill in AI-native teams is verification: proving AI output is good enough before trusting it. Risk-tiered delivery is that same judgment, built into the pipeline.

Six screening questions for senior DevOps and SRE candidates

Ask candidates to design the controls, not to recite the tools. These are our questions, not OpenAI’s. They’re built to test the reasoning behind the controls, so they don’t require production experience with coding agents.

# Question Strong answer includes Red flag
1 “Your PR volume grows 10x in six months. What breaks first in CI, and what do you change?” Queueing and flaky tests as first failures; test selection, merge queues, caching, CI cost per change “Add more runners” as the whole answer
2 “Define a low-risk change that an agent may auto-approve.” Checkable criteria (paths, size, config vs code, reversibility), opt-in scope, a way to revoke A list of file types with no way to measure errors
3 “How would you know your risk classifier is wrong?” Rollback and incident rates per tier, sampling auto-approved PRs for human audit “We’d hear about it”
4 “An agent deploys 40 small changes today. One causes a latency regression. How do you find it?” Deploy markers, per-change signals, progressive rollout, automatic rollback on SLO breach Bisecting by hand from service dashboards
5 “What should an incident agent never do without a human?” Irreversible or data-affecting actions, anything outside a tested runbook, clear audit logging “Nothing, if it’s accurate enough,” or “It shouldn’t do anything”
6 “Throughput went up and change failure rate went up. What do you report to leadership?” Both numbers together, stability as the constraint, a plan tied to specific gates Reporting deploy frequency alone

Reality check: “agentic CI/CD” is a new label and few CVs will show it. Screen for the underlying experience instead: merge queues at scale, progressive delivery, SLOs and error budgets, and change-level observability.

What this doesn’t mean

It doesn’t mean every team needs agent auto-approval now, or a new job title. Three limits apply before you rewrite a job spec.

  1. OpenAI is an outlier. It builds Codex and has an unusual incentive to push agent autonomy early. The 10x figure applies to “some systems,” not the whole pipeline, and the article’s later sections are paywalled.
  2. Humans still hold the pager. Sevbot doesn’t mitigate on its own yet, and high-risk changes can still require human review. The design goal is fewer human bottlenecks, not zero humans.
  3. A new title isn’t required. “Agent ops engineer” isn’t a role you need to open. These skills belong in your existing senior DevOps, SRE, and platform specs.

For most teams, the practical step is smaller: add risk tiering, change-level observability, and agent boundaries to the senior interview loop before agent volume forces the issue.

Hiring for this in Romania and CEE

Look for engineers who have operated high-volume delivery, not only engineers who list AI tools. The profiles that transfer best have run merge queues or monorepo CI at scale, built progressive delivery with automated rollback, or owned SLOs and on-call for a busy service. Agent experience is a bonus you can teach; judgment about blast radius takes years. We cover this in how we run AI and DevOps searches in Romania.

Senior DevOps and SRE pay in Romania varies widely by stack and seniority, so check current ranges in our IT salaries in Romania 2026 guide before you set a budget. If you’re building the team remotely, our playbook for hiring remote developers in Romania covers contracts, timelines, and process. And if agents are also driving your AI bill up, the cost side of this is in our guide to model routing and inference cost.

Frequently asked questions

What is agentic CI/CD?

Agentic CI/CD is a delivery pipeline where AI agents do much of the work humans used to do by hand: writing changes, running and fixing tests, opening pull requests, reviewing code, and monitoring deploys. Humans design the rules, such as which changes need human review, and handle what the agents can’t.

How fast is CI/CD load growing?

It varies widely. OpenAI reported roughly 10x load on some systems in about six months. Across GitHub, Octoverse 2025 counted 23% more merged pull requests per month and 35% more Actions minutes in public projects year over year. Octoverse doesn’t isolate how much of that growth comes from AI.

Does AI code review replace human review?

Not at OpenAI. Specialist agent reviewers check changes, low-risk PRs in opted-in areas can be auto-approved, and high-risk changes can require a human review after the agents finish.

What should I ask a senior DevOps engineer about AI agents?

Ask how they’d define a low-risk change that an agent may approve, how they’d detect a wrong risk classification, and what an incident agent should never do without a human. The six questions in the table above cover CI capacity, observability, and reporting.

Is DevOps or SRE a better fit for agent-driven pipelines?

Both, for different parts. DevOps and platform engineers usually own pipeline capacity, merge flow, and deploy tooling. SREs usually own SLOs, observability, and incident boundaries, which is where agent autonomy carries the most risk.

The bottom line

OpenAI’s software factory shows where delivery is heading: agents produce changes faster than people can approve them, so the pipeline has to decide which changes need a person. That’s a design problem, and it lands on your senior DevOps, SRE, and platform engineers.

Three things to take from this:

  1. Hire for risk design, not pipeline upkeep. Agents can keep CI green. Someone still has to define what’s safe for them to auto-approve.
  2. Screen for change-level observability. When volume rises 10x, “which change broke it?” has to take minutes.
  3. Keep humans on the pager, with clear agent boundaries. Even OpenAI does.

If you’re hiring a senior DevOps, SRE, or platform engineer to own this, we can help you write the spec and test the judgment behind it. We’re former engineers, so we screen with questions like the ones above. You get a curated shortlist of three to five candidates, not a CV flood. Tell us what you’re hiring for.


Last updated: September 16, 2026. Sources verified September 16, 2026: The Pragmatic Engineer, “Inside OpenAI’s agentic software factory” (Gergely Orosz, Sep 15, 2026; free sections only, later sections paywalled); GitHub, Octoverse 2025 (Oct 28, 2025; period Sep 2024–Aug 2025); Google Cloud, 2025 DORA report (Sep 23, 2025); CloudBees, 2026 State of Code Abundance report (May 19, 2026; vendor survey of 200+ enterprise technology leaders); Stack Overflow Developer Survey 2025, AI section. The skills table and screening questions are Wise Step’s analysis, not OpenAI’s.