Should a recruiting team use coding agents? OpenAI went 0 to 90%
Should a recruiting team use coding agents? OpenAI went 0 to 90%
Yes, if the agent can finish whole tasks. OpenAI’s recruiting, finance, and legal teams went from about 0% to 90% Codex usage in four months, and nobody ordered them to. Coding agents used to be a tool for engineers. In 2026, at least one company’s recruiters use one as a matter of course, and the thing that pulled them in was not a policy. It was a feature that let the agent keep working on a task until it was done.
We run a recruiting firm on agents ourselves, so we read this with some skepticism and some recognition. This piece covers what OpenAI actually reported, why adoption moved when it did, what a coding agent does in recruiting work, where the real limits are, and a four-week test you can run before buying anything.
Key takeaways
- OpenAI’s non-engineering teams, including recruitment, went from roughly 0% to 90% Codex usage in about four months, without a mandate, per The Pragmatic Engineer (September 15, 2026). The same article reports that almost all OpenAI employees now use Codex and ChatGPT Work weekly.
- Adoption reached close to 40% while the app still showed code on screen. It jumped from 60% to 90% between April and May 2026, after OpenAI added a /goal setting that keeps the agent working until a task is complete.
- The pull was task completion, not chat. Non-engineering staff used it for research, documents, and spreadsheets, and some colleagues had it monitor Slack or update Airtable.
- Mandates are not useless. A July 2026 field study of an enterprise “2x” mandate found it acted as a catalyst, with the gain coming from actual use.
- In recruiting, the agent should take the scale work: sourcing passes, market scans, scheduling, first-pass screening drafts. Candidate judgment, the offer conversation, and anything the EU AI Act calls high-risk stay with a person.
What OpenAI reported about its non-engineering teams
Between roughly February and May 2026, OpenAI’s finance, recruitment, and legal teams moved from almost no Codex usage to 90%, and the company did not require it. The account comes from Gergely Orosz’s visit to OpenAI’s headquarters, published in The Pragmatic Engineer on September 15, 2026. We worked from the free sections; the later parts of the article are paywalled.
The specific claims, in the article’s own framing:
- “In a four-month period, non-engineering orgs like finance, recruitment, and legal went from ~0% usage of Codex to 90% usage.” As a separate point, the article says almost all OpenAI employees now use Codex and ChatGPT Work weekly.
- The move happened “without a mandate from above for it.”
- OpenAI shipped the Codex desktop app for Mac in February 2026 and for Windows in March. ChatGPT Work, built on the same harness, followed in July.
- Non-engineering adoption got “close to 40%” during February to April, when the app was, in the article’s words, “hostile to non-engineering users.” It still showed code on screen. People used it anyway because it could research and produce a presentation, a document, or a spreadsheet.
- “Between April and May, usage surged from 60% to 90%” after OpenAI added a /goal setting, “where you can set up a goal for the agent and it keeps working until it is complete.”
The percentages come from a chart of Codex token usage by department since August 2025; the article does not publish the underlying task-level data. One caveat the article makes itself: OpenAI’s internal Codex is wired into nearly every internal system, so it is more capable than the version anyone else can buy.
For the engineering side of the same visit, where OpenAI’s VP of Engineering for Applied Infra describes “roughly a 10x increase in load on some systems” in about six months, see our take on OpenAI’s agentic CI/CD and DevOps hiring.
Why adoption moved: utility, not policy
The jump came when the agent could run a long task to completion, and the article’s own numbers say so. Adoption climbed only to around 40% over two months with a product that was hard for non-engineers to use. It reached 90% within weeks of the /goal setting. OpenAI’s desktop lead, Andrew Ambrosino, told Orosz: “The number one thing that is changing is that people are starting to use threads for much longer, and this longer usage has been a breakthrough.” He added that people “often set a goal and then have the model crank.”
Two other mechanisms the article names are worth copying:
- Role-specific plugins. Each group packaged its own workflows as plugins and passed them around. Ambrosino’s reasoning: “You can’t just give everyone an empty box. Skills and plugins let teams adapt the agent to their work.”
- Word of mouth as the distribution channel. Akshay Nathan, Engineering Lead for Productivity, described an “awareness gap”: some people had worked out they could use Codex to monitor Slack, update Airtable, or create onboarding materials, while “many others still use it for one task and then discover more uses from teammates.”
None of that is a mandate. It is a product that finishes work, plus a way for a recruiter to hand a working recipe to the recruiter next to them.
The mandate playbook is not wrong, it is just slower
Top-down orders do get adoption, but the evidence says they work as a catalyst, and the gain still comes from use. Shopify’s CEO wrote in April 2025 that “reflexive AI usage is now a baseline expectation”, with AI use folded into performance reviews. A longitudinal study posted to arXiv on July 2, 2026 (a preprint) tracked 802 developers and 196,212 pull requests at a mid-sized, AI-forward company that mandated a doubling of merged pull requests per engineer from mid-2025. Throughput did double, reaching 2.09x the baseline by April 2026. The authors read the mandate as “a catalyst rather than a direct driver,” with the gain tied to adoption and to accumulated use.
Zapier’s route is the closest public comparison to OpenAI’s. Its CEO, Wade Foster, described the rollout as a company-wide hackathon, internal demos, and adoption tracked in engagement surveys: 63% of staff reported daily AI use in late 2023, 77% by the end of 2024, and hovering around 97% more recently. That took more than two years and a “Code Red.” OpenAI’s non-engineering teams covered similar ground in four months, largely because the product changed underneath them.
Our read: if your recruiting team is stuck at low usage, the fix is usually not another all-hands. It is a tool that completes a task they actually have, and one colleague who shows them.
What a coding agent does in recruiting work
A coding agent is not a chatbot with better manners. It is a worker that can read your systems, write and run code, loop until a result checks out, and hand back a finished artifact. For a recruiting team, “code” is mostly invisible. What you see is a market map, a cleaned candidate list, a shortlist brief, or a scheduling run.
The split we use, and the one OpenAI’s account implies, is the same one we argued for in how reliable AI resume screening actually is: AI handles the scale, people handle the judgment.
| Recruiting task | Agent does | Recruiter keeps |
|---|---|---|
| Market scan for a role | Pull public data, salary references, competitor postings, and build a first map | Decide which companies to target and what “good” looks like for this stack |
| Sourcing pass | Search, dedupe, enrich, and rank against a written brief | Choose who to contact and how to open |
| First-pass screening | Summarize CVs against the brief, flag gaps and inconsistencies for review | Read the flagged ones, decide who moves, own the decision |
| Interview logistics | Scheduling, reminders, notes cleanup, tracker updates | Nothing, once the recipe is trusted |
| Technical assessment | Draft questions from the stack, structure notes | Run the interview, judge depth, spot coached answers |
| Offer and counter-offer | Pull market bands, model the total cost | Have the conversation, read the candidate |
| Reporting | Weekly client update drafted from the tracker | Add the honest sentence the tracker does not show |
Two rules keep this workable, and they are not a compliance program. First, the agent drafts and a person signs; screening scores are advice, not decisions. Second, AI systems used to recruit or evaluate candidates are listed as high-risk in Annex III of the EU AI Act, Regulation (EU) 2024/1689, which brings documentation, oversight, and transparency duties. We cover the details in what the EU AI Act means for recruitment.
What we run on agents at Wise Step, and what we do not
We are a two-founder recruiting firm, so we cannot show you a 90% adoption chart; a sample of two is not a curve. What we can show is which tasks moved and which did not. We are former software engineers, which makes us an easy case, and we say so. The pattern still matched OpenAI’s.
In the search work, the agent takes sourcing, matching, and the first pass. A person takes the brief, the technical assessment, and the offer. Our average time to first candidates is 10 to 15 days, and average time to fill is 28 to 30 days. The agent side handles the volume behind the first number; the founders stay on the second.
What did not move, and will not: reading whether a senior engineer’s answer is deep or coached, telling a client the uncomfortable part of a market, and the counter-offer call. Those are the job.
The honest limit: long-running agents produce more output than a person can check. That is the problem He et al. document in the study above: AI writes faster than humans can review. In recruiting, the fix is a queue a person owns, not a faster agent.
A four-week test for a recruiting team
Do not roll a coding agent out to a recruiting team. Give three people one task each that they already hate, and measure whether it gets done. OpenAI’s curve says the pull comes from completion, so test completion. This test is tool-agnostic by design: Codex, Claude Code, and comparable agents all have a long-running mode now, and per-seat pricing changes too often to print here.
- Week 1: pick three tasks with rich output. A market scan for an open role, a weekly client report drafted from the tracker, and interview scheduling. Each should take a person over an hour today. Write a one-page brief for each in plain language: inputs, the output you want, and what “done” means.
- Week 2: run them with a goal, not a prompt. Use whatever long-running mode your tool has and let the agent iterate until the output matches the brief. Log the time a person spent reviewing versus the time the task used to take. Log every error.
- Week 3: package what worked. Turn the two best briefs into a reusable skill or plugin, the way OpenAI’s teams did. Hand it to a fourth person with no instructions beyond “run this.”
- Week 4: measure use, not sentiment. Count who ran an agent task each week, unprompted. Count review time per task and errors caught by the reviewer. Decide on the numbers.
The two numbers that matter at the end are unprompted weekly use and reviewer-caught errors per task. If use is high and errors are low, expand. If use is high and errors are high, you have a review problem, not an adoption problem. If use is low, the task briefs were wrong, or the tool cannot finish the task. Neither is fixed by a memo.
Where this argument is weakest
OpenAI is the least representative company in the world for this question, and its 90% should be read as a ceiling, not a benchmark. Three limits apply.
- The tool was built next door. OpenAI’s engineers, researchers, finance, and marketing staff work with an unlimited token budget, and the internal Codex is wired into nearly every company system; the article says both. A recruiting team at a 200-person company will hit permissions, data access, and cost limits that OpenAI staff never see.
- Most recruiting teams are far earlier. SHRM’s AI in HR 2026, from 1,908 HR professionals and published April 8, 2026, found 27% of organizations applying AI in talent acquisition, the highest of any HR function. Among engineers, by contrast, Stack Overflow’s 2025 survey of about 49,000 developers found 51% of professionals using AI tools daily. Recruiting is not starting from the engineering baseline.
- Intent is running ahead of practice. Korn Ferry’s 2026 Talent Acquisition Trends, from 1,674 talent leaders and published October 28, 2025, found 52% planning to add autonomous agents to their teams in 2026. Microsoft’s 2026 Work Trend Index found only 26% of AI users saying their leadership is clearly and consistently aligned on AI. Plans and alignment are not the same as a recruiter running an agent on a Tuesday.
There is also the part we could be wrong about. Weekly use is a low bar. An agent opened once a week to draft a report is “adoption” in OpenAI’s chart, and it is not the same as the recruiting workflow moving. The article does not publish task-level data, and neither, to be fair, do we.
Frequently asked questions
Should a recruiting team use coding agents?
Yes, for tasks with rich output that a person can check: market scans, sourcing passes, CV summaries against a written brief, scheduling, and report drafts. Keep candidate decisions, technical interviews, and offer conversations with a person. Test three tasks for four weeks before rolling anything out.
Do recruiters need to know how to code to use a coding agent?
No. OpenAI’s non-engineering teams reached close to 40% adoption while the Codex app still showed code on screen, because the agent produced documents, spreadsheets, and research, not because staff read the code. The skill that matters is writing a brief with a clear definition of done.
What made OpenAI’s non-engineering adoption jump from 60% to 90%?
According to The Pragmatic Engineer, the jump between April and May 2026 followed OpenAI adding a /goal setting to Codex, which keeps the agent working on a goal until it is complete. Role-specific plugins and word of mouth between teammates also spread use.
Is an AI mandate a bad idea for a recruiting team?
Not bad, just insufficient. A July 2026 field study of an enterprise “2x” mandate found it acted as a catalyst, with the productivity gain tied to actual adoption and accumulated use. Pair any expectation with a task the tool can finish and a colleague who can show it.
What does the EU AI Act say about AI agents in recruitment?
Annex III of Regulation (EU) 2024/1689 lists AI systems used for recruitment and candidate evaluation as high-risk, which brings requirements for human oversight, documentation, and transparency toward candidates. Agents that draft, schedule, and research are lower risk than agents that score people. Keep the scoring advisory and a person accountable.
The bottom line
Coding agents used to be for engineers. At OpenAI, recruiting, finance, and legal teams reached 90% Codex usage in four months, and the driver was a feature that finished tasks, not an order from above.
Three things to take from it:
- Test completion, not chat. Give three recruiters one long task each and measure whether it gets done.
- Split scale from judgment. The agent drafts and sources. A person decides, interviews, and signs.
- Measure unprompted use. Sentiment surveys flatter. Weekly runs and reviewer-caught errors do not.
If you are hiring engineers in Romania or the wider CEE region and want a recruiting partner that already runs this way, we are former software engineers who use agents for the scale and keep the judgment ourselves. You get a curated shortlist of three to five candidates, with the uncomfortable parts of the market said out loud. Tell us what you are hiring for.
Last updated: September 17, 2026. Sources verified September 17, 2026: The Pragmatic Engineer, “Inside OpenAI’s agentic software factory” (Gergely Orosz, Sep 15, 2026; free sections only); He et al., “AI Writes Faster Than Humans Can Review” (arXiv, Jul 2, 2026); Zapier, “How Zapier rolled out AI org-wide and drove 97% adoption” (Wade Foster, updated Jan 2026); Tobi Lütke, Shopify memo (Apr 2025); SHRM, “AI in HR 2026” (Apr 8, 2026; n=1,908); Stack Overflow 2025 Developer Survey, AI section (n≈49,000); Korn Ferry, Talent Acquisition Trends 2026 (Oct 28, 2025; n=1,674); Microsoft, 2026 Work Trend Index (May 2026). Regulation (EU) 2024/1689, Annex III (EU AI Act Explorer). Wise Step process figures are our own operating averages.