AI Technical Screening: 48% Flagged, 61% Still Passed
Across 19,368 AI-led interviews run between July 2025 and January 2026, 48% of technical assessments were flagged for AI-assisted cheating, and 61.1% of the flagged candidates still scored above the passing bar. AI technical screening still handles volume perfectly well. It no longer produces a score you can make a hire or no-hire decision on.
Automated technical screening used to be the filter. On the current numbers, it is the leak.
That reversal is the whole story, and it took two independent datasets published four weeks apart to make it undeniable. Neither is comfortable reading if you have quietly outsourced your first technical round to a score. We’re former software engineers who run Cloud/DevOps and Data/AI searches, and “can we just automate the tech screen?” is a question we get every quarter. The honest answer changed in 2026, and this piece shows exactly where.
What follows: what the numbers say, who counted them and why that matters, why coding tests break before every other format, and a redesign you can run in the next 90 days.
Key Takeaways
- Across 19,368 interviews (July 2025 to January 2026), 48% of technical roles were flagged for AI-assisted cheating versus 12% in sales (Fabric, January 2026)
- 61.1% of flagged candidates still cleared a 7.0 passing score, so the flag caught them and the process advanced them anyway
- A separate dataset from CodeSignal puts proctored-assessment fraud at 35% in 2025, up from 16% in 2024, with entry-level rising from 15% to 40%
- Both datasets come from companies selling detection, and neither publishes stated limitations, so treat the magnitude as real and the precision as marketing
- Detection is the weaker fix. Redesigning the assessment so AI use is assumed, and judgement is what gets scored, is the one that survives the next model release
What 19,368 interviews actually showed
Fabric, which runs an AI interview platform, analysed 19,368 interviews conducted on it between July 2025 and January 2026. Overall, 38.5% of candidates were flagged for cheating behaviour. The split by role is where it gets useful.
| Segment | Flag rate |
|---|---|
| Technical roles | 48% |
| Sales roles | 12% |
| Junior candidates (0 to 5 years) | roughly 2x the senior rate |
| All candidates | 38.5% |
Technical roles were flagged at four times the rate of sales. The trajectory is just as sharp: flag rates roughly tripled between July and September 2025, then stayed elevated through January.
The number that should change your process is none of those. It is 61.1%, the share of flagged candidates who still scored at or above 7.0, the threshold Fabric treats as passing. The detection worked. The funnel advanced them regardless.
Run the arithmetic on a technical pipeline of 100 assessed candidates. Roughly 48 get flagged, and around 29 of those clear the bar anyway. That means close to three in every ten people reaching your next round carry an unresolved integrity flag that nobody acted on.
If you are hiring ML, data, or DevOps engineers, that maths is the reason a shortlist needs a human name attached to it. Our AI and data science recruitment work exists on exactly that boundary.
A second dataset says the same thing from a different angle
Four weeks later, CodeSignal reported that cheating and fraud attempt rates on proctored assessments more than doubled in a year: 16% in 2024, 35% in 2025. Entry-level assessments went from 15% to 40%, nearly tripling.
Different company, different product, different method. Similar magnitude. That convergence is stronger evidence than either number on its own.
CodeSignal also published the behavioural breakdown of what triggered flags, which is the most practically useful detail either dataset gives you:
| Signal in flagged assessments | Share |
|---|---|
| Frequent off-screen referencing during the session | 35% |
| Unusually linear typing, complex solutions with minimal pauses or debugging | 23% |
| Elevated similarity to known answers or leaked content | 15% |
One more finding deserves its own line, because it settles the take-home debate: score increases on unproctored assessments were more than four times larger than on proctored ones. Unsupervised formats are not slightly softer. They are a different measurement entirely.
Read the numbers like an engineer: who counted, and how
Here is the part nobody else publishing this data will tell you. Both datasets come from vendors selling the remedy. Fabric sells an AI interview platform with detection built in. CodeSignal sells proctoring. Neither report states its limitations.
That doesn’t make the numbers wrong. It does mean you should discount the precision and keep the direction.
Three specific caveats we would apply before quoting any of this in a board deck:
- A flag is a signal, not a verdict. Fabric flags at a cheat probability above 40%, using gaze tracking, response timing, and keystroke dynamics across 20-plus signals. That is a threshold someone chose, not a confession.
- Neither dataset breaks out Europe. CodeSignal’s regional split is Asia-Pacific 48% versus North America 27%. Europe is absent from both. Importing either figure as a CEE benchmark would be inventing a number.
- Some widely-quoted figures are anecdotes. The “80% of candidates use LLMs on code tests” line doing the rounds traces back to Karat quoting one unnamed tech leader who suspected it about one region’s candidates. We won’t cite it, and neither should you.
Karat’s own figure from the same piece is the one worth keeping: only 30% of organisations rank updating technical assessments for AI as a top priority. Detection budgets are growing faster than assessment redesign. That is backwards.
Reality check: if a vendor tells you their tool catches 48% of cheaters, ask what happens to the 61% who get caught and pass anyway. That is a process question, not a product one.
Why the coding assessment breaks before every other format
The 48% versus 12% gap is not a story about engineers having worse ethics than salespeople. It is a story about what an LLM is good at.
Coding questions have clean, checkable answers. That is the exact shape of problem a language model solves best, in seconds, from a screenshot. A sales role-play has no single correct output to generate, which is why the same tooling barely moves the number there.
Three structural factors compound it:
- The format is asynchronous and unsupervised. Score inflation on unproctored assessments runs more than 4x proctored, per CodeSignal.
- Juniors are flagged at roughly double the senior rate, and juniors are precisely where volume screening gets pointed.
- The volume pressure is real. CodeSignal puts applications per role up 182% since 2021, with an average posting drawing 340 applicants (CodeSignal, vendor marketing figures, not a published study, so treat them as directional). Nobody automated the tech screen for fun. They automated it because the queue became unmanageable.
That last point is why “just stop using automated screening” isn’t advice. The scale problem didn’t go away. Only the assumption that a score equals a judgement did.
Want a straight read on how many qualified engineers actually exist for your role in this market, before you design a funnel around filtering them? That is what our market intelligence work covers.
The technical screening redesign, assuming AI is in the room
Stop trying to prove the candidate did not use AI. Assume they did, and score the thing AI cannot fake: judgement about code they have to own in front of you.
| Format | AI-resistance | What it actually measures |
|---|---|---|
| Unproctored take-home, classic | Very low | Whether they can prompt |
| Proctored coding test | Medium | Recall under supervision |
| Live problem extension (working code given, they change it) | High | Reading, tradeoffs, debugging |
| Defend-your-own-code round | High | Whether the work is genuinely theirs |
| AI-permitted task, judgement scored | High | How they direct and verify a model |
| Production-context probing on past work | Very high | Real experience, unfakeable |
Four changes worth making in the next 90 days:
- Replace “write this function” with “here is working code, change this behaviour.” Extension tasks require reading and tradeoff reasoning, which is where AI-assisted candidates stall.
- Add a defend round. Fifteen minutes on a submitted solution: why this data structure, what breaks at 10x load, what you would delete. Off-screen referencing shows up fast under follow-up questions.
- Permit AI explicitly, then score the verification. Ask what they checked and what they rejected. This is the same signal we screen for in AI-native engineering hires, and it separates orchestration from blind trust.
- Keep automation where it belongs. Sourcing, scheduling, first-pass triage, deduplication. AI handles the scale. Humans handle the judgement. Both matter, neither replaces the other.
What this changes if you’re hiring in Romania or CEE
Do not import the headline percentages into a Romanian funnel. Neither dataset reports Europe, let alone CEE, and a borrowed number is still an invented one.
What does transfer is the structural read. Romania’s market is mid-to-senior heavy in the roles we work, and the senior end is where flag rates are lowest and production-context probing works best. The Eastern European hiring market in 2026 is a speed play. A screening step you can’t trust is a speed problem before it’s an integrity problem, because every bad advance costs a hiring-manager hour downstream.
First-hand: the searches where automated screening still holds up in our pipeline are the ones with a narrow, verifiable stack requirement, where a wrong advance gets caught in the next conversation. The searches where it does not are senior hires, where the whole question is judgement, and where the CV and the assessment both look fine right up until someone asks how the system behaved in production. That is the round we do not automate, and it is why we send 3 to 5 vetted candidates rather than a filtered list.
Before you buy an AI screening tool: the compliance check
Every article on this topic recommends detection software. Almost none mention that deploying it in the EU has rules attached.
Recruitment and candidate-evaluation systems sit in the high-risk category under the EU AI Act, and gaze tracking plus keystroke dynamics is a more invasive class of processing than a scored coding test. Two dates matter right now:
- 2 August 2026, transparency obligations under the AI Act became enforceable.
- 2 December 2027, under the Digital Omnibus, in force since 27 July 2026, the high-risk obligations for employment AI systems were deferred to this date.
The deferral buys documentation time. It doesn’t make candidate-facing surveillance a decision you take without counsel. The European Commission’s guidance on high-risk AI systems is the primary source, and we walk through the hiring-specific detail in our piece on what the AI Act means for hiring.
Frequently asked questions
Can you trust AI technical screening for engineering hires?
For triage, yes. For a hire or no-hire decision, no. Across 19,368 interviews, 48% of technical assessments were flagged for AI use and 61.1% of those candidates still cleared the passing score, so the score alone no longer separates candidates reliably.
What percentage of candidates cheat in technical interviews?
38.5% of candidates were flagged overall in Fabric’s dataset, rising to 48% for technical roles and falling to 12% for sales. CodeSignal separately reported a 35% fraud attempt rate on proctored assessments in 2025, up from 16% in 2024.
Can interviewers tell if a candidate is using AI?
Sometimes, through behavioural signals rather than certainty. The most common triggers in flagged assessments were frequent off-screen referencing (35%), unusually linear typing with minimal debugging (23%), and similarity to leaked content (15%). A flag is a probability, not proof, which is why a follow-up round matters more than a detection dashboard.
Do take-home coding tests still work in 2026?
Not in their classic unproctored form. Score increases on unproctored assessments ran more than four times larger than on proctored ones. A take-home becomes useful again only when paired with a live round where the candidate defends and extends their own submission.
Is AI proctoring legal in the EU?
It is regulated, not banned. Recruitment systems are high-risk under the EU AI Act, transparency obligations have applied since 2 August 2026, and the high-risk obligations for employment systems were deferred to 2 December 2027 by the Digital Omnibus. Take legal advice before deploying gaze or keystroke analysis on candidates.
The bottom line
AI technical screening did not stop working. It stopped measuring what you think it measures. The scale it delivers is real, and the score it returns is now a triage signal rather than a verdict, on evidence from two vendors with every commercial reason to prefer a tidier story.
Three things to take away:
- Assume roughly half of technical assessments in your funnel involve AI assistance, and that most flagged candidates advance anyway unless someone intervenes.
- Redesign before you buy: extension tasks, a defend round, and AI-permitted work where verification gets scored beat any detection subscription.
- Keep automation on scale, keep humans on judgement, and check the EU AI Act position before pointing a proctoring tool at candidates.
The engineers worth hiring will pass an interview that assumes AI exists. The process that cannot tell the difference is the thing that needs fixing, not the candidates.
Hiring Cloud/DevOps or Data/AI engineers? If you want a shortlist vetted by people who’d run the technical round themselves, brief your search. We’ll give you a straight read on what it takes to fill it.
Data sources: Fabric (State of AI Interview Cheating in 2026, 19,368 interviews, January 2026); CodeSignal (assessment fraud press release, February 2026); Karat (Why Technical Interviews Must Evolve for the AI Era); European Commission (guidelines for providers and deployers of high-risk AI systems). Last updated: August 21, 2026.