Skip to main contentGet the benchmark report: Where does your team stand on AI adoption?
CONTACT SALESSTART BUILDING

Peer Benchmark Report:

How Engineering Leaders Are Restructuring Around AI

What engineering and product leaders told us about AI adoption, where it stalls, and what actually changes delivery.

46%
Engineers using AI tools
at >75% team usage
21%
At agentic maturity
generating prod code with oversight
43%
Velocity improvement
report moderate 20–50% gain
32%
Review process unchanged
in 18 months despite rising volume
72%
Lose 2+ hrs/week to handoffs
coordination tax intact
61%
Made no org changes
org chart has not moved

Executive Summary

AI adoption in software development is near total. The more useful question is what happens after a team adopts these tools, when the easy wins are spent and the work that used to define a developer's day starts moving somewhere else.

We surveyed a group of engineering and product leaders to find out. They span software engineering, executive leadership, design, and engineering management, sitting at very different points on the AI journey. That spread is what makes the group useful as a benchmark; it lets you find the rung you are standing on and see what the teams a rung or two above you did differently.

Six findings run through the data, and together they tell a more complicated story than the adoption numbers suggest.

01
Adoption is nearly universal. Maturity is not.

46% of respondents report that more than three-quarters of their engineers use AI tools. At the same time, only about a fifth describe their organization as running agentic workflows where AI generates and ships production code under human oversight. The gap between "we use it" and "it changed how we ship" is the gap this report is about.

02
Coding got faster. The bottleneck moved downstream.

The largest share of respondents reports moderate improvement of 20% to 50%, a smaller group reports only slight improvement of 1% to 20%, and a handful report no measurable change at all. The teams seeing significant gains of 50% or more are a minority, and they share a specific characteristic we will get to.

03
Code review is the new constraint — and it's straining.

Coding got faster, so the constraint relocated to the steps coding feeds. Leaders now name code review, architectural oversight, and product strategy as the things slowing them down. One executive put it plainly: the bottleneck is now in design and strategic PM thinking, not coding.

04
Non-engineers have the keys. The workflow hasn't caught up.

Most respondents grant designers access to AI tools; a similar share grants it to product managers, and many extend it to QA. Access is widespread, and the structured workflow that would let those roles ship safely is mostly absent, which is why "working across tools" and "lack of governance" repeatedly show up as failure points.

05
Handoffs still cost hours every week.

Even with faster individual steps, about 72% of respondents report losing two or more hours per week to handoffs between tools and roles, and roughly a fifth lose more than five. The coordination tax survived the productivity revolution intact.

06
Hiring is mostly steady, with a quiet shift at the edges.

Most teams are not changing headcount or structure yet. The minority making changes is adjusting who they hire, not how many, shifting toward AI-native engineers, AI-focused roles, and judgment-first candidates.

Underneath all six sits one pattern. Every respondent has access to capable models, which makes the tools a constant across the sample, while the workflow is the real variable. The teams reporting significant velocity gains and the cleanest review processes are the ones that redesigned the process around the tools. The teams that stuck at moderate gains added powerful tools to the workflows they already had. That distinction is the throughline of this report, and it is the single most useful thing the data has to say.

Here is the same picture at a glance

What the data showsThe numberWhat it means
AI tool usage among engineers46% at >75% usageAdoption is effectively universal
Self-placed at agentic maturity~21%Few have rebuilt the workflow
Velocity impact43% moderate, 25% slight, 11% noneReal gains, capped by unchanged process
Review process unchanged in 18 months~32%Review reshaped more than coding
Weekly time lost to handoffs72% lose 2+ hoursCoordination tax intact
Made no hiring or structure change~61%The org chart has not moved
Finding 01

Adoption is nearly universal. Maturity is not.

The bulk of the field sits in assisted and integrated workflows, where AI helps individuals but the overall process stays roughly as it was. Less than a quarter have rebuilt the workflow around what the tools make possible.

Engineers actively using AI

More than 75% of engineers46%
50 to 75% of engineers18%
25 to 50% of engineers32%
Less than 10% of engineers4%

Self-placed maturity stage

Agentic21%
Integrated29%
Assisted29%
Experimentation18%
All stages4%

Time from trial to broader rollout

Less than 3 months32%
3 to 6 months21%
6 to 12 months25%
Still in pilot21%

The practical read: Adoption percentage tells you almost nothing about whether AI is helping your organization ship. The maturity stage tells you most of it. If your engineers all use AI and your delivery metrics haven't moved, you're in the company of roughly a third of this sample — getting individual productivity without organizational throughput.

Finding 02

Coding got faster. The bottleneck moved downstream.

For most teams, writing code is 25–75% of the job. A tool that speeds up coding by a real margin produces only a moderate change in overall velocity — the math caps the upside, and survey results land right where the math predicts.

AI impact on development velocity

Significant (50%+)21%
Moderate (20–50%)43%
Slight (1–20%)25%
No measurable change11%
43%
report moderate gains — the most common outcome

Share of week spent coding

64%spend ≥50%
More than 75% coding 32%
50–75% coding 32%
25–50% coding 29%
Less than 25% coding 7%

"Fundamentally shifted our focus from the manual mechanics of writing syntax to high-level architectural design."

Executive leadership

"The bottleneck is now design and strategic PM thinking, not coding."

Executive leadership, agentic workflows
Finding 03

Code review is the new constraint, and it's straining.

Only 32% say their review process looks the same as 18 months ago. For more than two-thirds, review is one of the things AI has most visibly reshaped, often ahead of coding itself.

How review changed vs. 18 months ago

Faster via AI pre-checks29%
Essentially the same32%
More architecture-focused25%
Higher volume reviewed11%
Other / unclear3%
68%
changed their review process in 18 months as code volume increased

Roles spending more time on strategic work

Senior & staff engineers46%
Engineering managers & architects32%
Product managers14%
Designers11%
None4%
46%
of senior & staff engineers now spend more time on strategic work
Approval gates
Right people sign off before a PR is even opened
Automated QA
Change is exercised in a real browser before a human sees it
Severity triage
High-severity issues block merge; lower ones get flagged
Finding 04

Non-engineers have the keys. The workflow hasn't caught up.

Most respondents grant designers access to AI tools, and a similar share extends it to PMs. Access is widespread. The structured workflow that would let those roles ship safely is mostly absent.

Roles beyond engineering with AI access

Designers & UX64%
Product managers57%
QA & test engineers39%
Anyone / unsure11%
Marketers4%

Biggest barrier to broader adoption

Code quality concerns29%
Security concerns21%
Integration with existing systems21%
Lack of governance / review workflows14%
Other (training, compliance, cost)14%
What works

Non-engineers given a shared environment, working in the real codebase and design system, with engineering review as the gate rather than the bottleneck.

What fails

Broad access without structure: designers, PMs, and engineers generating output in isolated silos that have to be reconciled through the same handoffs AI was supposed to remove.

Finding 05

Handoffs still cost hours every week.

72% of teams lose at least two hours every week to handoffs between tools and roles. AI sped up the time inside each step. It did nothing to the arrows between them.

72%
lose 2+ hours per week to handoffs
Even among teams where nearly every engineer uses AI and coding has gotten meaningfully faster.

Weekly time lost to handoffs

2–5 hours lost/week52%
Less than 2 hours28%
5–10 hours lost/week16%
More than 10 hours4%

Work moved off engineering's plate

UI scaffolding & boilerplate43%
Tests & documentation39%
Design-to-code implementation36%
None yet25%

The mechanism: When AI generates code that doesn't use your real components, follow your conventions, or fit your architecture, every piece becomes a rework cycle. The faster the generation, the more rework cycles per week, which is how a tool that speeds up coding can leave delivery flat.

The production-readiness gap

We asked how often AI-generated code reaches production without major rework. The answers were sobering. Only a minority said most of the time. A large share said about half the time, occasionally, or rarely. Several teams reporting the slowest velocity gains also reported that AI code rarely reaches production without significant rework. One respondent was blunt: "poor code generation that needs a rewrite in the worst case and fixing blatant errors in the best case."

How often AI output works with the team's real codebase and design system maps almost perfectly onto maturity and velocity. Teams answering most of the time are concentrated among agentic and integrated respondents reporting significant gains. Teams answering sometimes or rarely are at experimentation, reporting slight or no gains. The correlation is not subtle: whether AI output respects your real system is close to the whole ballgame for whether it produces value or rework.

The implication ties Findings 3, 4 & 5 together

Review strains because generated code that ignores the design system arrives as raw work to verify. Cross-functional access fails because a designer's AI output that doesn't use real components has to be rebuilt by an engineer. Handoffs cost hours because the production-readiness gap turns every handoff into a correction cycle. One root cause, the disconnect between what AI generates and what the team's real system requires, propagates into three of the report's headline problems.

Finding 06

Hiring is mostly steady, with a quiet shift at the edges.

The minority making changes is adjusting who they hire, not how many, shifting toward AI-native engineers and judgment-first candidates.

Changes to team structure or hiring

No changes61%
Slowed hiring for some roles21%
Shifted toward AI-native hires7%
Introduced AI-focused roles4%
Added AI fluency criterion4%

What this signals

The job changed
Coding is no longer the scarce skill. Judgment is, the ability to lead architecture, review generated code, and direct agents.
Composition, not headcount
Slowing some roles while adding AI fluency as a screen reshapes the team without shrinking it.
The org chart follows the workflow
Teams that changed how they build are adjusting who they hire. Teams that haven't changed the workflow have no reason to change hiring yet.

The Throughline

Workflow change, not tool adoption

Every respondent has access to the same capable models. The variable is workflow. Teams reporting significant gains, clean review, and low handoff loss are consistently the teams that redesigned the process around the tools, not the ones that put powerful tools into an unchanged process.

18% of sample
Experimentation

Individual developers try AI tools. Velocity gains are slight or unmeasured. Nothing has changed structurally.

What to build nextCommit to a chosen workflow, experimentation that doesn't resolve is where "it's a mess" starts.
29% of sample
Assisted

AI is a regular part of engineering. Individuals move faster while the team holds steady.

What to build nextExtend beyond the individual. Ensure AI output uses your real system so it stops generating rework.
29% of sample
Integrated

AI appears across product, design, and engineering. Cross-functional friction begins here.

What to build nextBuild governance and a shared workspace, the control layer that turns broad access into throughput.
21% of sample
Agentic

AI generates and ships production code under structured oversight. Senior engineers are on architecture.

What to build nextScale it: push more routine review to agents so senior judgment isn't the bottleneck.

What to do this quarter

Stop measuring adoption. Start measuring workflow.

01
Measure your production-readiness rate
What fraction of AI-generated changes merge without major rework? This single metric tracked velocity and maturity more tightly than any other in the sample.
02
Audit handoffs explicitly
Pick two recent features and trace every point where work passed between people or tools. The total tends to surprise, and it points straight at the handoffs worth collapsing.
03
Check who owns review and strategy
When senior engineers carry both rising review volume and architectural strategy, they are the constraint. Automated first-pass review frees senior judgment for where it counts.
04
Verify access produces throughput
Does a designer or PM with AI access move work into your codebase through a governed path, or does their output still land on an engineer's desk to be rebuilt?
05
Start where AI output is most reliable
Begin with UI scaffolding and design-to-code translation. These are checkable by non-engineers, and errors are recoverable. Sequence matters.

What this means for the year ahead

Here is where this sample lands, hold your organization against it.

The single recommendation that follows from all of it: stop measuring adoption and start measuring workflow. Count the hours lost to handoffs. Watch where review time actually goes. Notice whether broad tool access produced shared throughput or parallel chaos. Track whether your velocity gain on coding reaches your delivery metrics or evaporates in the steps around it.

If you are like this sampleThen
Your engineers almost certainly use AI
That is the baseline now, shared with everyone, rather than an edge
Your velocity gain is probably moderate on the coding step
It shrinks further on overall delivery once the unchanged surrounding steps are counted
Your bottleneck has likely moved from coding to review and strategy
Judgment is a scarce resource, and more generated code aimed at the same senior reviewers will not relieve it
Your review process is probably straining or has already been redesigned
A review process unchanged in 18 months, while volume rose, is a warning worth investigating
Your non-engineers probably have access without team-level throughput
A shared, governed workflow for those roles is the differentiator, and most of this sample lacks it
Your team is probably losing two to five hours a week to handoffs
That coordination tax is where coding-speed gains leak out before they reach delivery
Your headcount has probably held steady
The hiring profile may tilt toward judgment and AI fluency, and the org-chart revolution has not arrived in this data

Where a platform closes the gap

Build as a whole team, not a team of ones.

The gap is drawn from what these leaders reported. Here is what the data says closes it, broken down by where the friction actually shows up.

AI output that uses your real system

The production-readiness gap that drives the handoff tax comes from AI that does not know your system. Builder generates against your real codebase, design tokens, and components through design system intelligence, so output is built to merge through your normal review rather than being rewritten.

A shared environment for every role

The cross-functional chaos that follows broad access stems from the lack of a shared environment. Builder gives PMs, designers, and engineers a single place to work on the same branch, with engineering review as the gate rather than the bottleneck.

Review structured before a human engages

The review strain comes from volume hitting an unchanged process. Builder's quality review agent runs through the change in a real browser before a human engages, and approval steps mean the right people vet a change before it reaches a PR.

Governance built into the workflow

The governance barrier that shows up in the data as the thing blocking broader adoption is built into the workflow rather than bolted on after, so broad access is safe, not chaotic.

Build the workflow layer

Turn AI adoption into mergeable work.

Builder gives teams one governed place where AI builds with real components, brand tokens, and design systems, then ships through the review workflow engineers already trust.