Skip to main content
Jason Sonderman, UXMC, CPACCJason Sonderman

From Delivery to Direction

What the efficiency revealed, and what it means when your documented process becomes a training set.

Rebrand + light mode

1 wk

Down from a full 2-week sprint
Sprint recovery

½–1 sprint

Per milestone cycle within a 6-week cadence
Story point shift

5–8 → 1–3

Fibonacci UI estimates after design system hardening
Rollover rate

Full cycle → ≤20%

Capped under Shape Up by a Product/Engineering scope-renegotiation trigger
  • UX Leadership
  • UX Team Scaling
  • Design Systems
  • AI-Accelerated Delivery
  • Token Architecture
  • Cross-functional Leadership
  • Enterprise SaaS

The terrain

Control Tower and Arcos Field ran on two stacks, served two kinds of users, and opened every six-week cycle the same way: four days of planning sessions where product, design, and a development pod walked every defined outcome together before a line of code was written. Three senior designers covered that work across three pods. I carried the connective tissue role: coherence between the two surfaces, continuity into the legacy tools underneath, and the thread that kept what got built on one side from quietly drifting away from the other. The team was small by design. Constraints tend to produce clarity, and clarity, as it turns out, became the whole story.

For most of my time there, those planning sessions were where the friction lived. This is the story of what changed when the friction went away, and what it revealed when it did. For the technical story of how the delivery infrastructure was built, see The Handoff Is the Bug. This study picks up where that one ends.


What the efficiency revealed

The pipeline worked. Half a sprint to a full sprint recovered per six-week cycle. UI stories that pointed at 5s and 8s came back as 1s and 2s. A rebrand and light mode implementation that would have taken two weeks took one. The work was understood before it began.

Assuming a 65-hour per-person sprint capacity, front-end delivery started reaching cycle goals in about 30 hours, helped along by a shift to single-engineer ownership of a single outcome instead of shared ownership across a pod. The other 35 hours didn’t sit idle. They went to platform support, backend stories, technical overflow carried over from earlier cycles, and tech debt that had been waiting for room. That capacity moved sideways into engineering health, not into new front-end scope, which mattered for what the efficiency gain was actually worth.

Design capacity moved somewhere different. Once a designer could take a concept to delivery-ready in roughly two and a half days, the time that freed up didn’t just go back into reviewing what got built. It went into the next iteration’s outcomes while they were still being shaped, not after the shaping was already done. That produced a real overlap: the designer who’d just delivered one outcome carried what they’d learned building it directly into shaping the next one, instead of receiving someone else’s already-shaped outcome cold. Delivery capacity turned into shaping capacity, not only review capacity.

Working inside it closely enough revealed something the original hypothesis hadn’t accounted for. The ambiguity that remained wasn’t in the handoff. It was earlier: in whether the purpose behind a layout decision had ever been made explicit, in whether the shape of the data matched what a user actually needed to see, in the distance between what the interface asked someone to do and what they were actually trying to accomplish. Those questions don’t come from the design file. They come from being close enough to real users, in research, in testing, in the accumulated evidence of watching people work, that the intent behind a decision is grounded in something observed rather than assumed. The smarter the agent got, the more plainly it showed us how much that grounding mattered. Speed was compressing the cost of skipping it, not eliminating it.

We came looking for a better bridge. We found that the other side needed work first.

Then the mobile pod showed us what documentation is actually for.

The mobile surface had always been the harder one to plan for. There is no clean way to prototype in React Native, and the pod had never been part of the original delivery model rollout. During a planning session, we learned they had connected the Zeplin MCP to pull design context and user flows directly into a Claude Code workflow, using it to visualize shaping outcomes mid-discussion before the formal handoff. They had also built a UX Designer agent, trained on the ways-of-working documentation our team had published. Not a tool handed to them. Not a process they’d been asked to follow. Our own documented values and process norms, used as a behavioral frame for a model working alongside their developers. They built it themselves because the documentation was precise enough to make it possible.

I hadn’t written that documentation for this. I’d asked for it because unclear process wastes senior people’s judgment on questions that shouldn’t still be open, the same instinct behind Clarity of Purpose, just aimed inward at how the team’s own knowledge gets recorded. It turned out to double as something else entirely: the discipline that makes documentation useful to a new hire and the discipline that makes it useful to a model training on it are the same discipline. Say the actual reason a thing is done, not just the rule. If the ways-of-working docs had been checklists instead of reasoning, the mobile pod would have had nothing to build on. Good documentation was never just about onboarding humans. It’s what makes your judgment portable to people, and processes, you never planned for.

The testing story tracks the same before-and-after split. Convoy tracking was the one initiative we tested end to end before the pipeline existed, and its original scope tells you why: create a convoy, define its roster, track its travel, confirm arrival. Once we cut that down to just creating and tracking a convoy, roster and arrival check-in fell out of scope, and testing went with them. Login, mobile and Backoffice both, and the first AI expense agent were never tested at all. Once the pipeline was running, that changed. We tested working coded prototypes with real users on Roster Requests, Roster Response, Timecards Mobile, Timecards Backoffice, and Demobilization. Timecards Mobile is the clearest case of testing changing the build: storm-response users told us they cared less about what kind of hours they’d logged, travel, idle, working, resting, than about role level and pay type, regular, overtime, doubletime. The feature’s granularity moved to match what they actually needed to see.


Clarity of Purpose

Oracle’s Redwood Design System named two things that have stayed with me. Clarity of Purpose: every design decision grounded in a specific outcome before execution begins. And the shape of the data: not what the system holds, but what form a user needs it to take to act with confidence. Both principles are older than any library. Both became more urgent when the time between a decision and its built expression collapsed to days.

AI is a domain I’ve worked in long enough and closely enough that other leaders inside the organization began bringing me into the room when strategic decisions about it were being made, not as a title but as a point of view that had been tested against real work. Those conversations kept returning to the same reframe: the efficiency story is the smaller story. What AI-assisted delivery actually creates, if the recovered time goes somewhere intentional, is the conditions for design to do the work it was always supposed to be doing. Not generating UI. Holding purpose accountable. Defining the right shape for the right data. Staying close enough to what users actually need that the thing being built is worth building.

AI doesn’t make UX faster. It makes UX more valuable, if the recovered capacity goes back into the questions that determine whether velocity creates anything worth having.


The rollover, and what replaced it

Before the pipeline, rollover was the number nobody wanted to say out loud. Work a cycle didn’t finish wasn’t absorbed into the next one, it consumed it: an entire six-week release cycle could go to nothing but clearing the previous cycle’s backlog, with no room for anything new. Shape Up turned that into a ceiling instead of a slide. Rollover was capped near 20%, and the cap was enforced by a real mechanism rather than a hope: if projected rollover looked likely to exceed it, Product and Engineering had to meet and agree on a new scope before the work continued.

The ceiling held, and holding it changed what Shape Up itself needed to be. The mass of rollover it was managing kept rescaling every Milestone and Shaping session around it, and by April 2026 the formal versions of those phases were largely retired in favor of a shorter, more responsive planning cadence. The bet shifted from what a pod could commit to across six weeks to what one engineer could own across four. Individual engineers, not product owners, became accountable for a single outcome.


Early, and honest about it

One sprint in. That’s also where this account stops: the four-week, dev-owner cycle had just started when my time at Arcos ended, so what follows is what I saw, not what came after. The volume of UX updates to screens had dropped. More tellingly, the nature of the work had shifted. Designers were no longer the primary producers of UI in Figma, then advocates for its faithful translation into code. They had become reviewers: informed, principled, checking that what had been built held to the UX decisions that were made upstream, that accessibility hadn’t been traded away in the handoff, that the experience arriving at user acceptance testing was one the design team could stand behind. The production work moved to the agents. The judgment work stayed with the people.

That split looked different depending on which designer owned the domain, because I assigned ownership by skillset rather than spreading everyone across everything. Ales Karoza owned field and mobile experiences, including login. Oswaldo Vasquez owned workforce management, command-center flows, and the heaviest data visualizations. Erin Gold owned user-flow strategy and back-office data manipulation. The coded prototypes followed the same lines: Oswaldo built the Roster feature’s prototypes once he had requirements from Erin and product management behind him, plus a couple of rounds of Figma ideation. Ales built most of his prototypes in Figma Make, but Roster’s mobile surface ran on React Native, which didn’t lend itself to in-tool coded prototyping the way the web surfaces did. That work happened in Figma Design against the Harmony Mobile theme instead, then exported to Zeplin.

Whether that held, and what it looked like at scale across all three pods, was still being learned when my time there ended. What was visible by then was a design function that had moved upstream without shrinking: AI handling the translation layer, developers owning architecture and integration, UX with its attention on the questions that don’t resolve themselves.

What the user is actually trying to accomplish. Whether the thing being built gets them there. Whether the purpose behind the decision was ever made clear enough to survive the distance between intent and code.

That last question turns out to be the one that matters most. It was there before the AI. It will be there after whatever comes next. The tools changed what it costs to ignore it.

Agent view

This page has a machine-readable twin, generated from the same source rather than written alongside it.

See how an agent reads this page