# The Notebook in the Truck

Subject: Arcos, Inc
Role: Head of UX, Arcos, Inc
Timeline: Aug 2025 – Jun 2026
Canonical URL: https://jason.sonderman.info/case-studies/arcos-timecards/

Time entry for utility field crews at Arcos was taking two to four minutes against a forty-five to sixty second target, and slow entry gets deferred until it’s remembered instead of recorded. I traced the cause to the data model, not the interface: legacy timecard never read the crew and equipment records customers had already paid for. I held a Milestone phase two weeks to validate a crew-level entry model at the Linemen’s Rodeo with roughly 400 field users, then confirmed it again with an instrumented test during a Duke Energy storm simulation before a line of production code shipped. Timecards Mobile and Timecards Management shipped in June 2026, in beta.

## Key outcomes
- The gap: 45–60 sec target vs. 2–4 min observed (Target set from industry norms and direct field observation before the work started; 2.5 minutes was the typical figure on a standard job)
- Caught before build: 2 corrections to the data model (Entry-per-person to entry-per-crew, and an optional work order to a required one, both surfaced at the Linemen’s Rodeo before a line of production code was written)
- Validated at scale: ~400 field users, twice, independently (The Linemen’s Rodeo in October, then an instrumented test during a Duke Energy storm simulation in December, at a different utility)
- Shipped: June 2026, in beta (As Timecards Mobile and Timecards Management; one of two at-risk enterprise renewals closed on the presented capability before anything shipped)
- Alpha result: 60 seconds per entry (Internal experts. Bulk crew editing meant one entry covered a shift’s worth of a crew, beating the roughly-triple target set for a crew lead, not just meeting it)

## Summary
As Head of UX at Arcos, I traced a mobile time-entry defect back to the data model rather than the screen: the legacy timecard never read the crew and equipment records customers had already bought through Crew Manager, so a foreman retyped his crew by hand every entry. I held Shaping two weeks to validate a crew-level entry model at the Linemen’s Rodeo with roughly 400 field users, confirmed it again with an instrumented test during a Duke Energy storm simulation, and shipped Timecards Mobile and Timecards Management in June 2026, in beta, with one of two at-risk enterprise renewals closing on the presented capability before anything shipped.

## Facts
- The defect, benchmarked: Target of 45-60 seconds and 3-5 taps, set from industry norms and direct field observation before the work started. Observed: 2-4 minutes, with 2.5 minutes the typical figure on a standard job. A crew lead entering on behalf of a crew of three was targeted at roughly triple the individual benchmark, since the work really is three times the size
- Aggregate cost (self-reported, not a measured average): A job ran 2-4 entries and a shift ran 8-15 jobs, putting a typical day at roughly 60-90 minutes of entry alone and a storm day, at the heavy end of both ranges, closer to 4 hours inside a shift already running 10-12 hours
- Why it was a payroll problem, not a UX complaint: Slow entry gets deferred, and deferred means remembered instead of recorded by the end of a long shift. The lineman who prompted this work kept a paper notebook in his truck to check the app’s entry against, because payroll had gotten it wrong before: hours missing, hours billed to the wrong rate
- What was at stake: Two enterprise renewals were in trouble, timekeeping among the named reasons. Customers had already bought Crew Manager, which models individuals, crews, equipment, vehicles, and lodging, but the legacy timecard read only the logged-in user’s own record, so a foreman retyped his crew by hand out of a system his company had already paid for. A prior time-entry attempt had shipped before Arcos had a UX function, and it missed
- How the problem got scoped (Milestone phase): Outcome: a crew lead can submit time for their crew for the day, from the truck. Signal: the 45-60 second, 3-5 tap benchmark. No-gos: a month view, and Crew Manager integration, deferred rather than dropped, on the belief that the crew record would need to be rebuilt inside the new platform and synced back later
- Domain ownership: Ales Karoza owned mobile time entry and Oswaldo Vasquez owned the back office, each the authority on their end of the same data model rather than routing disagreements through me. Harmony, the Arcos design system, made that autonomy safe by settling spacing, states, token semantics, and destructive-action patterns before either of them had to escalate a decision
- Promotion case (honest gap): That domain ownership was documented as the basis to move both Ales and Oswaldo to Lead Product Designer the following cycle. A layoff reached them first, so neither promotion happened. It’s a case I built, not an outcome I delivered
- The Linemen’s Rodeo (October): Product’s Gainsight and Pendo reports showed time entry breaking down worse than almost anywhere else in the app, which told me where to point scarce field time, not why. One AEP site visit in August surfaced a mismatch between the scoped entry-per-person model and what a foreman actually did with his crew. That single site wasn’t a sample, so it went to leadership as a concern, and the CPO used an already-scheduled Linemen’s Rodeo to validate it at scale: roughly 400 linemen and foremen, plus Arcos’s Storm Response Action Committee, for a total incremental cost of one plane ticket and about two weeks of schedule held on Shaping
- What the Rodeo corrected, before a line of code was written: Entry-per-person moved to entry-per-crew (individual entry was never wrong, just incomplete for the highest-stakes case, storm response). A work order attachment moved from optional to required, including in storm, on a since-corrected assumption that storm was the exception to nearly everything else about storm. Caught at the Rodeo, both were a prototype revision; caught in build, they would have been a schema change
- Why research ran three times: Problem validation at Milestone, before a solution existed to bias the answer; solution validation during Shaping at the Rodeo; and a third pass in December, opportunistic rather than planned, once a mid-build instrumented test became possible
- The December instrumented test: A browser analog of the in-development React Native app, built agentically against Harmony’s component library, instrumented with Pendo, and run unmoderated during a Duke Energy storm simulation surfaced through the Storm Response Action Committee. Measured against baseline on time to goal, click-backs, and error corrections. Two negative findings surfaced and were deferred to v2 rather than blocking v1: users wanted the work order auto-attached to the job, and one utility’s crews call the same equipment a cherry-picker where the product said bucket truck
- What shipped: Timecards Mobile and Timecards Management shipped June 2026, in beta, not a full commercial launch. One of two at-risk enterprise renewals closed on the presented capability before anything shipped; the second remained open, with Timecards one factor among several, not the deciding one. In alpha, with internal experts, a single entry came in at 60 seconds, and because crew edits are bulk, that same 60 seconds covers a full crew on a typical shift, beating the roughly-triple crew-lead target rather than just meeting it
- What I don’t have (honest gap): A field-validated number that the crew-level fix held. Product owned Pendo and the quantitative side of the beta, and that tracking never firmed up once real crews were working long stretches with no connection; I hadn’t defined what I needed to see in that data before calling the fix proven, and I hadn’t built a manual fallback for when the telemetry couldn’t reach me. Whether entry speed holds with a gloved hand and no connection is a separate, still-open question
- What I’d change: Agree on the after-number before build starts rather than after, and read what Product already tracks on a set cycle. Separately, bring an engineering lead into the end of Milestone: the Crew Manager integration was scoped around data we believed was unreachable, and a more granular Crew API was already in progress on the legacy side. Nobody in the Milestone room was positioned to know that, which is a structural gap, not an individual one

## Scope
I left Arcos in July 2026 and no longer have access to dashboards, the wiki, or the ticket system, so figures here are reconstructed from memory and the materials I kept, not pulled from a live report. The Rodeo participant count, the aggregate daily time-cost figures, and Nick Nelson’s estimate of development time saved are stated as such below and shouldn’t be read as measured figures with the same footing as the alpha benchmark or the shipped date. Pair with UX as Organizational Strategy for how this project fit into the broader case for UX at Arcos, and with The Handoff Is the Bug for the Harmony-based delivery model this project’s coded prototypes used.

**Team:** 2 Senior Designers (mobile and back office)

<script type="application/ld+json" set:html={JSON.stringify({
  "@context": "https://schema.org",
  ...frontmatter.structuredData
})} />

<section aria-labelledby="notebook-heading">
  <h2 id="notebook-heading">The notebook in the truck</h2>

  <figure style="margin: 2rem 0;">
_Figure: Illustration of a lineman in a hard hat and hi-vis shirt writing in a notebook resting on a rugged tablet, standing near storm debris and a downed line with a bucket truck in the background._
    <figcaption style="font-size: 0.85rem; color: var(--color-text-muted, #666); margin-top: 0.5rem;">Illustration.</figcaption>
  </figure>

  <p>I was at an American Electric Power site doing field research for a different product entirely, emergency preparedness. Timekeeping wasn’t on my list. But I had the access, so I asked about it, and watched a lineman do it.</p>

  <p>He kept a notebook in the truck to check it against.</p>

  <p>Not instead of the app. Ahead of it. In the truck, job by job, he’d jot the hours, the job code, the work type on paper. The real entry happened later, on a laptop back at the office. And he kept the paper anyway, because payroll had gotten it wrong before: hours missing, hours billed to the wrong rate.</p>

  <p>That single detail turned out to hold the whole project. Everything that follows traces back to it: what was actually broken, why fixing the screen wasn’t enough, and why a foreman still didn’t fully trust the number until he’d checked it himself.</p>

</section>

<section aria-labelledby="defect-heading">
  <h2 id="defect-heading">The defect</h2>

  <p>Here’s what one entry actually took. You find the job, by job number. Then you pick the work type, and that’s not one choice, it’s a tree: travel, assess, repair, then which kind of repair, or wait time. That tree exists because it sets the hourly rate. Get it wrong and it bills wrong. Then you find the work order number and attach it. None of it was filled in ahead of time. None of it was filtered down to the person already logged in.</p>

  <p>The target was forty-five to sixty seconds and three to five taps, set from industry norms and from what we’d already watched people do before the work started. What we observed was two to four minutes, with two and a half minutes the number that kept coming up on a standard job. For a crew lead entering on behalf of three people, we set the target at roughly triple, because the work really is three times the size.</p>

  {/* IMAGE:
    What: A drawn decision-chain diagram matching the interview deck’s version: find job by number, then choose work type (a branching tree: travel / assess / repair / repair subtype / wait), then find and attach work order number. Annotated with the two benchmark figures (target: 45-60s, 3-5 taps; observed: 2-4 min, 2.5 typical).
    Why: The chain is the whole argument for "this is a data-model problem," and it’s easier to see as a diagram than read as prose.
    Alt: A flowchart showing three sequential steps to log one time entry: find job by number, choose work type from a branching list, then find and attach a work order number. Two callout boxes below show a 45-60 second target versus a 2-4 minute observed time.
  */}

  <p>The system knew his role. It knew which crew he was on. It knew there was a storm. None of that was used. That isn’t a bad screen. That’s a system that didn’t use what it already knew.</p>

  <p>A job runs two to four entries. A shift runs eight to fifteen jobs. On a typical day that’s somewhere around sixty to ninety minutes spent just putting time in. On a storm day, at the heavy end of both ranges, it’s closer to four hours, inside a shift that’s already ten to twelve. Slow entry doesn’t just cost time. It gets put off, and by the end of a long shift, put off means remembered instead of recorded. That’s not a UX complaint. That’s a payroll problem waiting to happen, and it’s the reason the notebook existed in the first place.</p>

</section>

<section aria-labelledby="stakes-heading">
  <h2 id="stakes-heading">What was at stake</h2>

  <p>That’s the user side. Here’s the business side: two enterprise renewals were in trouble, and timekeeping was one of the named reasons, raised more than once by the customers involved. That came to me through a standing forum I ran with Sales Engineering, where Tino Mathew, our VP of Field Engineering, brought concerns straight from the field back to product.</p>

  <p>Most of those same customers had already bought Crew Manager, which models the entire operation: individuals, crews, equipment, vehicles, lodging. It’s the answer to who’s working today and what they’re working with. The legacy timecard never touched any of it. It read the logged-in user’s own record and nothing else. So a foreman sat in a truck and retyped his crew by hand, every entry, out of a system his own company had already paid us for.</p>

  <p>The other thing in the room was history. Arcos had built time entry once already, before there was a UX function. It shipped. It missed. So the question I pushed into intake wasn’t whether we could build timesheets. It was why the last one missed, and what had to be different this time. A feature that already exists and isn’t trusted costs more than one that doesn’t exist yet.</p>

</section>

<section aria-labelledby="outcome-heading">
  <h2 id="outcome-heading">Writing the outcome</h2>

  <p>Arcos ran an adapted Shape Up, with a phase ahead of Shaping we called Milestone. Milestone takes a raw business ask and breaks it into capabilities, not features. It doesn’t scope and it doesn’t set requirements. What came into intake read like a list of things the team would do: build a mobile time-entry screen, add a week view, ship Timecards Mobile. None of those lines said what would become true for anyone, and none of them could be wrong, only late.</p>

  <p>The way I write an outcome instead has three parts, and none of them is a feature. The outcome: a crew lead can submit time for their crew for the day, from the truck. The signal: the forty-five to sixty second, three to five tap benchmark, the actual number we were trying to move. The no-gos: what we ruled out on purpose. A month view was one. Crew Manager integration was the other, deferred rather than dropped, on the belief that we’d need to rebuild the crew record inside the new platform and sync it back later. That belief turned out to be wrong, and I’ll come back to why.</p>

  <blockquote>
    <p>”Jason and I partnered together from day one&hellip; Jason was easy to work with, not afraid to disagree but always in a constructive way.”</p>
    <cite>Nick Nelson, Sr Product Director, Arcos</cite>
  </blockquote>

</section>

<section aria-labelledby="ownership-heading">
  <h2 id="ownership-heading">Domain ownership, not assignments</h2>

  <p>Ales Karoza owned mobile time entry. Oswaldo Vasquez owned the back office. Neither of those was an assignment I handed out for this one project. Ales and Oswaldo owned opposite ends of the same data model, mobile capture on one side and back-office reconciliation on the other, and when those two ends disagreed, I didn’t want both of them waiting on me to arbitrate. A domain owner has to be the authority on their end, or the model doesn’t distribute, it just adds a layer of asking permission.</p>

  <p>That only works if there’s a system underneath it. Ales and Oswaldo could make experience calls without checking with me, not because I told them to be bold, but because Harmony, the Arcos design system, already answered most of what they’d otherwise have escalated: spacing, states, token semantics, what a destructive action looks like. A design system is what makes autonomy safe. Without one, distributed ownership is just distributed guessing. For the token architecture underneath that, see <a href="/case-studies/arcos-harmony-design-system/">Harmony: A Design System Built for Machines</a>.</p>

  <p>A domain is also what makes somebody promotable. I had it documented as the basis to move both Ales and Oswaldo to Lead Product Designer the next cycle. A layoff reached them first, so neither promotion happened, and I want to be straight that it’s a case I built, not an outcome I delivered.</p>

  <blockquote>
    <p>”You can’t build a promotion case out of a list of assignments. You can build one out of a surface somebody owns.”</p>
  </blockquote>

</section>

<section aria-labelledby="rodeo-heading">
  <h2 id="rodeo-heading">What the Rodeo confirmed</h2>

  <p>The Milestone scope had every crew member entering their own time, which is what Product’s Gainsight and Pendo reports pointed to: entry-per-person, sessions and counts all clean, and time entry breaking down worse than almost anywhere else in the app. That told me where to point scarce field time. It didn’t tell me why. Screen data can show you what happens inside the app. It can’t show you what happens in a truck, at a pole, or at the end of a twelve-hour shift when somebody’s rebuilding hours from memory.</p>

  <p>In August, on the same AEP visit that opened this project, I watched a foreman put in time for his whole crew, and saw the mismatch right away. One utility isn’t a sample, so it went to leadership as a concern, not a finding. Arcos’s Linemen’s Rodeo was six weeks out, and our CPO asked whether it could do double duty. Shaping held its start about two weeks to use it. Executives signed off on validating at scale before any development resource committed, and no build cycle started until we had the answer. Roughly four hundred linemen and foremen, plus Arcos’s Storm Response Action Committee, for a total incremental cost of one plane ticket.</p>

  <p>It held: team leads enter for the crew, and that’s always been true. Worth saying that individual entry was never wrong, only incomplete for the case that matters most. On a blue-sky day, a lineman working alone on a meter install or a damage assessment really is one person. Storm response is where the crew model is the one that counts, and it’s the one we’d scoped a default around by accident.</p>

  <p>The second finding was the work order. Every entry needed one attached, including in storm, on an assumption that storm was the exception to nearly everything else about storm. Neither of these was a usability finding. Both were in the data model. Caught on a convention floor, they’re a prototype revision. Caught in build, they’re a schema change.</p>

  {/* IMAGE:
    What: A simple before/after data-model diagram: "Entry-per-person" with individual records feeding in separately, versus "Entry-per-crew" with one crew record and per-member rows nested under it. A second pair for the work-order field moving from optional to required.
    Why: These two corrections are described as data-model changes, not screen changes, and a diagram makes that distinction visible rather than asserted.
    Alt: Two before/after diagrams. Left pair: "Entry-per-person," showing separate disconnected records per crew member, versus "Entry-per-crew," showing one crew record containing all members. Right pair: a work order field shown as optional, then as required, both labeled "including in storm."
  */}

  <p>What changed: the crew shows up already filled in, with add and remove for whoever’s out that day. Hours go in for everybody at once, or you switch over and enter them one at a time. The work order requirement got a notes field in the first release, which wasn’t the real answer, but was the honest one for a v1.</p>

  <blockquote>
    <p>”He suggested and then built a vibe coded prototype so that we could conduct user testing as quickly as possible&hellip; which ultimately saved us weeks of development time and put our product on a successful path.”</p>
    <cite>Nick Nelson, Sr Product Director, Arcos</cite>
  </blockquote>

</section>

<section aria-labelledby="measurement-heading">
  <h2 id="measurement-heading">Why twice, and how we measured</h2>

  <p>We ran research twice on purpose, and a third time because the chance showed up. Problem validation happened at Milestone, before a solution shape existed to bias the answer. Solution validation happened during Shaping, at the Rodeo.</p>

  <blockquote>
    <p>”Jason made a deliberate effort to bring frontline perspectives into the UX process, not just as a validation step at the end, but while we were still defining the problem and shaping the experience.”</p>
    <cite>Tino Mathew, VP of Field Engineering, Arcos</cite>
  </blockquote>

  <p>The third pass came in December, mid-build, and it was opportunistic rather than planned. My team built a browser analog of the React Native mobile app already in development, coded agentically against Harmony, reproducing the same components in React and MUI. Because it was built from the same tokens as the real app, the numbers it produced carried over. I put Pendo tags on it and ran an unmoderated test during a Duke Energy storm simulation, surfaced through Arcos’s Storm Response Action Committee, the only setting where you get storm conditions and a willing audience at the same time. We measured against baseline on time to goal, click-backs, and error corrections.</p>

  <p>The numbers told us the flow worked. They didn’t tell us why. One field user, using the app during the simulation, put it this way:</p>

  <blockquote>
    <p>”It’s like it knows who I am and where in the process I am at&hellip; it just drops me there, and this is what I need to do right now.”</p>
    <cite>A field user, Duke Energy storm simulation</cite>
  </blockquote>

  <p>Two negative findings came out of the same session, and reporting them is what makes the positive result credible. One user wanted the work order auto-attached to the job rather than found manually, which the notes-field stopgap from the Rodeo hadn’t actually solved. Another told us our equipment naming was wrong: we called it a bucket truck, his crew calls it a cherry-picker, and the mismatch would confuse people in the field. When the interface is a list you pick from, naming decides whether someone finds the right row on the first try. Neither finding changed v1. Both went into v2 planning.</p>

</section>

<section aria-labelledby="shipped-heading">
  <h2 id="shipped-heading">What shipped</h2>

  <figure style="margin: 2rem 0;">
_Figure: Mobile Timecards screen showing a storm crew’s timecard with two time segments, each covering all six crew members at once, with an option to add another time segment and a notes field before sending the timecard._
    <figcaption style="font-size: 0.85rem; color: var(--color-text-muted, #666); margin-top: 0.5rem;">Timecards Mobile: bulk entry for a full crew, submitted from the truck.</figcaption>
  </figure>

  <figure style="margin: 2rem 0;">
_Figure: Mobile Timecards confirmation screen showing a submitted timecard with total hours calculated across two time segments and crew sizes, plus a crew lead’s note explaining two crew members' partial hours._
    <figcaption style="font-size: 0.85rem; color: var(--color-text-muted, #666); margin-top: 0.5rem;">Submitted: hours calculated per person, with the notes field standing in for the work-order fix still ahead of it.</figcaption>
  </figure>

  <p>Timecards Mobile and Timecards Management shipped June 2026, in beta. One of the two at-risk renewals closed on the presented capability, before we’d shipped anything. That’s the outcome I’d point to first. The second was still open when I left; Timecards was a factor in it, but I’d put it third, behind API extensions and platform speed. I don’t want to overclaim that one.</p>

  <p>A crew lead could now submit time for their crew for the day, from the truck, which was the outcome we wrote at Milestone, in the same words. In alpha, with internal experts, a single entry came in at sixty seconds, the top of the original target range. Because crew edits were bulk, that same sixty seconds covered a whole crew on a typical shift, where everyone worked the same job, the same work type, the same hours. The target for a crew of three was roughly triple the individual benchmark. Bulk editing meant one entry beat that target rather than just meeting it.</p>

  <blockquote>
    <p>”It’s bearing fruit in our recent production rollout and the positive feedback coming back from customers.”</p>
    <cite>Rusty Davis, Engineering Manager, Arcos</cite>
  </blockquote>

  <p>Whether the lineman who prompted this whole project still keeps his notebook, I don’t know. Faster entry is why he might not need it in the field anymore. Whether payroll stopped getting it wrong often enough that he’d trust the app without paper is a different question, and I don’t have that number either.</p>

</section>

<section aria-labelledby="different-heading">
  <h2 id="different-heading">What I’d do differently</h2>

  <p>I never got a hard number proving the crew-level fix held in the field. Product owned Pendo and the quantitative side of the beta, and that tracking never firmed up once real crews were working long stretches with no connection. I didn’t push early enough to define exactly what I needed to see in that data before calling the fix proven, so when the numbers didn’t show up, I had nothing else lined up to check. I can tell you what the problem cost, and that we hit the target in a room with experts in it. I can’t tell you it held in the field, and the field is the only place that counts.</p>

  <p>What I’d change is agreeing on that number before build starts, not after. If there’s real doubt that Product’s tracking will reach you from the field, that’s a reason to plan a manual check alongside it, not a reason to skip planning the number at all. I skipped it. The fix isn’t a bigger analytics platform. It’s a habit: agree on the after-number before build starts, read what Product already tracks on a set cycle, and don’t call an outcome done without a before and an after attached to it.</p>

  <p>The second thing cuts both ways. Product and UX assumed a mobile analytics tool would behave on a low-connection field device; an engineer would likely have known better, or run a spike and found out in a day. That one’s on us for not asking. But it runs the other direction too. The Crew Manager no-go from Milestone was written on a belief that we couldn’t reach the crew data we needed, so we scoped around it. Late in Shaping, an engineering lead mentioned that a more granular Crew API was already coming from the legacy side. We’d spent a whole phase scoping around data we thought was unreachable, and it was already on the way. That’s not an engineering failure any more than it’s a UX one. It’s a structural one, and putting one engineering lead in the room at the end of Milestone fixes most of it.</p>

</section>
