The Notebook in the Truck
A lineman kept a paper notebook in his truck to check his time entry against, because payroll had gotten it wrong before. Getting him to trust the app meant tracing his notebook back to what was actually broken: not the screen, the data model underneath it.
45–60 sec target vs. 2–4 min observed
Target set from industry norms and direct field observation before the work started; 2.5 minutes was the typical figure on a standard job2 corrections to the data model
Entry-per-person to entry-per-crew, and an optional work order to a required one, both surfaced at the Linemen’s Rodeo before a line of production code was written~400 field users, twice, independently
The Linemen’s Rodeo in October, then an instrumented test during a Duke Energy storm simulation in December, at a different utilityJune 2026, in beta
As Timecards Mobile and Timecards Management; one of two at-risk enterprise renewals closed on the presented capability before anything shipped60 seconds per entry
Internal experts. Bulk crew editing meant one entry covered a shift’s worth of a crew, beating the roughly-triple target set for a crew lead, not just meeting itThe notebook in the truck

I was at an American Electric Power site doing field research for a different product entirely, emergency preparedness. Timekeeping wasn’t on my list. But I had the access, so I asked about it, and watched a lineman do it.
He kept a notebook in the truck to check it against.
Not instead of the app. Ahead of it. In the truck, job by job, he’d jot the hours, the job code, the work type on paper. The real entry happened later, on a laptop back at the office. And he kept the paper anyway, because payroll had gotten it wrong before: hours missing, hours billed to the wrong rate.
That single detail turned out to hold the whole project. Everything that follows traces back to it: what was actually broken, why fixing the screen wasn’t enough, and why a foreman still didn’t fully trust the number until he’d checked it himself.
The defect
Here’s what one entry actually took. You find the job, by job number. Then you pick the work type, and that’s not one choice, it’s a tree: travel, assess, repair, then which kind of repair, or wait time. That tree exists because it sets the hourly rate. Get it wrong and it bills wrong. Then you find the work order number and attach it. None of it was filled in ahead of time. None of it was filtered down to the person already logged in.
The target was forty-five to sixty seconds and three to five taps, set from industry norms and from what we’d already watched people do before the work started. What we observed was two to four minutes, with two and a half minutes the number that kept coming up on a standard job. For a crew lead entering on behalf of three people, we set the target at roughly triple, because the work really is three times the size.
The system knew his role. It knew which crew he was on. It knew there was a storm. None of that was used. That isn’t a bad screen. That’s a system that didn’t use what it already knew.
A job runs two to four entries. A shift runs eight to fifteen jobs. On a typical day that’s somewhere around sixty to ninety minutes spent just putting time in. On a storm day, at the heavy end of both ranges, it’s closer to four hours, inside a shift that’s already ten to twelve. Slow entry doesn’t just cost time. It gets put off, and by the end of a long shift, put off means remembered instead of recorded. That’s not a UX complaint. That’s a payroll problem waiting to happen, and it’s the reason the notebook existed in the first place.
What was at stake
That’s the user side. Here’s the business side: two enterprise renewals were in trouble, and timekeeping was one of the named reasons, raised more than once by the customers involved. That came to me through a standing forum I ran with Sales Engineering, where Tino Mathew, our VP of Field Engineering, brought concerns straight from the field back to product.
Most of those same customers had already bought Crew Manager, which models the entire operation: individuals, crews, equipment, vehicles, lodging. It’s the answer to who’s working today and what they’re working with. The legacy timecard never touched any of it. It read the logged-in user’s own record and nothing else. So a foreman sat in a truck and retyped his crew by hand, every entry, out of a system his own company had already paid us for.
The other thing in the room was history. Arcos had built time entry once already, before there was a UX function. It shipped. It missed. So the question I pushed into intake wasn’t whether we could build timesheets. It was why the last one missed, and what had to be different this time. A feature that already exists and isn’t trusted costs more than one that doesn’t exist yet.
Writing the outcome
Arcos ran an adapted Shape Up, with a phase ahead of Shaping we called Milestone. Milestone takes a raw business ask and breaks it into capabilities, not features. It doesn’t scope and it doesn’t set requirements. What came into intake read like a list of things the team would do: build a mobile time-entry screen, add a week view, ship Timecards Mobile. None of those lines said what would become true for anyone, and none of them could be wrong, only late.
The way I write an outcome instead has three parts, and none of them is a feature. The outcome: a crew lead can submit time for their crew for the day, from the truck. The signal: the forty-five to sixty second, three to five tap benchmark, the actual number we were trying to move. The no-gos: what we ruled out on purpose. A month view was one. Crew Manager integration was the other, deferred rather than dropped, on the belief that we’d need to rebuild the crew record inside the new platform and sync it back later. That belief turned out to be wrong, and I’ll come back to why.
”Jason and I partnered together from day one… Jason was easy to work with, not afraid to disagree but always in a constructive way.”
Nick Nelson, Sr Product Director, Arcos
Domain ownership, not assignments
Ales Karoza owned mobile time entry. Oswaldo Vasquez owned the back office. Neither of those was an assignment I handed out for this one project. Ales and Oswaldo owned opposite ends of the same data model, mobile capture on one side and back-office reconciliation on the other, and when those two ends disagreed, I didn’t want both of them waiting on me to arbitrate. A domain owner has to be the authority on their end, or the model doesn’t distribute, it just adds a layer of asking permission.
That only works if there’s a system underneath it. Ales and Oswaldo could make experience calls without checking with me, not because I told them to be bold, but because Harmony, the Arcos design system, already answered most of what they’d otherwise have escalated: spacing, states, token semantics, what a destructive action looks like. A design system is what makes autonomy safe. Without one, distributed ownership is just distributed guessing. For the token architecture underneath that, see Harmony: A Design System Built for Machines.
A domain is also what makes somebody promotable. I had it documented as the basis to move both Ales and Oswaldo to Lead Product Designer the next cycle. A layoff reached them first, so neither promotion happened, and I want to be straight that it’s a case I built, not an outcome I delivered.
”You can’t build a promotion case out of a list of assignments. You can build one out of a surface somebody owns.”
What the Rodeo confirmed
The Milestone scope had every crew member entering their own time, which is what Product’s Gainsight and Pendo reports pointed to: entry-per-person, sessions and counts all clean, and time entry breaking down worse than almost anywhere else in the app. That told me where to point scarce field time. It didn’t tell me why. Screen data can show you what happens inside the app. It can’t show you what happens in a truck, at a pole, or at the end of a twelve-hour shift when somebody’s rebuilding hours from memory.
In August, on the same AEP visit that opened this project, I watched a foreman put in time for his whole crew, and saw the mismatch right away. One utility isn’t a sample, so it went to leadership as a concern, not a finding. Arcos’s Linemen’s Rodeo was six weeks out, and our CPO asked whether it could do double duty. Shaping held its start about two weeks to use it. Executives signed off on validating at scale before any development resource committed, and no build cycle started until we had the answer. Roughly four hundred linemen and foremen, plus Arcos’s Storm Response Action Committee, for a total incremental cost of one plane ticket.
It held: team leads enter for the crew, and that’s always been true. Worth saying that individual entry was never wrong, only incomplete for the case that matters most. On a blue-sky day, a lineman working alone on a meter install or a damage assessment really is one person. Storm response is where the crew model is the one that counts, and it’s the one we’d scoped a default around by accident.
The second finding was the work order. Every entry needed one attached, including in storm, on an assumption that storm was the exception to nearly everything else about storm. Neither of these was a usability finding. Both were in the data model. Caught on a convention floor, they’re a prototype revision. Caught in build, they’re a schema change.
What changed: the crew shows up already filled in, with add and remove for whoever’s out that day. Hours go in for everybody at once, or you switch over and enter them one at a time. The work order requirement got a notes field in the first release, which wasn’t the real answer, but was the honest one for a v1.
”He suggested and then built a vibe coded prototype so that we could conduct user testing as quickly as possible… which ultimately saved us weeks of development time and put our product on a successful path.”
Nick Nelson, Sr Product Director, Arcos
Why twice, and how we measured
We ran research twice on purpose, and a third time because the chance showed up. Problem validation happened at Milestone, before a solution shape existed to bias the answer. Solution validation happened during Shaping, at the Rodeo.
”Jason made a deliberate effort to bring frontline perspectives into the UX process, not just as a validation step at the end, but while we were still defining the problem and shaping the experience.”
Tino Mathew, VP of Field Engineering, Arcos
The third pass came in December, mid-build, and it was opportunistic rather than planned. My team built a browser analog of the React Native mobile app already in development, coded agentically against Harmony, reproducing the same components in React and MUI. Because it was built from the same tokens as the real app, the numbers it produced carried over. I put Pendo tags on it and ran an unmoderated test during a Duke Energy storm simulation, surfaced through Arcos’s Storm Response Action Committee, the only setting where you get storm conditions and a willing audience at the same time. We measured against baseline on time to goal, click-backs, and error corrections.
The numbers told us the flow worked. They didn’t tell us why. One field user, using the app during the simulation, put it this way:
”It’s like it knows who I am and where in the process I am at… it just drops me there, and this is what I need to do right now.”
A field user, Duke Energy storm simulation
Two negative findings came out of the same session, and reporting them is what makes the positive result credible. One user wanted the work order auto-attached to the job rather than found manually, which the notes-field stopgap from the Rodeo hadn’t actually solved. Another told us our equipment naming was wrong: we called it a bucket truck, his crew calls it a cherry-picker, and the mismatch would confuse people in the field. When the interface is a list you pick from, naming decides whether someone finds the right row on the first try. Neither finding changed v1. Both went into v2 planning.
What shipped


Timecards Mobile and Timecards Management shipped June 2026, in beta. One of the two at-risk renewals closed on the presented capability, before we’d shipped anything. That’s the outcome I’d point to first. The second was still open when I left; Timecards was a factor in it, but I’d put it third, behind API extensions and platform speed. I don’t want to overclaim that one.
A crew lead could now submit time for their crew for the day, from the truck, which was the outcome we wrote at Milestone, in the same words. In alpha, with internal experts, a single entry came in at sixty seconds, the top of the original target range. Because crew edits were bulk, that same sixty seconds covered a whole crew on a typical shift, where everyone worked the same job, the same work type, the same hours. The target for a crew of three was roughly triple the individual benchmark. Bulk editing meant one entry beat that target rather than just meeting it.
”It’s bearing fruit in our recent production rollout and the positive feedback coming back from customers.”
Rusty Davis, Engineering Manager, Arcos
Whether the lineman who prompted this whole project still keeps his notebook, I don’t know. Faster entry is why he might not need it in the field anymore. Whether payroll stopped getting it wrong often enough that he’d trust the app without paper is a different question, and I don’t have that number either.
What I’d do differently
I never got a hard number proving the crew-level fix held in the field. Product owned Pendo and the quantitative side of the beta, and that tracking never firmed up once real crews were working long stretches with no connection. I didn’t push early enough to define exactly what I needed to see in that data before calling the fix proven, so when the numbers didn’t show up, I had nothing else lined up to check. I can tell you what the problem cost, and that we hit the target in a room with experts in it. I can’t tell you it held in the field, and the field is the only place that counts.
What I’d change is agreeing on that number before build starts, not after. If there’s real doubt that Product’s tracking will reach you from the field, that’s a reason to plan a manual check alongside it, not a reason to skip planning the number at all. I skipped it. The fix isn’t a bigger analytics platform. It’s a habit: agree on the after-number before build starts, read what Product already tracks on a set cycle, and don’t call an outcome done without a before and an after attached to it.
The second thing cuts both ways. Product and UX assumed a mobile analytics tool would behave on a low-connection field device; an engineer would likely have known better, or run a spike and found out in a day. That one’s on us for not asking. But it runs the other direction too. The Crew Manager no-go from Milestone was written on a belief that we couldn’t reach the crew data we needed, so we scoped around it. Late in Shaping, an engineering lead mentioned that a more granular Crew API was already coming from the legacy side. We’d spent a whole phase scoping around data we thought was unreachable, and it was already on the way. That’s not an engineering failure any more than it’s a UX one. It’s a structural one, and putting one engineering lead in the room at the end of Milestone fixes most of it.
Team
- 2 Senior Designers (mobile and back office)
Agent view
This page has a machine-readable twin, generated from the same source rather than written alongside it.
See how an agent reads this page