We built NavigateAI with a specific set of assumptions about how field crews work and what they need from a guidance tool. Some of those assumptions were correct. Others turned out to be wrong in ways that we would not have discovered without putting the product in front of real crews on real jobs. This post is an account of what we expected, what happened, and what we changed as a result.
We are writing this publicly because we think the pattern of field software assumptions failing in contact with reality is common enough that other teams building in this space will recognize it. If one of these is a near-miss you avoided, that is useful to know. If one of them is a mistake you are currently making in your own product, we hope this saves you time.
What We Expected Going In
Our design assumptions before the first pilot were three. First: field crews would find the step-by-step guidance useful immediately, because it reduces the cognitive load of keeping a mental sequence. Second: the offline mode would matter most during inspections in basements and mechanical rooms. Third: the photo capture requirement would feel natural to crews who already take photos during jobs and attach them to emails or text messages.
Two of these were roughly right. One was wrong in an instructive way.
What the First Day Actually Looked Like
We ran the first pilot session with a two-person maintenance crew doing preventive maintenance rounds at a mixed-use property in the Denver area. The crew had been informed about the pilot but had received no formal training, which was intentional. We wanted to see what happened when someone opened the app for the first time with a job queue in front of them.
The first thing that happened was that one of the technicians asked if they could keep using their clipboard alongside the app until they knew how to use it. We said yes, which turned out to be the right answer. By the third job that morning, the clipboard was in the truck. By the end of the first week, it was not being taken out at all.
What we saw that day confirmed the first assumption: the guided sequence reduced the number of times a technician had to think about what came next. On an HVAC preventive maintenance job that the more experienced technician had done dozens of times, the sequence did not add new information. It removed friction: they did not have to hold the sequence in memory while handling the equipment. For the less experienced technician, it was more obviously useful as a reference for what to check and in what order.
Where the Offline Assumption Was Both Right and Wrong
The assumption that offline mode mattered most for basement and mechanical room inspections was correct in the obvious case. But we had missed a more common scenario: signal dropout at street level in buildings with significant steel and concrete, which covers most of the urban multifamily properties we were piloting in. Signal loss was not occasional. It was frequent, unpredictable, and had nothing to do with basement depth.
This meant our write-local architecture was more important than we had framed it. We had positioned it as the edge-case solution for specific environments. The pilots showed us it was the default operating mode for a significant portion of the job queue. We updated our product documentation and our onboarding language to reflect this. The architecture did not need to change. Our description of why it mattered did.
The more significant offline lesson was about sync timing. We had assumed crews would want photos to sync immediately when connectivity resumed. Two of the pilot participants told us they found the sync notifications disruptive during a job. They did not want their phone to do background work during a job run. They wanted sync to happen after the job was submitted. We added a "sync on job close" setting that several pilot participants used consistently. The default sync behavior remains aggressive, but the option is available.
The Photo Capture Assumption: Where We Were Wrong
The assumption that photo capture would feel natural to crews who already take photos during jobs turned out to be wrong in a specific way. Crews do take photos during jobs. But they take photos of problems, not of normal states. The pattern was: inspect, note any issues, photograph the issues, move on. The step-locked photo requirement in NavigateAI asked them to photograph every step, including steps where everything looked fine.
The reaction from several pilot participants was consistent: "Why am I photographing this if it looks normal?" The answer, which we had not communicated clearly enough, is that the completeness record requires documentation of normal state as much as documentation of problems. A photo that shows a component in normal condition is part of the record for the same reason a photo that shows a defect is: it documents the state at a specific point in the sequence.
We addressed this in two ways. First, we added optional annotation to the capture interface so a technician can mark a capture as "normal, no action" without adding a note. This reduced the friction of photographing a step where nothing was wrong. Second, we changed the onboarding language to explain the evidentiary purpose of the complete record, not just the completeness requirement. Pilots who understood why they were capturing normal-state photos adopted the behavior more quickly than pilots who understood only that the step required a photo.
What We Changed After the Pilot and Why
The three concrete changes that came out of the first pilot were the sync-on-close setting, the normal-state annotation, and a restructuring of the job type library to better match the variety of PM job types crews actually ran.
The job type library issue was the most significant. We had launched the pilot with nine job types covering common preventive maintenance and inspection categories. The pilot revealed that the categories were too broad. A "HVAC PM" category that covered both split systems and packaged rooftop units produced sequences that were not well-matched to either job type specifically. Technicians either skipped steps that did not apply to the equipment they were servicing or flagged them as inapplicable so frequently that the inapplicable flag became noise.
We broke the broad categories into narrower ones, adding job-type specificity at the cost of more job types in the library. Technicians who had been selecting "HVAC PM" and mentally filtering the sequence now select "RTU Rooftop PM" or "Split System PM" and get a sequence that is well-matched to the equipment. The inapplicable flag rate dropped significantly in the follow-on pilot sessions.
What We Are Not Claiming About the Pilots
We are not saying the pilots proved that NavigateAI works for all field operations. The pilot participants were maintenance teams at properties where the crew was willing to participate and the management was open to the experiment. That is a favorable selection. Organizations where crews are resistant to new tools, where management is not invested in making the change succeed, or where job types are not well-covered by the current library are not represented in our pilot data.
We are also not claiming that the changes we made based on the pilot are final. Product iteration based on field feedback is ongoing. The pilot gave us enough signal to make specific, targeted changes. It did not give us a complete picture of everything that needs to change. Running more pilots with different crew types, different property types, and different organizational contexts will surface more of what we got wrong. We are actively seeking those pilots.