The deskless majority
Most of the workforce has never had a keyboard, a login, or an hour to spare for training. The interface problem that kept them off enterprise software is the one this generation of AI removes.

In shortDeskless workers are about 80 percent of the global workforce and receive roughly 1 percent of enterprise software spend, and they use AI about three times less often than desk workers. The gap is not enthusiasm, it is that everything was built for someone sitting down. Voice, camera and messaging need no app and no training, which turns the hardest group to serve into one of the easiest. The blockers that remain are identity, shared devices, connectivity and language, none of which are model problems.
Two numbers describe the strangest gap in enterprise technology. Deskless workers are roughly 80 percent of the global workforce, and they receive about 1 percent of the 300 billion dollars a year that enterprises spend on workplace software.
AI is reproducing the gap rather than closing it. Surveys this year put regular generative AI use at 75 percent among leaders and 51 percent among frontline staff, and find frontline workers about three times less likely than desk workers to use AI tools regularly. The explanation is not reluctance. Only 28 percent of frontline workers say their employer encourages AI use, against 55 percent of desk workers, and only three in ten say they have been given clear guidance, against half of desk workers. Nobody built it for them, told them to use it, or showed them how.
What makes this worth writing about now is that the reason enterprise software never reached these people is the specific reason it can reach them today. The barrier was never the value of the software. It was the interface, and the interface is the part that changed.
Why software never reached them
Enterprise software assumes four things: a keyboard, a persistent login, a few hours of training, and a person who can stop doing their job in order to use it.
A technician on a ladder has none of those. Neither does a driver at a loading dock, a nurse on a round, an inspector on a roof, a field engineer in a basement with no signal, or a merchandiser working a store aisle. For those jobs, the software has always arrived after the fact: the work happens, the person remembers it, and at the end of the shift, or the week, they sit down and type some fraction of it into a form. What gets typed is thinner than what happened, later than it happened, and shaped by what the form asks rather than by what mattered.
Everything downstream inherits that. Scheduling is built on estimates rather than measured durations. Parts forecasting runs on part numbers that were guessed from memory. Billing lags because the proof is missing. Compliance records are reconstructions. The expensive problem in deskless industries is rarely that people do the work badly. It is that the record of the work is thin, and every system that depends on it degrades accordingly.
What actually changed
Three interfaces arrived that require no application, no training and no hands.
Voice. A person with their hands inside a machine can narrate what they are doing. The account of the work is created while the work is happening, by the person doing it, at no extra time cost. We wrote separately about where voice actually works, and this is the category where it is least contested: the alternative is not a better app, it is nothing.
The camera. Photographing a thing is faster and more accurate than describing it. A model reads the nameplate, the serial number, the damaged corner, the meter reading, the shelf. The worker takes a picture they would have taken anyway, and a structured record comes back.
Messaging. The most underrated of the three, because it requires no rollout at all. The application is already installed, the person already knows how to use it, it works on a five-year-old handset, and it needs no training for a workforce that may turn over twice a year. Sending a message is not a new skill for anyone.
What the three share is that the learning curve is zero and the work product is produced at the moment of work rather than reconstructed afterward. That is the whole argument.
What it is good for
Four patterns cover most of the value, and they are not the same patterns as desk work.
Capture at the moment of work. A photograph and two sentences become a structured record in the system of record: the asset, the fault, the parts used, the time on site, the follow-up. This is the one that changes the downstream numbers, because it replaces a reconstruction with an observation.
Answers at the point of need. The manual, the history of this specific asset, what the last three people who touched it did, the policy on this exception. Field service first-time fix rates have sat around 75 percent industry-wide for years, and most of the misses are not skill, they are arriving without the right information or the right part. Vendors now claim meaningful improvements from automated pre-job briefings; treat the specific figures as vendor claims and measure your own, but the mechanism is sound.
Dispatch and coordination. Check-ins, ETAs, exceptions, reschedules, proof of delivery. This is high-volume repetitive phone and radio work that consumes a dispatcher's entire day, and it is structured enough to hand over while keeping the exceptions with a person.
Closing the loop. Sign-off, compliance evidence, the billing trigger. In most deskless operations this lags the work by days, and the lag is pure working capital.
The four things that actually block it
None of them are model problems, which is why programs that focus on the model stall.
Identity. Deskless workers frequently have no corporate email, no single sign-on account, and no device the company manages. But an agent that acts on someone's behalf needs to know whose behalf, because permissions, approvals and the audit trail all hang off that. This is the first thing to solve and the most commonly skipped: a person, with a role, tied to the messaging number or the device session, not a shared inbox that everyone uses.
Devices. Shared handsets passed between shifts, personal phones, gloves, cracked screens, no device management. The practical consequences are specific: sessions have to be short and explicit, sensitive data should not persist on the device, and handing the phone to the next shift must not hand over the last person's access.
Connectivity. Basements, rural routes, steel warehouses, tunnels. The requirement is not offline everything, it is that a capture is never lost. Queue it locally, sync when there is signal, tell the person plainly what has and has not been recorded.
Language. Frontline workforces are frequently multilingual, and the record needs to land in one canonical language while the person works in theirs. Voice makes this easier than typing did, and it is an argument for conversation over forms.
Behind all four sits turnover. In workforces that turn over quickly, any interface requiring training is a permanent tax. That is the strongest practical argument for messaging and voice: they are the only interfaces that a new hire already knows how to use on their first shift.
What not to build
Another app. If it requires an install, an account and a tutorial, adoption stops at the people who were already going to adopt.
Something that answers but cannot act. A frontline assistant that returns instructions rather than doing the thing is a manual with a chat window. The agent has to be able to order the part, book the revisit, update the work order and trigger the invoice, which is the same resolution rather than conversation argument we made about voice.
Surveillance with a helpful face. This is the one that kills programs. If the first visible feature is monitoring rather than help, adoption collapses and does not recover, and in a group already getting 28 percent encouragement you cannot afford it. Give people something that saves them the end-of-shift paperwork before you ask anything of them.
A desk workflow on a small screen. A seven-field form is not better on a phone. If the interaction is not conversational, photographic or spoken, it is going to be done badly or not at all.
How we think about it
The deskless case is the clearest illustration of the argument we made about meeting the work where it happens: the same agent, with one identity, one company context, one policy and one record, reachable wherever the person actually is.
In practice that means Fig on mobile as a full client rather than a viewer, with voice and vision so a spoken request or a photograph becomes structured work, and Fig in WhatsApp and Telegram so the people who will never install anything can use it from the app already on their phone. It means the agent can act in the systems of record through connectors, including the ones without APIs, because a record that does not reach the system of record has not been captured. It means approval gates on anything irreversible, since the field is exactly where an unrecoverable action is easiest to trigger by accident. And it means the same audit trail regardless of whether the request came from a phone, a channel or a desktop.
There is a second-order effect worth naming. The work captured this way is precisely the material that has always been missing from the institution's model of itself: what actually happens on site, how long it really takes, which exception occurs how often, what the experienced technician does that the procedure does not describe. That is the ontology and decision model of the operational business, and for most companies it has never existed in any system. Serving the deskless majority is not only the largest unserved group. It is also the fastest way to learn how your operation actually runs.
Key takeaways
- Deskless workers are about 80 percent of the workforce, get about 1 percent of workplace software spend, and are roughly three times less likely to use AI. The gap is provision, not appetite.
- The old failure was the interface: software assumed a keyboard, a login, training time, and the ability to stop working. Voice, camera and messaging assume none of them.
- The value is not acceleration, it is capture. A record made during the work replaces a reconstruction made after it, and everything downstream improves.
- The blockers are identity, shared devices, connectivity and language, plus turnover. None of them are model problems, and all of them have to be designed for.
- Do not ship another app, an assistant that cannot act, or surveillance with a helpful face.
The largest group of workers in the world has been waiting for an interface that does not require them to sit down. It finally exists.


