Pick almost any service today, and there is probably a startup trying to turn it into software: an AI accountant, AI recruiter, AI support rep, AI lawyer, AI travel agent. Some of those products will work. Some will disappear quietly. Most will land somewhere in the middle, where a part of the service becomes software, and the rest still needs a human being.In my opinion, the demo is a weak filter. A capable model can now produce a plausible answer for almost any white-collar task, so the practical test is whether someone can check the result. A quick, objective check lets the service move much further toward a real product. A slow or subjective check, especially one that arrives after the damage is done, means the AI can still help but the operating model has to stay much more human.Coding assistants were an early win for this reason. Software has a natural feedback loop: tests run, code compiles or fails, and the result comes back almost immediately. A model pricing a house, recommending a cancer treatment, or making a hiring judgment lives in a very different world. The answer may sound confident, while the business may wait months to learn whether it was right, and sometimes it never gets a clean read.Why the usual question is too broadResearch on AI exposure already points in this direction. The task is the right unit of analysis, not the job title. Eloundou and colleagues estimate that about 80 percent of US workers could have at least 10 percent of their tasks affected by language models, while about 19 percent could have half or more of their tasks affected. That sounds broad, and it is. But exposure does not automatically mean a product should exist, or that the economics make sense. Acemoglu puts a stricter economic filter on the same idea: once cost is counted, only around 23 percent of exposed tasks look profitable to automate within a decade. The surface area is large. The near-term product opportunity is narrower.This is where the simple “can AI do it?” question becomes misleading. AI may be able to write the answer, summarize the call, draft the document, suggest the treatment, or rank the candidates. But if the company then needs an expensive expert to check every output line by line, the promised software margin starts disappearing. The company has built an AI-assisted service, not a product that can run on its own.There is an older version of the same idea in Autor’s discussion of Polanyi’s paradox: we often know more than we can clearly explain. Work that depends on tacit judgment is harder to automate fully, because even people may struggle to state exactly how the correct answer should be checked. Where AI helps most today, it often speeds people up inside the existing workflow. In one study, support agents handled 14 percent more cases per hour, with the largest gains going to less experienced workers. That is valuable. It still leaves much of the human operating model in place.The real test: can you check the answer?The usual test asks whether the model can do the task. The more useful product question is whether the result can be checked at low cost and soon enough to matter.By verification, I mean a practical check on the result. Can you confirm that the answer is correct? Will two reasonable people agree on the check? Can you run it before the mistake creates cost, risk, or customer harm? When that check exists, AI can act with more autonomy. Without it, every result has to be reviewed by a person, and the product starts looking like a service again.The BCG experiment by Dell’Acqua and colleagues shows the problem clearly. Consultants using AI performed much better on tasks that fit the tool, with output about 40 percent higher in quality. But on tasks that looked similar from the outside and depended on harder-to-read judgment, the same people performed 19 percentage points worse, often without realizing it. That is the dangerous part. Good answers and confident wrong answers can look the same until the workflow has a way to check them. Acemoglu makes a similar point from theory: the hardest productivity gains come from work where there is no objective measure of whether the result is good.I would sort service work into three buckets. A simple check that is cheap and quick enough can let the work productize and run with little supervision. If the check exists only for part of the work, AI becomes a copilot: it automates the verifiable steps and leaves the judgment-heavy part to a person. If the check is expensive or unclear, the core of the service stays human and AI belongs around the edges. The important thing is to apply the test to each slice of work, not to the company label. Most services split once you look closely.Level one: functions inside a companyThe easiest wins usually sit inside back-office functions, where a result is either right or wrong and the business finds out quickly. Software engineering has tests and compilers. Finance operations have reconciliations that need to balance. First-line support has a visible signal in whether the ticket stayed resolved or came back. Data extraction, document processing, scheduling, and transcription against audio all have some version of the same built-in correctness signal.This is why these categories fill up quickly. The model matters, but the check matters just as much. If a billing agent can draft a refund, compare the action against the customer account and the refund policy, and act only when the numbers reconcile, the product has a chance to run on its own. Without that verification layer, the company is really selling a human review process with an AI draft in the middle.Level two: whole services that become productsSometimes the entire service clears the verification test. Transcription is the clean example. A person used to listen to a recording and type it out; now software can do most of the job because the output can be checked against the audio. Bookkeeping and accounts-receivable collections are moving in the same direction. For bookkeeping, the ledger has to balance; for collections, the invoice either gets paid or it does not. The feedback loop is not perfect, but it is concrete and arrives soon enough to manage the work as a product.When the verification loop is early and reliable, the service has a path toward becoming a product. A late or muddy check changes the risk profile completely and can make the same idea expensive very quickly.The expensive failures usually have the opposite structure. Zillow’s home-buying business used AI to estimate home values, make instant cash offers, buy houses, and resell them. The product answer was the price. The check came only after the home was resold. By then the market could move, and the error could already sit on Zillow’s balance sheet. The wind-down involved a $304 million inventory write-down and another $240 million to $265 million in expected losses.IBM’s Watson for Oncology ran into a related problem in a much higher-stakes setting. The recommendations sounded like the kind of work AI should help with, but the system was trained largely on hypothetical cases and had limited real-world patient data. Its advice proved unreliable, raising serious questions about whether this kind of medical recommendation could be safely productized. In both cases, AI could produce an answer, and the business lacked a safe enough way to check that answer at the moment of action.Level three: services that productize in partsMany services split into productized, copilot, and human layers. The line is usually drawn by the quality of the check.Most services are a mix, and the line is often not where people expect it to be. Tasks that feel personal can still productize when the output leaves a useful signal. Personal styling looks like pure taste, yet Stitch Fix turned much of it into a product by letting algorithms propose clothes from a customer’s sizes and past choices, with human stylists finishing the selection. The model has a useful signal to learn from: what the customer keeps. The taste does not become fully objective, but part of the service becomes checkable enough to productize.Travel works the same way. Building an itinerary checks against dates, budget, location, inventory, and availability, so a lot of that work can become software. Rescuing a stranded traveler after a missed connection is different. There is judgment, money, urgency, and human stress in the loop, so the service is harder to reduce to a clean product.The same split shows up in professional services. A legal opinion may resist full productization, while document review and clause extraction can verify cleanly. A hiring decision remains a judgment call, while resume parsing, deduplication, and matching against explicit criteria are much easier to check. Even therapy productizes only in pieces: structured exercises, mood tracking, and skills practice can move into software, while diagnosis, risk calls, and the therapeutic alliance carry the kind of accountability that should stay with a person. The practical move is to automate the slices with a reliable check and leave the liability-heavy judgment with a person. That is outcome-driven innovation in practice.What this means for buildersIf verification is the gate, the core product work includes the checking layer around the model. Take an AI agent that resolves billing tickets. The impressive part may look like the generated answer, but the product depends on everything around it: checking the customer account, applying the refund policy, making sure the numbers reconcile, confirming that the rule allows the action, and watching whether the ticket reopens.That layer is also the defensibility. Anyone can call the same model. Fewer teams can design the workflow that proves the outcome is correct early and reliably enough to run without constant human supervision. This is close to what Foundation Capital describes as vendors selling the outcome and standing behind it. Once a company promises the outcome, verification becomes part of the product itself.Where to startBefore scoping an AI build, I would take the service in front of me and choose one concrete result. Then I would work backward from the evidence: what would prove the result was right, and would that evidence arrive cheaply enough and early enough to prevent real cost?If the answer is easy to confirm and arrives soon, there may be a real AI product hiding there, and the check deserves as much engineering attention as the work itself. If the answer is slow, subjective, or available only after the customer has already suffered the consequence, the honest product shape is probably a copilot or a service with AI inside it.The model’s raw ability will keep improving on its own. The harder strategic question is where that ability can be turned into a product with a working operating model. For services, I would start with verification in plain operational terms. Until you know how the work will be checked before a mistake gets expensive, you do not yet know whether you are building a product or a service with AI inside it.