Sal has sat through enough of these to know the pattern before the deck even loads. The room is enthusiastic. The screenshots look sharp. Someone says "we're seeing real traction," and nobody in the room can name the one thing in the business that actually changed because of it.
That is not a failed pilot. A failed pilot gets killed and everyone moves on. This is worse. It is a pilot that never had a mechanism for succeeding or failing, because it was never connected to anything that mattered in the first place.
Theater pilots are expensive in a way that does not show up on the invoice. They consume a quarter of leadership attention, they let a company report "we're doing AI" without producing anything a board or a P&L would recognize, and they quietly convince the organization that AI does not really move the numbers, when the honest conclusion is that this particular pilot was never built to.
The difference between a pilot and a real deployment is not how long it has been running or how much the team likes it. It is whether the thing ever touched a system the business actually depends on.
Here are the five signs Debsan looks for, and a concrete test for each. Any one of these alone is a yellow flag. Two or more means the pilot is theater.
Ask whoever owns the pilot for the one number that changed in the business because of it. Not a usage stat. A cycle time, an error rate, a dollar figure, something a P&L or an ops review would recognize.
If the answer comes back as logins, queries run, or "the team loves it," that is the tell. Activity metrics measure whether people opened the tool. They say nothing about whether the business is better off, and a pilot can post rising activity numbers for years without ever producing one.
Ask what data the pilot actually ran against. A real deployment eventually has to read from, and often write to, the ERP, the CRM, or the financial system the business actually runs on.
A pilot that only ever ran against an exported CSV, a sandbox copy, or a curated sample set has been testing whether the AI works in general, not whether it works here. That is a legitimate first step. It stops being one the moment it has run for more than a quarter without a plan to touch the live system. The excuse is almost always the same: the real system is too messy, too regulated, or too risky to connect yet. Sometimes that is true. More often it is a reason nobody has scheduled the harder work of actually connecting it.
An AI operator is an AI system with write access inside a company's systems of record, positioned to execute the next step itself instead of surfacing a suggestion for a person to relay by hand. Most pilots never get anywhere close to that. They stay a copilot suggesting an answer inside a sandbox, a smaller and different thing than what got approved in the original budget request.
The test is simple: ask whether the pilot has write access to anything real. If the honest answer is no, the pilot is still auditioning, and it has probably been auditioning longer than anyone wants to admit in the steering committee update.
Find someone who runs an operational function the pilot is supposed to help, someone who is not the pilot's sponsor and was not in the original kickoff meeting. Ask them what the pilot does.
If they cannot answer, or answer with something that does not match what the pilot's own team would say, the pilot has never actually been adopted by the part of the business it was meant to serve. It has been adopted by the team that built it, which is a much lower bar and a much smaller win.
Ask for the calendar date phase two was originally supposed to start, and whether that date has already passed once. Almost every theater pilot has a phase two. It is always close. It is never scheduled against a real date with a real owner attached.
A pilot that keeps expanding its own runway without ever expanding its actual footprint is not slowly succeeding. It is a project that has found a comfortable way to keep existing without ever being evaluated against the outcome it was funded to produce. A real phase two has a start date on someone's calendar and a named owner who is accountable for hitting it. A theater phase two has neither, just a slide that says "coming soon" at every quarterly review.
This is the most common sign, and the easiest to miss because it looks like normal work. The pilot produces an answer, a score, a draft, a recommendation, and someone copies it into the real system themselves.
That gap has a name: the relay gap, the distance between where an AI tool produces an answer and the system of record where that answer has to end up, closed only by a person copying it over by hand. A relay gap that has existed for more than a few weeks is not a temporary rollout step. It is the permanent shape of the pilot, and it means the coordination cost the AI was supposed to remove has just moved one stage downstream, onto whoever does the copying.
An AI Forward Deployed Engineer is an AI implementer who works inside a client's actual ERP, CRM, and financial systems to deploy AI where it reduces coordination cost or decision latency, rather than delivering a generic product or a demo from a distance. That definition is not a marketing description. It is a checklist, and it is the direct answer to all five signs above: a real system, a named business outcome, an operational owner who did not build the tool, a scheduled handoff instead of a rolling phase two, and no person standing in the middle relaying the output by hand.
None of this means the pilot phase itself was a mistake. Starting narrow, on a bounded task where the output can be checked, is exactly right. The mistake is letting a narrow pilot stay narrow indefinitely while calling it deployment in every internal update. A pilot that is honestly labeled a pilot, with a real date and a real metric attached to its next step, is not theater. It is just the first stage of a deployment that has not happened yet.
If you want the fuller picture of what real deployment work looks like day to day, we laid it out in What an AI Forward Deployed Engineer Actually Does. And if you are still deciding whether your company is even ready to start, the five-sign readiness check in Is Your Company Ready for an AI Forward Deployed Engineer? is the natural place to start before you greenlight a pilot at all.
How many of these five signs mean we should be worried?
One is worth watching. Two or more means the pilot has stopped being a pilot and started being a permanent exception to how the business actually runs.
Should we just kill a pilot that shows several of these signs?
Not necessarily. Often the fix is smaller: give it a real system to touch, name an outcome metric, and set a hard date for the handoff, rather than starting over.
Is a long pilot phase always a bad sign?
No. A pilot that stays narrow on purpose, with a scheduled expansion date and a named outcome it is being measured against, is doing exactly what a pilot should do. The problem is a pilot with no schedule and no metric, not a pilot that simply takes its time.
Who should be accountable for closing the relay gap?
Whoever owns the system of record the output needs to land in, not whoever built the AI tool. That handoff is usually the first thing theater pilots skip.