Data FirstPilots
Data FirstPilotsPhysical AI

Why Your Physical AI Demo Works in the Lab and Fails in the Field

The demo worked on your test data. On a real customer's machines, it falls apart. It's rarely the model's fault — it's the gap between lab data and field data. Here are the five gaps we see most, and how to close them.

By Rightshift Team

·

September 15, 2026

·

7 min read

"The demo worked on our test data. On a real customer's machines it falls apart."

We hear some version of this in almost every first call. The founder has a working model, a good demo, and an excited customer — and then the pilot results come back flat. The instinct is to blame the model and start tuning. Usually, the model is fine. The data isn't the same data.

Here are the five gaps between the lab and the field that we see most often.

1. Clean samples vs. noisy reality

Lab data is collected carefully: sensors freshly calibrated, mounted correctly, recorded in controlled conditions. Field data comes from sensors that were installed in a hurry, knocked loose, painted over or replaced with a different model.

What to do: Treat data quality as a first-class signal. Detect flatlines, spikes and impossible values before they reach the model, and track how much of each site's data is usable. A model that's 90% accurate on clean data and fed 30% garbage won't be 90% accurate.

2. One machine vs. many machines

A model trained on a handful of machines learns those machines — including their quirks. Every customer site has different models, ages, maintenance histories and operating patterns.

What to do: Make sure your training and test data are split by machine and by site, not by random rows. If you test on data from the same machines you trained on, your accuracy number is flattering you.

3. Missing context

In the lab you know exactly what the machine was doing. In the field, a spike in vibration might mean a failing bearing — or just a heavier part on the line that day.

What to do: Capture operating context alongside the signal: load, speed, product, shift, ambient conditions. Context is often what turns an alarm that technicians ignore into one they trust.

4. Labels that don't exist yet

Your demo was labelled by your team. In the field, the "ground truth" lives in maintenance tickets, technicians' notes and spare-parts orders — in another system, written in free text, often days after the event.

What to do: Plan the label pipeline as seriously as the sensor pipeline. Agree with the pilot customer up front how failures and repairs will be recorded, so you can actually measure whether the model was right.

5. Connectivity and timing

Lab data arrives complete and in order. Field devices go offline, buffer data, send it late and occasionally twice. Clocks drift. A model that expects a perfect stream will make strange predictions when the stream isn't perfect.

What to do: Design ingestion for late, duplicate and out-of-order data from day one, and make sure the model can tell "no data" apart from "normal data".

The pattern behind all five

Every one of these is a data problem, not a model problem. That's why we tell founders to make the right shift: data first, then AI. Fixing what you collect and how it moves is usually faster — and far more durable — than another round of model tuning.

Before your next pilot

Ask three questions:

  1. Is our test set split by machine and site, so it reflects a new customer?
  2. Do we know how much of the pilot customer's data will actually be usable?
  3. Have we agreed how outcomes will be recorded, so we can prove the model worked?

If any answer is "not sure", it's worth fixing before the pilot starts, not after. Our approach to designing pilots that convert is in How to Scope a Predictive Maintenance Pilot That Converts to Paid.

Building AI for machines, vehicles or energy?

Free 30-min call. You leave with a first read on your data, even if we never work together.

Tell us what you're building →