Data FirstPhysical AI
Data FirstPhysical AIFounders

What Data to Collect Before You Build Physical AI

Most physical AI startups collect too little of the right signal and too much of the wrong one. Here's how to decide what to capture from your machines, vehicles or sensors before you train a single model.

By Rightshift Team

·

September 29, 2026

·

7 min read

The most expensive mistake in physical AI isn't picking the wrong model. It's realising, six months in, that the data you need was never collected — and that the devices in the field can't be updated to collect it without a site visit.

Models can be retrained in an afternoon. Missing history can't be recovered at all. That's why we start every engagement with the data, not the AI.

Start from the decision, not the sensor

Founders usually start with the question "what data do we have?". The better question is "what decision should the product help someone make?"

  • A maintenance technician deciding which part to replace
  • A fleet manager deciding which vehicle to pull off the road
  • A plant operator deciding whether to shut a line down

Each decision tells you what you need to know, how early you need to know it, and how wrong you're allowed to be. Work backwards from there to signals. You'll usually find that half the data you planned to collect doesn't help the decision — and one or two signals you never thought of are essential.

The five things to capture, beyond the raw signal

Raw sensor values are rarely enough. Without context, a vibration reading is just a number. For every signal, plan to capture:

  1. Timestamps you can trust. Device clocks drift. Record when a reading was taken on the device and when it arrived, and know which one your models use.
  2. Identity. Which machine, which component, which site, which firmware version. A model that can't tell two machines apart can't learn either of them.
  3. Operating context. Load, speed, temperature, shift, product being made. "High vibration" means something different at full load than at idle.
  4. Events and outcomes. Maintenance tickets, part replacements, failures, operator notes. These are your labels — and they usually live in a different system from your sensor data.
  5. Data about the data. Gaps, sensor faults, calibration changes. A flatlined sensor looks like a very healthy machine unless you know it's broken.

The fourth item is the one most startups miss. Sensor data tells you what happened; service history tells you what it meant. Joining the two is where most of the value in physical AI comes from.

Sample rate: collect for the question, not the maximum

It's tempting to stream everything at the highest rate the hardware supports. That multiplies storage and bandwidth costs, and it rarely improves the product.

Ask what the fastest meaningful change in your signal is. Bearing faults may need high-frequency vibration data; tank levels may need one reading a minute. Many teams end up with a mixed approach: high-rate data processed at the edge into features, with summaries sent to the cloud and raw data kept only around interesting events.

Decide early what stays at the edge

Some data should never leave the device: it's too large, too sensitive, or only useful for a few seconds. Some must reach the cloud to train models across your whole fleet. Making that split deliberately — before you deploy hundreds of devices — saves a painful migration later.

We cover this trade-off in more depth in Edge or Cloud?.

Write it down as a data contract

Before you build, write a one-page data contract for each source: what's collected, at what rate, in what units, with what identifiers, and who owns it. It sounds bureaucratic. It's the cheapest insurance you'll buy. When a firmware update silently changes units from Celsius to Fahrenheit, the contract is how you'll notice.

A simple checklist

Before you train your first model, you should be able to answer yes to these:

  • We know which decision the product supports, and what data that decision needs
  • Every reading has a trusted timestamp, a device identity and its operating context
  • We can join sensor data to maintenance and outcome records
  • We know which data stays at the edge and which goes to the cloud, and why
  • Each source has a written contract, and we'd notice if it changed

If you can't tick these yet, that's normal — it's exactly what our 2-week Data & AI Diagnostic is for. Data first. Then AI.

Building AI for machines, vehicles or energy?

Free 30-min call. You leave with a first read on your data, even if we never work together.

Tell us what you're building →