What Running 1M+ Connected Endpoints Teaches You About Startup Data
Our founder built a data platform with more than a million endpoints connected at once. Startups are far smaller — but the same lessons apply from the very first device. Here are six of them.
Our founder, Paranjay Menaria, spent 15 years building data systems — including a platform with more than a million endpoints connected at the same time, collecting data across the US, Europe and China.
A physical AI startup won't have a million devices on day one. But the lessons from that scale show up surprisingly early — often by the time you have your first few hundred devices in the field. Here are six worth learning before you need them.
1. Devices will disconnect, constantly — design for it
At scale, "some devices are offline" isn't an incident; it's the normal state of the system. Vehicles drive into tunnels. Factories lose Wi-Fi. Gateways reboot.
The lesson for startups: buffer data on the device, send it when the connection returns, and make sure your pipeline can handle data that arrives hours late. If your system assumes a perfect stream, it will be wrong every day.
2. Every device is slightly different
Different hardware revisions, firmware versions, sensor suppliers and installation quirks mean that no two devices behave exactly alike. At a million endpoints, rare edge cases happen every minute.
The lesson for startups: record which hardware and firmware every reading came from. When something strange appears in your data, the first question will be "which devices?" — and you'll want an answer in seconds, not days.
3. Schemas change — plan for it
Over time, devices gain new sensors, fields are renamed and units change. With a large fleet, you'll always have old and new versions running side by side.
The lesson for startups: version your data formats from the first device, and make your pipeline tolerant of both old and new versions at once. Retrofitting this later is painful.
4. Send less, but send the right thing
At scale, sending everything at full resolution is unaffordable. You learn to decide what is processed on the device, what is summarised, and what is sent raw only when something interesting happens.
The lesson for startups: decide deliberately what stays at the edge and what goes to the cloud — we cover how in Edge or Cloud?. Costs that look small with ten devices grow with every customer.
5. Real-time and batch are different jobs
Some questions need answers in seconds: is this machine about to fail? Others are about history: which machines fail most, and why? Trying to serve both from one pipeline usually serves neither well.
The lesson for startups: separate the fast path (live alerts and dashboards) from the slow path (storage, analysis and model training) early, even if both are simple to begin with.
6. Data quality is a product, not a clean-up job
At large scale, there is no way to "fix the data later". Quality has to be checked continuously, automatically, as data arrives — missing values, impossible readings, silent sensors.
The lesson for startups: build simple quality checks into ingestion from day one, and make the results visible. Your models — and your customers — will only trust predictions built on data you can vouch for.
The common thread
None of these lessons require a million devices to matter. They matter from the first pilot, because that's when the habits — and the architecture — are set. It's much cheaper to start right than to migrate later.
That's the idea behind Rightshift: give physical AI startups the judgment that usually only comes from having run systems like this — years before they could hire someone who has. Make the right shift. Data first. Then AI.
If you're about to scale from your first devices to your first customers, tell us what you're building.
Building AI for machines, vehicles or energy?
Free 30-min call. You leave with a first read on your data, even if we never work together.
Tell us what you're building →