After more than twenty-two years around manufacturing and industrial systems, I’ve asked myself this more than once: if a plant already has PLCs running the machines and SCADA or a DCS supervising them (often decades of it, deeply reliable) why did anyone ever buy an MES on top of that?
The honest answer is that control and execution are two different jobs. A PLC keeps the process running correctly right now. ISA-95 draws the layers: sensors and actuators at the bottom, PLCs and SCADA/DCS doing real-time control above them, MES sitting at Level 3, ERP at the top. MES was never about the millisecond control loop. It existed to answer a slower, harder set of questions: how well are we operating, how do we make it better, and how do we stay agile when the same line has to build many different things? It turned machine reality into operations: scheduling, dispatch, traceability, OEE, quality, genealogy, the eleven functions the MESA model lays out. The essence of MES was visibility, improvement, and flexibility, not execution.
Control runs the process; MES (Level 3) exists to run the operation, and every layer is a different technology, joined by standardized seams.
That last one has only gotten louder. Manufacturing has moved hard from single-product lines toward mixed-model production (many variants intermixed on one line, lot size one) and from build-to-stock toward build-to-order and mass customization, where a unit isn’t made until someone orders that exact configuration. Picture a paint line running a mix of car models and colors in sequence, or an assembly line where consecutive units are each a different build, delivered just-in-sequence. That’s the environment MES has to run in now, and it’s a fundamentally different problem from a line that makes one thing all day.
And the people who got the most out of it were the plant managers and shift supervisors running the floor day to day, and the process and continuous-improvement engineers behind them. Their whole job is to make the operation better: more yield, more throughput, less scrap, tighter quality, and the flexibility to absorb a changing product mix without the line falling over. MES gave them the data to see what was happening.
But it never gave them the one thing they actually needed: a safe place to try what might work. To test a change, you changed the physical line: took the downtime, risked the scrap, and hoped. So improvement was gated by the cost of experimenting on a running plant. The MES could show you the problem in high resolution. It couldn’t let you rehearse the fix.
Two assumptions that became the ceiling
I think Physical AI is a genuine step change rather than a smarter feature bolted onto Level 3. MES was built on two assumptions that were completely reasonable when they were made, and that quietly became its ceiling.
The first assumption: model everything explicitly, up front. Every entity, every route, every exception had to be defined in advance and wired together. That’s why MES projects drowned in integration: the standards even name the handshakes (OPC UA between control and MES, B2MML between MES and ERP) because the layers were built by different people in different technologies and never spoke natively. It’s also why the data models got enormous: I worked with a FactoryTalk ProductionCentre deployment whose schema ran well past a thousand relational tables, because that’s what it takes to model every operation explicitly. And mixed-model, build-to-order production multiplies that, since every variant, option, and route has to be defined up front, per unit, before the line can run at all. And it’s why the reporting sprawled. Every persona (operator, quality, maintenance, supervisor, plant manager, corporate) needed a different view, so through the 2010s we built dashboards until there were more than a hundred of them, then stood up SAP BusinessObjects Universes to keep them fed (a Universe is a hand-built semantic layer: you define every dimension and measure before a single report can render). At one customer it got complex enough that we added dedicated BI engineers to the MES team just to keep the reporting alive.
The industry’s own answer had exactly the same shape. Rockwell acquired an analytics company called Incuity in 2008 and evolved it into FactoryTalk VantagePoint; Siemens had XHQ Operations Intelligence, GE had Proficy, and Wonderware (later folded into AVEVA) shipped its own Intelligence layer. Across roughly 2010 to 2020, nearly every major vendor sold an “enterprise manufacturing intelligence” product whose entire job was to normalize MES, ERP, historian, and LIMS data into one “Unified Production Model” and unify the sprawl the other systems had created. The tooling got better every year. The underlying problem (a model you had to build and maintain by hand) never went away.
The second assumption: show the data, and a human decides. Improvement was report-and-decide by design. The MES surfaced the number; a person read it and judged what to do. Which means the human was the throughput limit on improvement: you could only get better as fast as people could read reports, meet about them, and act. The BI engineers weren’t a bug; they were what it took to keep feeding a fundamentally human decision loop.
MES sat where two kinds of fragmentation met, and the industry answered each by adding more layers.
Physical AI and agentic AI can loosen both assumptions, but only if they are working from shared operational context. That does not mean replacing every plant system with one model. It means linking the spatial model, asset identities, live state, orders, and history well enough that a simulation or an agent is not guessing which version of the plant is real. Decisions can then move through software where the risk allows it, with people still holding the gates that matter.
Walk the floor
I can see the change most clearly in the modules I know best. Physical AI and agentic AI are not two separate initiatives there. They often meet in the same operating loop, using the same linked context for different jobs.
Five places where the operating pattern could change. The system boundaries will still depend on the use case.
Material movement and the warehouse. Today the MES tracks moves, reconciles locations, and keeps WIP honest in its tables. A more autonomous version adds AMRs that perceive and navigate, with software re-routing the fleet against live demand rather than a fixed plan. Synthetic data can reduce how much footage has to be collected by hand, and edge systems such as Jetson Thor can keep perception and planning near the vehicle. MES does not become the robot controller. It remains one source of operational truth while the fleet manager, map, orders, and live plant state are tied together around it.
Scheduling and JIS/JIT sequencing. Today MRP and heuristics produce a plan, and when reality diverges, a planner often re-sequences it. That is hardest on a mixed-model line where the order decides whether the right parts arrive just in sequence and whether paint or tooling changes stay efficient. An agent could propose a new sequence from live orders, machine state, and labor availability, then test it against a digital twin before a planner commits it. OpenUSD can hold the composed spatial scene for that simulation in Omniverse; MES and ERP still own the production orders and genealogy that make the scenario meaningful.
Recipe and process design. Today recipes are authored and then validated on real equipment under change control. Simulation offers a cheaper place to explore some parameter changes before a line trial, with an optimizer or agent proposing candidates against the twin. The real process still decides whether the model was faithful enough. I would use the twin to arrive at the controlled trial with fewer bad options.
Quality inspection. Today, on many lines, quality still combines sampling plans, SPC charts, and defects that are flagged after the process has already moved on. In-line vision AI can inspect more of the work, while an agent can assemble the evidence for root-cause analysis or propose an upstream adjustment. The authority to change the process is a separate design decision. One current example is Foxconn’s Bianca SOP assistant: NVIDIA reports that DeepHow used Metropolis video search and summarization with Cosmos models, and that the system helped improve first-pass yield by 3%. NVIDIA is the source for that result, so I treat it as a vendor-reported deployment figure. It is still more useful than a generic accuracy claim because it ties the system to an operating measure.
Maintenance and asset health. Today many plants mix time-based maintenance with reactive work. Condition data such as vibration, temperature, acoustics, and current can help estimate risk before a failure, while an agent can gather the history, check parts, draft a work order, and propose a downtime window. I would still keep diagnosis approval and work authorization explicit. An automatically written work order still needs an approved intervention behind it. An older U.S. Department of Energy operations guide cites industry averages of 25–30% lower maintenance cost and 70–75% fewer breakdowns for predictive maintenance compared with preventive maintenance. I would not apply those broad historical figures directly to a particular plant. Separately, Deloitte reports 10–20% lower maintenance cost and 20–25% more uptime. The useful shift is from logging the event afterward to seeing enough evidence to act earlier.
These five cover only part of MES. The MESA model counts eleven functional areas: detailed scheduling, resource allocation and status, dispatching production units, document control, data collection and acquisition, labor management, quality management, process management, maintenance management, product tracking and genealogy, and performance analysis. I picked the five where the change is easiest to feel. Across them, the direction is from report-and-decide toward systems that can perceive, simulate, and sometimes act from linked operational context.
What ties all five together
None of those five works if every system means something different by the same asset, order, or location. What is needed is shared operational context, not one giant replacement database. The spatial model describes the plant, line, cell, geometry, and physical relationships. MES and ERP remain authoritative for orders, genealogy, and business state. Historians and asset services keep the time series, identity, and live condition data they are built to manage. A semantic and identity layer links those representations so a simulation, a robot, a vision system, and an operations agent can refer to the same thing without pretending all of their data belongs in one file.
OpenUSD is a strong foundation for the spatial part. It began at Pixar as a way to compose large 3D scenes from many authored layers, and that same design is useful when CAD, layout, simulation, and robotics each need a view of the same physical environment. Omniverse can run simulations against that composed scene. It does not remove the work of mapping plant identities and semantics, and it should not be asked to replace transaction systems or historians. Its value is giving those systems a shared spatial frame they did not have before.
The spatial model, identity links, and operational systems form the context together. No one layer owns all of it.
That’s why I keep Physical AI and agentic AI in the same conversation. Physical AI perceives, simulates, and acts in the physical setting; agentic AI reasons over the evidence and coordinates work. They become useful together when the asset identities, spatial relationships, and operating state line up well enough that both are referring to the same plant.
Where I’d temper this
I don’t think the whole pyramid collapses at once, and it would be dishonest to pretend otherwise.
The safety-critical control edge is the hardest part to converge. Deterministic controller loops, safety interlocks, and AI perception workloads have different timing and assurance requirements. Putting a GPU at the edge means vision, condition monitoring, or planning can run close to the line instead of waiting on a distant data center. It does not mean moving a certified interlock into a probabilistic policy. The interesting architecture is the boundary between them: AI can propose or guide an action, while deterministic control and explicit safety functions keep authority over the machine where the risk demands it.
Most plants still do not have this linked context. The reason we needed a hundred dashboards and a hand-built Universe in the first place is that the data was never coherent: different sources, identifiers, and meanings. Point an agent at that mess and it may produce a confident answer from mismatched records; run a twin on bad state and it will give a precise answer to the wrong plant. Building the mappings, ownership, and update paths is still the work. Physical AI does not supply that foundation, and agentic AI does not magically infer it. Both depend on it.
So, is it the next level?
Yes, I think it is, but not because MES gets an AI feature or a chatbot. The change is that more of the operation can be perceived and simulated before a person commits the next move, and some low-risk actions can be coordinated automatically. MES does not disappear. It remains part of the operational record while a spatial twin, live plant services, and AI systems make that record more useful for improvement. The future is still a stack of systems. The difference is whether they share enough context to work as one operating loop.
That’s the version I’d bet on. I’ve been wrong about timing before, and the control edge and the data problem are real gates. But the direction feels less like a trend and more like the shape the whole thing was reaching for all along.
Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.