The Fragmented Data Problem Nobody in Industrial Safety Wants to Admit

Posted

The global market for AI-powered safety technology is projected to exceed $4 billion by 2033. Billions more are being spent every year on EHS platforms, compliance software, and monitoring infrastructure across construction, manufacturing, oil and gas, and mining. The technology investment is real. The outcomes are not keeping pace with it.

The U.S. Bureau of Labor Statistics recorded 5,070 worker fatalities in 2024, a 4 percent decline from the prior year, and a figure the industry rightly points to as evidence of progress.

But a 4 percent annual improvement, measured against a plateau in recordable injury rates that has persisted for several years, tells us something uncomfortable: we may have reached the ceiling of what the current approach can deliver.

The limiting factor is increasingly no longer the quality of the detection. It is the quality of the data that the detection is running on. That is a problem the industry has been remarkably reluctant to name directly. The answer is not more data. It is better-connected data.

The Organizational Reason Behind the Data Problem

I want to be precise about this, because it matters for how you solve it. Industrial safety data is not fragmented because the technology has not caught up. It is fragmented because safety responsibilities are divided across organizational silos that have no incentive to consolidate.

Procurement owns contractor records and induction documents. HR owns training certifications. Operations owns shift logs and work permits. The EHS team owns incident reports. Legal owns investigation records. Each of these functions has its own systems, its own data governance, and its own definition of what constitutes a safety event. When something goes wrong, you end up doing archaeology across all of them.

The downstream consequence is something the industry rarely discusses openly. The AI safety systems being used across high-risk sectors are being trained and operated on incomplete inputs.

The predictive models that are supposed to flag risk before it escalates are drawing from whatever slice of data a single department happens to own. Across construction and industrial sites, one of the most common friction points is not camera placement, network connectivity, or edge processing. It is the question of what happens after a system flags a risk. Who receives the alert? Does it integrate with the permit-to-work system? Does it update the contractor’s safety record? Can the EHS manager see it alongside the inspection log from this morning?

In most cases, the alert goes to a centralised dashboard that the safety team can see, while the rest of the data ecosystem carries on in parallel, unaware. The risk is logged. The risk context is not.

This creates a situation where organizations invest heavily in detection capability and very little in data connectivity. The result is sophisticated tools generating signals that nobody can act on at the system level, because the system was never designed to receive them.

How AI Amplifies the Problem

There is a version of this problem that was tolerable in the era of paper-based safety management. You could live with fragmentation if you were not asking your data to do very much. The arrival of predictive AI has turned a manageable inconvenience into an active liability.

Predictive safety models work by identifying patterns across large volumes of historical and real-time data. The more complete the data, the more reliable the prediction. When that data is fragmented, the model does not know what it does not know. It will confidently identify risk in the areas it can see and stay silent about the areas it cannot.

There is also a subtler issue around data bias. When AI safety systems learn from historical incident records, they inherit every underreporting distortion baked into those records. For example, workers who stayed silent for fear of blame, incidents classified differently across sites because definitions were never standardized, and contractor events that fell outside the internal reporting system entirely.

Algorithms trained on past incident data that underrepresent certain hazards or worker groups will consistently underestimate risk in those areas. This is not a flaw you can patch. It is a consequence of building intelligence on fragmented foundations.

Solving the Data Fragmentation Gap 

I am not going to pretend this is simple. Consolidating safety data across organizational silos requires decisions about ownership, accountability, and budget authority that cannot be resolved at the EHS manager level. It requires a C-suite decision that safety data is an enterprise asset, not a departmental resource.

What I have seen work in practice is a three-part approach.

First, map the data before buying the technology. Most organizations cannot tell you with confidence where all their safety-relevant data lives, who owns it, or what format it is in. That map is the prerequisite for everything else.

Second, standardize definitions before standardizing systems. The reason safety data is inconsistent across sites is almost always definitional before it is technical. What counts as a near miss? What triggers a permit-to-work? Until those questions are answered uniformly, no amount of integration produces reliable data.

Third, expand the range of safety evidence available to decision-makers. While reported incidents and inspections remain essential, they capture only part of the operational picture. Many organizations are beginning to complement these records with additional observational data, such as connected equipment, environmental sensors, wearable devices, or AI-assisted visual monitoring.

These emerging sources should not be viewed as replacements for existing safety systems, but as complementary inputs that can help validate, enrich, and contextualize operational data.

The Uncomfortable Question for Industry Leaders

If you are a safety leader reading this, the question I would put to you is not whether you have an AI safety platform. Increasingly, everybody does.

The question is: what is it actually seeing? Does your AI system have access to contractor records, shift patterns, permit histories, inspection findings, and live site observations simultaneously? Or is it working from whatever slice of data your EHS team loaded into it?

The industry has spent the better part of a decade talking about making AI work for safety. The conversation that has been consistently avoided is making the data worthy of the AI. Those are not the same conversation, and confusing one for the other is why so many organizations are sitting on sophisticated platforms that are, in practice, running blind.


Gary Ng is the CEO and Co-Founder of viAct, bringing more than a decade of experience implementing technological innovations across the construction industry. 

Environment + Energy Leader