Why Environmental AI Needs Context, Not Just More Data

Posted

Environmental telemetry seems like a perfect match for AI. Buildings, industrial sites and cities already generate streams of measurements such as CO2, PM2.5, temperature, humidity, occupancy, energy use, ventilation status and outdoor conditions. The more sensors are added, the easier it is to assume that context will take care of itself. If a model can see enough readings across enough spaces and enough time, it should theoretically be able to detect hidden patterns, flag anomalies before tenants complain, and translate the same data into an engineer’s diagnosis, an owner’s summary or an ESG metric.

Unlike typical corporate data, which is often scattered across emails, spreadsheets, reports and human memory, environmental data is dense, continuous and increasingly cheap to collect.

The problem is that measurements are not the same as context. Indoor air quality, ventilation performance and energy use are produced by systems that change with weather, occupancy, building design, maintenance history and human behavior. A 2025 study from London illustrates this well. Researchers collected almost 150,000 hours of indoor air-quality data from 258 households and found that giving residents access to real-time pollution readings reduced indoor PM2.5 by 17% overall, and by 34% during the hours when people were at home. The information changed behavior, and that behavior changed the physical environment.

For AI systems, this is where the difficulty begins. The model may have the measurements, but it still has to understand what changed around them, why it changed and whether the same recommendation would work under different physical conditions.

The Problem of Inherited Errors

The first context problem is informational. Because general-purpose AI systems are trained on large bodies of public and licensed information, a simplified explanation or an outdated reference may appear far more often than the correct technical detail.

One clear example sits inside a common myth in indoor air quality: that ASHRAE sets a hard 1,000 ppm limit on indoor CO2. A 1989 version of the standard did include a concentration-based limit of approximately 1,000 ppm, but that limit was removed in later revisions. Despite this, the figure has been repeated so often that ASHRAE had to address the confusion directly, clarifying that Standard 62.1 should not be cited for a number it has not contained in more than three decades.

Some general-purpose AI systems may now answer this specific question correctly, in part because the “1,000 ppm limit” myth generated enough public discussion and dedicated corrections to shift the balance of information available to the model. The correction became reliable because it was visible, authoritative and sufficiently represented in the data a model could learn from. Many narrower technical questions in environmental monitoring never receive that kind of visible correction, which means the same imbalance can persist elsewhere.

This is one factor that can contribute to hallucination: a confident, well-cited answer built on a claim that is not actually correct. A recent paper by Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala and Edwin Zhang describes one part of this problem mathematically. The authors explain that standard training rewards models for producing an answer rather than acknowledging uncertainty, while statistical frequency helps shape the output. In a field where the correct technical detail may appear once in a specialist document while a simplified, inaccurate version appears hundreds of times across the web, that mechanism can work directly against accuracy.

Correct Answers Can Still Lead to the Wrong Decisions

Factual accuracy is only part of the problem. In environmental systems, the harder question is often whether the model has enough hyperlocal context to interpret a correct measurement properly.

Consider two identical office buildings next to the same busy road. Both record a spike in indoor CO2 during the morning rush hour. Opening the windows could lower CO2 in both buildings. But suppose Building A has windows facing the street, while Building B has windows facing an inner courtyard. The same recommendation would produce different consequences. In Building A, opening the windows may reduce CO2 while bringing more traffic pollution indoors. In Building B, it may solve the CO2 problem without the same exposure.

Handling this well requires information no dataset provides by default: where the sensor is, what the space is used for, how the HVAC system is configured, what is happening outdoors, and whether doors or windows are open. With that extra context, a system might recommend keeping windows closed despite elevated CO2, as in Building A. That recommendation can sound counterintuitive until the model accounts for outdoor pollution, the building’s air intake configuration and the current ventilation state. With enough reliable data, the counterintuitive recommendation can be the more accurate one.

The same need for context extends to monitoring itself. A few hours of measurements can show that something happened, but they will not show whether it happens every morning, only when a particular room is occupied, only under certain weather conditions, or because a piece of equipment is behaving differently from its intended operation.

Continuous monitoring can help reconstruct what happened before, during and after an event. It can also reveal cases where a recommendation looked correct on paper but failed in practice: a damper was not fully open, a filter was installed incorrectly, or someone manually changed a setting. Without monitoring what happens afterward, software can report a successful intervention while air quality remains unchanged or improves only marginally. Continuous monitoring closes the loop between recommendation and reality.

There is a broader business case for getting that loop right. Harvard’s COGfx research, which examined three ventilation strategies and four HVAC systems across seven U.S. cities, found that the indoor conditions associated with a doubling of cognitive function test scores could be achieved at an energy cost of roughly $14-$40 per person per year. The researchers estimated that the resulting improvement in productivity could be worth as much as $6,500 per person per year. Of course, these are research estimates, not guaranteed returns, but they show why indoor environmental conditions matter well beyond a facilities-management budget.

What Reliable Environmental AI Requires

Reliable AI-supported environmental decision-making should reflect these physical realities in its design. It should be able to say when the available information is not enough to distinguish between two possible causes, and it should lay out the alternatives rather than guess. If condition A is confirmed, response X may be appropriate. If condition B is confirmed, response Y may be safer. Calculations should be performed deterministically in code rather than left to the language model itself. Outputs should be expressed as ranges or probabilities where appropriate, not as one precise-looking number. The system should also check what happened after an intervention instead of assuming that its recommendation worked.

It is not always easy to tell from the outside whether a system is equipped to handle these physical realities. For business leaders evaluating AI for buildings, industrial sites or other physical environments, six questions are a useful starting point:

  1. Source of truth. What is the underlying knowledge base, and how many independent primary sources support it?
  2. Filtering. How does the system identify outdated standards, repeated claims and unsupported information?
  3. Uncertainty. What happens when the available data is incomplete or ambiguous?
  4. Model or rules. Is there a genuine learning model behind the system, or are fixed thresholds simply presented through an AI interface?
  5. Metadata. What physical and spatial information does the system need before it can give a meaningful recommendation?
  6. Verification. How does it establish that a recommended action actually worked after it was implemented?

Vitalii Matiunin is the Co-Founder & CEO of Airvoice. A physicist by training with degrees in Physics and Space and Engineering Systems, Vitalii combines a scientific background with more than 10 years of experience in international business development and technology commercialization. Before founding Airvoice, he worked in scientific research and healthtech.

Environment + Energy Leader