Skip to main content
KŌJŌ Stack logo
KŌJŌ Stack
All Use Cases

Data Centers

A data hall is six protocols and three monitoring systems pretending to be one facility. Meters, PDUs, UPSs, switchgear, chillers, and BMS controllers each speak their own dialect, and BMS, EPMS, and DCIM each keep their own point list. KŌJŌ Stack reads every one of them natively-SNMP, Modbus TCP, BACnet/IP, DNP3, OPC UA, MQTT-normalizes at the point of ingestion, and publishes one canonical model of the facility that every consumer subscribes to. The same model deploys unchanged to the next site, so a portfolio finally compares to itself.

OneModel of the Facility

Architecture Highlights

Structured at the Source

Native Protocol IngestionISA-95 Facility NamespaceEdge RBE & CELDurable Ordered ReplayEdge MCP Server
Industry Challenges

The Problem

1

Every Subsystem Speaks a Different Protocol

Branch meters talk Modbus TCP, UPSs and PDUs talk SNMP, chillers and CRAH units talk BACnet, switchgear talks DNP3 or IEC 61850, and newer gear talks MQTT. Each dialect arrives with its own per-vendor gateway and translation layer, and every one of them is another box to license, patch, and debug.

2

Siloed Point Lists, Duplicated Polling

BMS, EPMS, and DCIM each hold their own point list, so no single canonical model of the facility exists. Worse, each polls the same devices independently-three systems hitting one PDU over a fragile serial or TCP link, with polling load scaling by the number of consumers rather than the number of devices.

3

Raw Values With No Units, Scale, or Provenance

The same measurement arrives as watts from one device family and kilowatts from another, with per-vendor scaling factors and naming that share nothing. Nothing on the value records which device answered, when, or whether the reading was good-so every downstream consumer re-derives that context, differently.

4

Blind Spots When the Link Drops

When the connection to the historian or the cloud goes down, the record simply stops. Without local buffering and ordered replay, the outage becomes a permanent hole in the facility history-exactly across the window an operator most needs to reconstruct afterwards.

5

Alarms Batched Instead of Processed at the Edge

Threshold crossings and state changes are routed through batch-oriented collection before anyone evaluates them. A breaker trip or a thermal excursion is an event that needs a decision now, not a row that shows up in the next collection cycle.

6

On-Site AI Cannot Reach the Facility

Facility data sits behind vendor UIs with no machine-readable, scoped way in. AI tooling and agents brought in for capacity planning, anomaly detection, or diagnostics have nothing to read-so they either get a screen-scraping integration or they get nothing.

What Breaks Without This

What Fails in Traditional Architectures

Without structured, prepared data at the first mile, downstream systems inherit every inconsistency, gap, and limitation of the raw source data.

1

Every Site Models Racks Differently

Left to itself, each facility invents its own naming for racks, rows, and feeds. Comparing PUE across sites, benchmarking capacity, or rolling a fleet report then means reconciling models by hand-work that scales linearly with every facility added and never stays done.

2

You Pay to Ship Raw Telemetry

Sending full-rate, unfiltered telemetry upstream turns bandwidth and cloud egress into a line item that grows with every device installed. Most of that volume is values that never meaningfully changed, paid for at full price on the way to a store that will compress them anyway.

3

Point-to-Point Integration Sprawl

Each consumer-Kafka, the lakehouse, the historian, an ML pipeline-gets wired separately to each site. The connection count is consumers times sites, and every protocol change, credential rotation, or schema tweak has to be chased across all of them.

4

WAN Drops Take History With Them

Site links fail, and without durable buffering with ordered replay the gap is permanent. Fleet reporting quietly runs on incomplete records, and nobody finds out until an analysis depends on the window that went missing.

5

Per-Site Setup and Unmanaged Drift

Connectors configured by hand at each facility start similar and diverge. Two years in, no one can say what any given site is actually running, and a fleet-wide change becomes a site-by-site expedition rather than a push.

6

Models Trained on Telemetry Nobody Trusts

Analytics and ML built on gap-filled, unlabeled, per-site telemetry inherit every inconsistency underneath. Features engineered at one facility do not transfer to the next, and a model that cannot tell a real reading from an interpolated one cannot be trusted with a capacity or cooling decision.

KŌJŌ Stack Solution

How KŌJŌ Stack Helps

Every Domain Read Natively, No Gateway Layer

SNMP v1/v2c/v3 reaches UPSs, PDUs, CRAC units, switches, and servers through HOST-RESOURCES-MIB. Modbus covers circuit-level power metering, BACnet covers chillers and building control, and DNP3 covers switchgear and the utility tie. OPC UA, MQTT, and Kafka cover the rest. Device templates map a family's OID tree to named tags rather than raw numbers, and an attestation flow records what the hardware actually answered beside what the template claimed.

Poll Once, Serve Every Consumer

The runtime becomes the single reader on the OT network. BMS, EPMS, DCIM, capacity planning, analytics, and AI tooling all subscribe to the namespace instead of adding another poller to the same fragile link. Polling load stops scaling with the number of systems that want the data.

Normalized and Attributed at Ingestion

CEL expressions apply scaling, offsets, and unit conversion at the edge, so a kilowatt is a kilowatt regardless of which device family reported it. Every value carries tag_id, timestamp, value, quality, and source context, and protocol-specific metadata is preserved rather than discarded on the way through.

One Namespace: site.hall.row.rack.device

Every facility maps into the same ISA-95 hierarchy, and MQTT, Kafka, and lakehouse topics are generated from that one definition. A rack in hall B addresses the same way as a rack three sites away, which is what makes portfolio-level questions-PUE by site, capacity trend by row-a single query instead of a reconciliation project.

Sub-10ms Locally, Durable When the Link Drops

Threshold and state-change logic runs in local pipeline execution paths with sub-10ms processing latency, so events are evaluated on site rather than upstream in a batch. When the WAN or historian link fails, persistent local buffering holds the record and replays it in order on reconnect-the outage costs continuity, not data.

One Delivery Layer, Every Destination

Report-by-exception filtering and edge transforms cut telemetry volume by 90%+ before anything crosses the WAN, then one configuration fans the result out to InfluxDB or TimescaleDB for time-series history, Kafka or AWS Kinesis for event consumers, and S3, S3 Tables with Iceberg, or Google Cloud Storage for the lakehouse. Adding a consumer is a destination, not another integration at every site.

An MCP Server on Site for Agents

A fleet-aware MCP server deployed at the edge gives AI agents a machine-readable path into live and historical facility data. OAuth 2.0 with Dynamic Client Registration establishes per-agent identity, a three-axis scope model bounds what each agent can reach, and 22 read-only tools cover discovery, health, UNS graph queries, and live subscriptions.

One Runtime, Deployed the Same Way Everywhere

The same managed edge runtime installs as a host service, a Docker container, or a Kubernetes workload, and fleet management pushes configuration, namespace models, and pipeline definitions across sites with controlled rollout and rollback. New halls come online from a known-good baseline instead of a per-site connector build.

Technical Depth

Why This Requires First-Mile Data Structuring

A data hall is one of the densest protocol environments in industry, and unusually, all of it describes one physical system. Compute reports over SNMP, meters over Modbus, chillers over BACnet, switchgear over DNP3-different data models, update rates, and quality semantics, each claimed by a different monitoring system running its own polling loop against the same devices. So the question operators most want answered, how a workload change moved cooling load and power draw, becomes a join across incompatible schemas and three unaligned time bases. KŌJŌ Stack reads the devices directly and normalizes at ingestion: CEL applies scaling and units, device templates resolve OID trees to named tags, and every value carries timestamp, quality, and source. What lands is one namespace, site.hall.row.rack.device, that consumers subscribe to instead of poll-and it is identical at the next site, which is what makes a portfolio comparable. None of this is recoverable downstream: a gap-filled value with no provenance stays that way.

Measurable Results

Expected Outcomes

90%+
Cut Before Egress

RBE filtering and edge transforms reduce volume before it crosses the WAN

Zero
Data Lost to Outages

Durable buffering with ordered replay across unreliable site links

Sub-10ms
Local Processing Path

Threshold and state-change logic evaluated at the edge, not upstream in a batch

Own the First Mile

Owning the first mile ensures data centers data is consistent, contextualized, and usable across the enterprise.