Home/Technology/Helios Grid
Framework · Fleet Operations

Helios Grid

The control plane for distributed intelligence — roll out models, monitor health, and respond to events across thousands of nodes, streaming only telemetry and signal upstream, never raw data. Structured by Tangram, governed by Origi.

Helios Grid
4 capabilities
Lifecycle · rollout · observability · response
6 invariants
Load-bearing, everywhere it runs
Signal
Not data — the defining discipline
Sovereign
Self-hosted, works with the WAN absent
Overview

Declare centrally. Reconcile locally. Report by exception.

A distributed intelligence fabric succeeds or fails on a problem that is neither sensing nor inference: operating the fleet. Hundreds to thousands of edge nodes — carried by teams, mounted on vehicles, embedded in autonomous platforms, or bolted to poles — must be enrolled, configured, kept healthy, updated, and recovered when they fail, across intermittent, contested, and bandwidth-poor links. Doing this by hand does not scale past a handful of nodes; doing it with a conventional cloud device-management service violates the sovereignty, classification, and disconnection constraints these fabrics live under.

Helios Grid is Enkidu's answer: a sovereign, self-hosted control plane, and the third foundational framework of the stack. To the established premise — determinism where it matters, reasoning where it helps — Helios adds its own corollary: declaration where it scales. Fleet state is declared, signed, and reconciled by pull, so a node that has been dark for a day converges to the same state as a node that never lost its link.

    The four capabilities
  • Fleet lifecycle at scale — zero-touch enrolment, node twins, pull reconciliation
  • Safe artifact distribution — signed, staged, health-gated, reversible
  • Fleet-scale observability — within a strict telemetry envelope
  • Autonomous event response — bounded remediation, human escalation
Position in the stack

Three frameworks, three questions

The three interlock rather than overlap: Helios is structured by Tangram, governed by Origi — and Origi and Tangram are operated by Helios, since policy bundles and zone definitions are themselves versioned artifacts. No single component may both decide and grant: Helios observes and executes; Origi judges; Tangram bounds. A compromised control plane can, at worst, propose.

Structure

Tangram

Where does work happen, and under what boundary conditions? Zones, continuums, contractual handoffs, error zones.

Trust

Origi

May this action happen, and can we prove what happened? Trust zones, deny-by-default policy, chain of custody.

Operations

Helios

Is the fabric itself healthy, current, and recoverable? Enrolment, desired state, rollout, telemetry, remediation.

The load-bearing distinction: the data plane is the mission fabric — sensing, inference, fusion. The control plane is Helios. The two share the physical bearers but never share content: mission data never enters a Helios channel, and Helios never adjudicates a mission result. Its refusal to touch mission data is not a limitation; it is the design.

The components

One component at every tier

Helios rides the same bearers as the data plane — with its own, strictly smaller traffic class. The hub is the unit of autonomy: a cluster cut off from the centre continues to reconcile, monitor, and remediate locally. And distribution follows the fabric: a 2 GB model update to a 40-node cluster costs the reachback link one transfer, not forty.

On every node

Helios Agent

The small-footprint control-plane process: identity and enrolment, the reconciliation loop, artifact fetch and verification, fail-safe A/B install, local telemetry collection, and the local runbook executor.

On every hub

Helios Warden

A local sub-control-plane: caches desired state and artifacts for its cluster, aggregates and downsamples telemetry, evaluates alert rules locally, coordinates rollout waves, and holds out-of-band recovery hooks — sustaining fleet operations through central-link loss.

At the centre

Helios Core

The fleet registry and node twins, the signed desired-state store and artifact registry, the rollout orchestrator, the telemetry lake, the event engine, and the fleet operator console — deployable forward and central.

The node twin

Helios's authoritative record of every managed node — identity, hardware tier, installed versions, bound policies, current health, desired state. The twin is what the fleet operator sees; the node is what exists in the field. Operators and the event engine reason over twins, never ad-hoc queries to the field.

The discipline

Six load-bearing invariants

They hold everywhere Helios runs — because the control plane must be as disciplined as the fabric it operates.

01Control never carries data

No mission payload ever transits a Helios channel — the telemetry schemas structurally cannot express mission content.

02State is declared and pulled

Desired state is signed at the centre; nodes reconcile by pull. Disconnection is a normal state, not a failure.

03Every artifact is signed and admitted

Helios distributes; Origi admits. A broken signature or provenance link stops installation at the node — regardless of what the centre ordered.

04Updates are fail-safe by construction

Every mutable install uses an A/B mechanism with automatic rollback — a failed update degrades to the previous working state, never to a bricked node.

05Remediation is bounded

Automated responses come from a signed runbook with explicit blast-radius limits. Anything outside it escalates to a human.

06The control plane is itself governed

Agent, Warden, and Core are agents under Origi — trust zones, deny-by-default policy, every action in the chain of custody. The operator of the fleet is as accountable as the fleet.

Staged rollout

Canary → cluster → theatre.
Gates between, rollback behind.

No artifact reaches the whole fleet at once. A rollout declares its waves and the health gates between them — latency and accuracy within bounds on the canary, no elevated error-zone entries, no trust demotions attributable to the new artifact. A gate failure freezes the wave and triggers automatic rollback. Models swap live through the serving layer without dropping availability; system images install to the inactive slot and commit only on a passing health check.

A model that cannot show its provenance cannot enter the registry; a model not in the registry cannot reach a node. Models are policy encodings — versioned and governed like policy changes.

Signal, not data

Fleet-scale visibility
without fleet-scale bandwidth.

The edge reduces; the fabric forwards signal; raw data never leaves the node on a control channel. Each node holds a signed telemetry envelope — its per-interval budget, priority classes, and degraded-mode behaviour — and each link upstream carries an order of magnitude less than the one below it. The Warden computes a per-node health score an operator can rank a thousand nodes by; alert rules run at the lowest tier that has the data, so a node raises its own alarm even when it is alone in the dark.

Model health is a first-class metric: confidence distributions, abstention rates, drift indicators — the operational feed for retraining triggers and Origi's drift monitoring.

The line that never moves

Automation may reduce capability instantly.
Nothing automated ever increases it.

Runbook autonomy is deliberately conservative — restart, roll back, isolate, shed load, request recovery — and asymmetric in the same direction as Origi's trust zones. Any action that widens capability, or reaches beyond a node's own boundary, takes the Origi verdict path and, where policy demands, a human. Everything outside the runbook escalates with the twin, the correlated event chain, and the telemetry already assembled in the operator's view.

One control plane, many fabrics: the reference deployment is Northwind — the hardest version of the problem — and the same architecture, with civilian bindings, operates Lamassu and fixed-site deployments such as smart health units.

Proven patterns, hardened & self-hosted
Fleet orchestrationFleet Command pattern, sovereign Core
ReconciliationGitOps pull · K3s-class orchestration
Fail-safe updatesA/B installs · OOB recovery
Live model servingTriton-class control API
TelemetryPrometheus-pattern · DCGM-class GPU metrics
Model productionTAO-pattern pipeline → signed registry
Enkidu

Operate the fleet, not the node

Book a walkthrough and watch a staged rollout cross its health gates — canary to theatre — with the WAN unplugged.

Request a briefingTalk to the teamDownload the architecture ↓