Files
Digitaltwin/docs/architecture.md
T

6.8 KiB

Architecture — and where your data plugs in

The one-pager for the engineer who asks "can it take our data?"

Shape

   ┌──────────────────── telemetry sources ─────────────────────┐
   │                                                            │
   │   SimulatedSource        MqttSource         OpcUaSource     │
   │   (working)              (stub)             (stub)          │
   │        │                    │                   │           │
   └────────┴────────────────────┴───────────────────┴──────────┘
                              │
                    TelemetrySource  ← the seam
                    emits "frames"
                              │
              ┌───────────────┴────────────────┐
              │                                │
        AnalyticsEngine                  WebSocket /ws  ──►  React dashboard
        · threshold alarms                     │              · 3D scene
        · learned baselines / z-scores         │              · trend charts
        · least-squares trend projection       │              · alarms
        · alarm latching                       │              · what-if controls
              │                                │
              └──────────► frame.analytics ─────┘
                              │
                        Copilot  ──►  OpenRouter / Ollama / rule engine
                        (context = analytics conclusions, not raw floats)

Everything above the seam consumes frames and has no idea where they came from. That is the whole design: swapping the simulator for your plant means implementing one class.

The seam

server/ingest/source.js defines TelemetrySource. A source does three things:

  1. start() — connect, subscribe, begin producing.
  2. emit(frame) — hand a frame upward. Called on your cadence.
  3. capabilities — declare what it supports (timeControl, faultInjection, setpointControl). The UI reads this and hides controls a read-only source cannot honour, rather than showing buttons that do nothing.

A frame is station states, signal values, WIP levels, KPI rollups and the analytics block. server/ingest/simulatedSource.js is the reference implementation; read it alongside the stub you are filling in.

Connecting your plant

Both adapters are written out as documented stubs with the real interface, the tag map, and the traps called out. They throw an honest "not configured" error rather than pretending to work.

MQTT — server/ingest/mqttSource.js

npm i mqtt

Fill TAG_MAP with your topics, implement start(), set TELEMETRY_SOURCE=mqtt and MQTT_URL. If your broker speaks Sparkplug B (most industrial ones do), add sparkplug-payload and map metric aliases from NBIRTH/NDATA instead of topic strings.

OPC-UA — server/ingest/opcuaSource.js

npm i node-opcua

Browse your server's address space, fill NODE_MAP with real NodeIds (never guess them — they are namespace-qualified and installation-specific), implement start(), set TELEMETRY_SOURCE=opcua and OPCUA_ENDPOINT.

Historian / CSV

Simplest of the three and often the best first step: read the export, sort by timestamp, and emit frames on a timer with a scrub control. You get real customer data in the twin without touching plant networks or waiting on IT.

Four things that bite when the data is real

These are called out in the stubs too, because each one is a way a twin starts lying:

Staleness. Track last-seen per tag and mark a station offline when its tags go quiet. A stale value rendered as live is worse than a gap. The demo already models this — the INS-04 sensor dropout fault holds the last value and flags the station rather than drawing zeros, and the charts stop rather than continuing a flat line that would read as a healthy steady state.

Status codes. OPC-UA gives you StatusCode per value. A Bad or Uncertain reading must not become a data point.

Timestamps. Prefer the source timestamp — when the PLC sampled the sensor — over the server timestamp, which is when the message reached you. Under load they diverge, and trend slopes computed on arrival times are wrong.

Units. Vibration in in/s, temperature in °F, pressure in psi are all common. Convert at the boundary, never downstream, or every threshold in the system is quietly wrong.

What the analytics layer needs from you

Nothing changes. TrendTracker and BaselineBank work on timestamped values regardless of origin. Two things worth knowing:

  • Baselines are learned then frozen. A continuously adapting baseline absorbs a slow ramp, so the exact failure mode this is built to catch would never raise a z-score. Freezing means "different from how this machine normally behaves".
  • Signals declare their own analytics eligibility. In server/sim/stations.js, cumulative: true (tool wear, part counters) and volatile: true (belt speed, setpoints) exclude a signal from anomaly testing — normal accumulation is not an anomaly, and normal state swings are not either. Both remain covered by thresholds and trend projection. When you add your signals, set these flags next to the definition.

Writing back to the plant

OpcUaSource is read-only on purpose, and its capabilities say so. OPC-UA can write to a PLC, and a twin that closes the loop is genuinely valuable — but it needs interlocks, an audit trail, rate limits and the customer's explicit sign-off. Do not enable it because the demo UI has a slider.

Deployment

One Node process plus a static bundle:

npm run build        # web/dist
node server/index.js # serves the API and the telemetry socket

Runs on a plant-floor box or an industrial PC. With Ollama or the rule-based copilot it needs no internet at all — which for most plant networks is not a preference but a requirement.

Honest limits of the demo

Worth saying out loud, because the engineer will work it out anyway:

  • The physics is first-order lags, a PID and accumulating wear — enough that signals move the way an engineer expects and stay coupled. It is not FEA or CFD, and it is not calibrated against a real machine.
  • The five stations are a plausible generic line, not the customer's process.
  • Prediction is trend extrapolation. There is no trained model and no failure history behind it. With real historical failure data you could do considerably better — and that is the natural next conversation, not a gap to hide.