Initial commit with Dockerfile and demo code

This commit is contained in:
2026-08-28 18:16:30 +05:30
commit cf2575fd4e
53 changed files with 11602 additions and 0 deletions
Binary file not shown.
+150
View File
@@ -0,0 +1,150 @@
# Architecture — and where your data plugs in
The one-pager for the engineer who asks *"can it take our data?"*
## Shape
```
┌──────────────────── telemetry sources ─────────────────────┐
│ │
│ SimulatedSource MqttSource OpcUaSource │
│ (working) (stub) (stub) │
│ │ │ │ │
└────────┴────────────────────┴───────────────────┴──────────┘
│
TelemetrySource ← the seam
emits "frames"
│
┌───────────────┴────────────────┐
│ │
AnalyticsEngine WebSocket /ws ──► React dashboard
· threshold alarms │ · 3D scene
· learned baselines / z-scores │ · trend charts
· least-squares trend projection │ · alarms
· alarm latching │ · what-if controls
│ │
└──────────► frame.analytics ─────┘
│
Copilot ──► OpenRouter / Ollama / rule engine
(context = analytics conclusions, not raw floats)
```
Everything above the seam consumes **frames** and has no idea where they came
from. That is the whole design: swapping the simulator for your plant means
implementing one class.
## The seam
`server/ingest/source.js` defines `TelemetrySource`. A source does three things:
1. **`start()`** — connect, subscribe, begin producing.
2. **`emit(frame)`** — hand a frame upward. Called on your cadence.
3. **`capabilities`** — declare what it supports (`timeControl`,
`faultInjection`, `setpointControl`). The UI reads this and hides controls a
read-only source cannot honour, rather than showing buttons that do nothing.
A frame is station states, signal values, WIP levels, KPI rollups and the
analytics block. `server/ingest/simulatedSource.js` is the reference
implementation; read it alongside the stub you are filling in.
## Connecting your plant
Both adapters are written out as documented stubs with the real interface, the tag
map, and the traps called out. They throw an honest "not configured" error rather
than pretending to work.
### MQTT — `server/ingest/mqttSource.js`
```
npm i mqtt
```
Fill `TAG_MAP` with your topics, implement `start()`, set `TELEMETRY_SOURCE=mqtt`
and `MQTT_URL`. If your broker speaks **Sparkplug B** (most industrial ones do),
add `sparkplug-payload` and map metric aliases from NBIRTH/NDATA instead of topic
strings.
### OPC-UA — `server/ingest/opcuaSource.js`
```
npm i node-opcua
```
Browse your server's address space, fill `NODE_MAP` with real NodeIds (never guess
them — they are namespace-qualified and installation-specific), implement
`start()`, set `TELEMETRY_SOURCE=opcua` and `OPCUA_ENDPOINT`.
### Historian / CSV
Simplest of the three and often the best first step: read the export, sort by
timestamp, and emit frames on a timer with a scrub control. You get real customer
data in the twin without touching plant networks or waiting on IT.
## Four things that bite when the data is real
These are called out in the stubs too, because each one is a way a twin starts
lying:
**Staleness.** Track last-seen per tag and mark a station offline when its tags go
quiet. A stale value rendered as live is worse than a gap. The demo already models
this — the `INS-04 sensor dropout` fault holds the last value and flags the
station rather than drawing zeros, and the charts stop rather than continuing a
flat line that would read as a healthy steady state.
**Status codes.** OPC-UA gives you `StatusCode` per value. A Bad or Uncertain
reading must not become a data point.
**Timestamps.** Prefer the *source* timestamp — when the PLC sampled the sensor —
over the server timestamp, which is when the message reached you. Under load they
diverge, and trend slopes computed on arrival times are wrong.
**Units.** Vibration in in/s, temperature in °F, pressure in psi are all common.
Convert at the boundary, never downstream, or every threshold in the system is
quietly wrong.
## What the analytics layer needs from you
Nothing changes. `TrendTracker` and `BaselineBank` work on timestamped values
regardless of origin. Two things worth knowing:
- **Baselines are learned then frozen.** A continuously adapting baseline absorbs
a slow ramp, so the exact failure mode this is built to catch would never raise
a z-score. Freezing means "different from how this machine normally behaves".
- **Signals declare their own analytics eligibility.** In
`server/sim/stations.js`, `cumulative: true` (tool wear, part counters) and
`volatile: true` (belt speed, setpoints) exclude a signal from anomaly testing —
normal accumulation is not an anomaly, and normal state swings are not either.
Both remain covered by thresholds and trend projection. When you add your
signals, set these flags next to the definition.
## Writing back to the plant
`OpcUaSource` is read-only on purpose, and its `capabilities` say so. OPC-UA can
write to a PLC, and a twin that closes the loop is genuinely valuable — but it
needs interlocks, an audit trail, rate limits and the customer's explicit
sign-off. Do not enable it because the demo UI has a slider.
## Deployment
One Node process plus a static bundle:
```
npm run build # web/dist
node server/index.js # serves the API and the telemetry socket
```
Runs on a plant-floor box or an industrial PC. With Ollama or the rule-based
copilot it needs no internet at all — which for most plant networks is not a
preference but a requirement.
## Honest limits of the demo
Worth saying out loud, because the engineer will work it out anyway:
- The physics is first-order lags, a PID and accumulating wear — enough that
signals move the way an engineer expects and stay coupled. It is not FEA or CFD,
and it is not calibrated against a real machine.
- The five stations are a plausible generic line, not the customer's process.
- Prediction is trend extrapolation. There is no trained model and no failure
history behind it. With real historical failure data you could do considerably
better — and that is the natural next conversation, not a gap to hide.
+188
View File
@@ -0,0 +1,188 @@
# Demo script — 6 minutes
One story, told once, with a clear beginning and end: *the twin sees a bearing
failing before it fails, explains why in plain language, and lets you test a fix
before touching the plant.*
Resist the urge to show every feature. The four faults exist so you can answer
"what else can it do?" — not so you can fire all of them.
**Before you walk in:** `npm run dev`, open `http://localhost:5173`, confirm the
header shows *Telemetry live* and the copilot badge shows a provider. Leave the
clock at **1×**. Press **Reset** if anyone has been clicking around — the line
comes back pre-warmed to a realistic ~88% OEE.
---
## 0:00 — What they are looking at (45s)
> "This is a live digital twin of a five-station production line — infeed,
> machining, curing oven, vision inspection, packing. Everything on this screen is
> being computed from a running model of the line, not replayed from a video."
Orbit the 3D view once, slowly. Click **CNC-02** in the 3D scene — the detail
charts below change with it.
> "OEE is 88%. Availability is 100% — nothing has broken down this shift.
> Performance is 91%, and that's micro-stops. Quality is 97.6%. That's a well-run
> line having a normal day."
**Why this beat matters:** the baseline has to be believable before the failure
means anything. If they accept 88%, they will accept everything that follows.
## 0:45 — Show that it is a model, not a dashboard (60s)
In **What-if controls**, drag **Oven setpoint** from 305 to 330 °C. Point at the
Zone temperatures chart, then at Burner duty.
> "Watch the response. It doesn't jump — it takes about half a minute of run time
> to settle, and ten seconds in it's only 44% of the way there. And look at burner
> duty: it spikes from 68 to 85% and then settles back to 74% as the controller
> finds the new equilibrium. That's a thermal model with a real time constant and
> a controller on top. A dashboard would just redraw a number."
Measured: 305 °C → 316 at 10 s → 328 at 25 s → within 1 °C of setpoint by 30 s.
Drag it back to 305.
> "That's the point of a twin: you can ask 'what happens if' without doing it to
> the actual plant."
## 1:45 — Inject the fault and compress time (45s)
Click **Inject** on *CNC-02 bearing degradation*. Then set the clock to **60×**.
> "I've just started a spindle bearing degrading. In a real plant this plays out
> over days. I'm running the clock at 60×, so we'll watch it in about thirty
> seconds. Everything else stays physical — the clock is the only thing I sped up."
Say the 60× out loud. Point at the sim clock. Never let them think this is
real-time.
Watch the **Bearing Vibration** chart climb toward the dashed **Warn 3.5** line.
## 2:30 — The prediction (60s)
A purple **Trend** card appears above the charts, and an alarm appears in the feed.
> "There it is. Vibration is at 2.9 and rising 0.18 per minute. Nothing is over
> limit yet — but it's telling me I hit the 4.5 alarm limit in about 27 minutes of
> run time."
If asked how — and someone always asks:
> "It's a least-squares fit over the last twenty minutes, extrapolated to the
> limit, and it only reports when the fit is good enough to mean anything — that
> r² of 0.65 is the fit quality. It's a regression, not a black box. I can show
> you the arithmetic."
**Do not call this AI.** Calling a straight-line fit "AI prediction" is the fastest
way to lose the engineer in the room.
## 3:30 — The copilot (90s)
Let it cross 4.5. OEE is now falling visibly. Click **Explain this** on the
bearing alarm.
While it thinks:
> "This is reading the live telemetry — states, readings against limits, the OEE
> breakdown, the measured trends. What it is *not* told is which fault I injected.
> It has to work that out from the data, same as your engineer would."
Read the answer out. It will trace vibration → tool wear → rejects → OEE.
> "Notice it separates what's measured from what's inferred, and it traced the
> whole chain: the vibration is raising tool wear, which is pushing parts out of
> tolerance, which is why the reject rate went from 2% to 5% — and that's what's
> dragging OEE down, not the machine stopping."
Then click **Draft a maintenance work order**.
> "And it writes the work order. That goes to your CMMS."
**If the cloud model is slow:** switch the selector to **Local** before the
meeting. See *Provider choice* below — a 20-second silence kills this beat.
## 5:00 — The intervention (45s)
> "So: I know the bearing is going, I know roughly when, and I know what it's
> costing me. Let me act on it."
Click **Clear** on the bearing fault, then **Tool change**. Set the clock to
**20×**.
> "Bearing replaced, fresh tooling. Watch the reject rate come back down and OEE
> recover. The twin closed the loop — it found it, explained it, and confirmed the
> fix worked."
## 5:45 — Land it (30s)
> "Two things to take away. This is running on a simulated line today, but the
> ingest layer is built to take your data — MQTT, OPC-UA, or a historian export —
> and nothing above that layer changes when we swap it. Second, the reasoning is
> reproducible: I can turn the language model off entirely and the same diagnosis
> comes out of a rule engine."
Switch the copilot to **Offline rules** and ask *"Why is OEE down?"* — it answers
instantly.
> "Same conclusion, no model, no network. The AI makes it conversational. It isn't
> where the analysis comes from."
That last move is worth more than it looks: it is the answer to "is the AI just
making this up?", and it lands better as a demonstration than as a claim.
---
## Provider choice — decide before you present
Open **AI ⚙** in the header, pick a model, and hit **Test**. It reports
time-to-first-token and tells you whether that model is presentable live. Do this
on the demo machine, on the demo network, before the meeting — cloud latency
measured anywhere between 2 s and 36 s on consecutive identical calls.
| Setting | First token | Use it when |
|---|---|---|
| **Cloud** (`stealth/ox-alpha`) | 2–35 s, variable | Best analysis. Good for the follow-up conversation, risky for the live beat. |
| **Local** (Ollama) | ~1–2 s | The live walkthrough. Needs `ollama serve` running. |
| **Offline rules** | instant | No network at all. Your safety net, and the credibility proof at the end. |
The copilot panel shows which one answered every message, so you always know. The
selector there overrides the default per question, so you can run the walkthrough
on Local and switch to Cloud for a deeper answer when someone pushes.
Your choice is saved, so set it once while preparing and it will still be there.
## Questions you will get
**"Is this our real data?"** No — it is a simulated line. The adapters for your
data are in `server/ingest/`, and `docs/architecture.md` shows exactly where they
plug in. Nothing above that layer changes.
**"How does it predict the failure?"** Least-squares fit over recent run time,
extrapolated to the alarm threshold, suppressed when r² is too low to trust. Open
`server/analytics/trend.js` if they want to see it.
**"Could it control the line?"** Technically yes over OPC-UA; deliberately
read-only here. Writing setpoints to a PLC needs interlocks, an audit trail and
their sign-off. Say that — it builds more trust than saying yes.
**"Why is OEE only 88% at the start?"** Because a line that reads 100% is a line
nobody believes. Micro-stops and unplanned stops are modelled, so the number moves
for reasons you can name.
**"Can it run on the plant floor / offline?"** Yes. It is one Node process and a
static front end, and with the local model or the rule engine it needs no internet
at all.
## If something goes wrong
- **Charts empty / "Disconnected"** — the server died. `npm run dev` again; the
dashboard reconnects on its own and the server comes back pre-warmed.
- **Copilot returns nothing** — switch to **Offline rules** and carry on. This is
why that mode exists.
- **The line is in a strange state** — press **Reset**. It clears faults, learned
baselines and history, and re-warms to a clean ~88% OEE in under a second.
- **The 3D is choppy** — drop the clock to 20×. The model is fine; it is the
render loop competing with the projector.
+289
View File
@@ -0,0 +1,289 @@
# Installation and running the demo
Everything needed to get the LINE-1 Digital Twin running on a laptop and present
it, including the copilot and the offline fallback.
If you only want the run-of-show — what to click and what to say — that is
[demo-script.md](demo-script.md) and the
[slide deck](LINE-1-Digital-Twin-Demo.pptx). This file is setup and operation.
---
## 1. Requirements
| | |
|---|---|
| **Node.js** | 20 or newer. Built and tested on **24.4.1**. |
| **npm** | Ships with Node. Tested on **11.13.0**. |
| **Disk** | ~250 MB for `node_modules`, plus models if you use Ollama. |
| **GPU** | Not required. The 3D view is primitives and runs on integrated graphics. |
| **Internet** | Only for `npm install` and the cloud copilot. The demo itself runs fully offline. |
Optional, for the fast local copilot:
| | |
|---|---|
| **Ollama** | Any recent version. Needed only for the **Local** copilot mode. |
Check what you have:
```bash
node -v && npm -v
```
## 2. Install
From the repository root — **not** from `server/` or `web/`. This is an npm
workspaces repo, and one install at the root covers both packages.
```bash
npm install
```
Expect ~240 packages and about 25 seconds. There should be no build step and no
native compilation.
## 3. Configure the copilot (optional)
**The demo works with no configuration at all.** With nothing set, the copilot
falls back to a deterministic rule engine that answers instantly, offline, and
never fails. Skip this section entirely if you just want to see the twin.
Copy the template and edit it:
```bash
cp .env.example .env
```
The server resolves a provider at startup in this order:
1. **OpenRouter** if `OPENROUTER_API_KEY` is set
2. **Ollama** if it answers at `OLLAMA_HOST`
3. **Rule engine** otherwise
```bash
# Cloud — best analysis, slowest and most variable to first token
OPENROUTER_API_KEY=sk-or-v1-...
OPENROUTER_MODEL=stealth/ox-alpha
# Local — fast, needs `ollama serve` running
OLLAMA_HOST=http://127.0.0.1:11434
OLLAMA_MODEL=qwen3.5-4b-32k:latest
# Force the rule engine, for rehearsing the worst case
# COPILOT_PROVIDER=fallback
```
`.env` is gitignored. Keep the key out of commits and out of screen shares.
**You do not have to edit this file to change models.** Once running, the in-app
settings page (**AI ⚙** in the header) lists every model each provider offers and
switches at runtime. See §7.
### If you are using Ollama
```bash
ollama serve
ollama pull qwen3.5-4b-32k
```
One gotcha worth knowing: Ollama commonly sets `OLLAMA_HOST=0.0.0.0` machine-wide.
That is a *bind* address, not a destination — you cannot connect to it. The server
normalizes it to `127.0.0.1` automatically, and the settings page shows both the
configured value and the address actually being dialled.
## 4. Run it
### For the demo (recommended)
```bash
npm run dev
```
Starts the API on **:8787** and the Vite dev server on **:5173**, prefixing each
process's output. One Ctrl+C stops both.
Open **http://localhost:5173**.
You should see, within a couple of seconds:
- Header: **Telemetry live** with a green dot, and a source badge reading `simulated`
- Header: an **AI ⚙** button naming the active provider
- **OEE around 87–90%** — not 100%
- All five stations green, **Alarms: all clear**
- Trend charts already populated with history
The line starts **pre-warmed**: 30 simulated minutes are run before the first
frame is served, so the rolling OEE window is full, the anomaly baselines are
learned, and the charts have history. You are never presenting a dashboard that
just booted.
### Single port, no dev server
For a plant-floor box, a clean rehearsal, or handing someone a link:
```bash
npm run serve
```
Builds the front end and serves everything from **http://localhost:8787** — one
Node process, one port, no Vite. Use `npm run build` and `npm start` separately if
you prefer.
Once `web/dist` exists, the API process serves it too, so `npm run dev` will log
`serving built front end`. That is harmless — but :8787 then shows whatever was
last built, which can be stale. In development always use **:5173**, which is
served live by Vite.
### Other scripts
| Command | What it does |
|---|---|
| `npm run dev` | API + dev server together (the normal way) |
| `npm run dev:server` | API only, on :8787 |
| `npm run dev:web` | Front end only, on :5173 (needs the API running) |
| `npm run build` | Build the front end to `web/dist` |
| `npm start` | Serve API, and `web/dist` if it has been built |
| `npm run serve` | `build` then `start` |
| `npm run simcheck` | Verify the simulation model (no browser needed) |
| `npm run apicheck` | Verify the server API (server must be running) |
## 5. Verify the install
Two suites. Neither needs a browser, and both are worth running once after
install so you know the machine is good.
```bash
npm run simcheck
```
Runs the model headless and asserts what the demo actually claims: signals stay
inside their physical ranges, faults move OEE the right way in a controlled A/B
against an unfaulted line, a packer stoppage propagates upstream as blocking, a
dropped sensor goes stale rather than to zero, setpoint changes are lagged, runs
are reproducible from a seed, and the prediction arrives *before* the threshold is
crossed. Ends with `ALL CHECKS PASSED`.
```bash
npm run apicheck # in a second terminal, with the server running
```
Exercises the WebSocket stream, the control API, fault propagation into the
stream, and the copilot in every provider mode — checking answers are grounded in
real readings rather than generic prose. Ends with `ALL API CHECKS PASSED`.
## 6. Running the demo
The full run-of-show with timings and wording is in
[demo-script.md](demo-script.md). This is the mechanical sequence.
### Ten minutes before
1. `npm run dev`, open **http://localhost:5173**.
2. Confirm the header shows **Telemetry live**.
3. **Pick your copilot and measure it.** Open **AI ⚙**, choose a provider, click
**Test**. Under 4 s to first token is presentable live; 4–12 s means a
noticeable pause; above that, keep it for follow-up questions only. Do this on
the demo machine, on the demo network — cloud latency has measured anywhere
between 2 s and 36 s on consecutive identical calls.
4. Press **Reset** if anyone has been clicking. The line re-warms to ~87% OEE,
all clear, in well under a second.
5. Leave the clock at **1×**. You speed it up during the demo, not before.
6. Rehearse once with **wifi off** so you know what the fallback looks like.
### The sequence
| Step | Action | What to watch |
|---|---|---|
| **1** | Orbit the 3D view. Click **CNC-02**. | Detail charts follow the selection. OEE ~87%. |
| **2** | What-if → drag **Oven setpoint** 305 → 330 °C | Zone 2 lags: 316 °C at 10 s, settles ~30 s. Burner duty spikes 68→85→74%. Drag back to 305. |
| **3** | What-if → **Inject** *CNC-02 bearing degradation*. Set clock **60×**. | Bearing vibration begins climbing toward the dashed Warn 3.5 line. |
| **4** | Wait ~10 simulated minutes | A purple **Trend** card appears: reaches the 4.5 limit in ~40 min of run time, while vibration is still ~2.8. |
| **5** | Click **Explain this** on the bearing alarm, then the **work order** chip | The copilot traces vibration → tool wear → rejects → OEE. It is never told which fault you injected. |
| **6** | **Clear** the fault, click **Tool change**, clock to **20×** | Reject rate falls, OEE recovers. |
| **7** | Copilot selector → **Offline rules**, ask *"Why is OEE down?"* | Same diagnosis in ~40 ms, no model, no network. |
### Resetting between runs
**Reset** in the header clears injected faults, learned baselines and client-side
history, then re-warms the line. Use it between back-to-back demos rather than
restarting the server.
## 7. Switching models at runtime
**AI ⚙** in the header opens the settings page. The dashboard keeps streaming
behind it, so nothing is interrupted.
- **Provider** — Auto / Cloud / Local / Offline rules
- **Model** — every model each provider offers. OpenRouter's catalogue arrives
live with context length and price per million; the Ollama list comes from your
local library with parameter size and quantization. Filter, or type an id.
- **Test** — measures time to first token and grades it
- **Generation** — temperature, max output tokens, reasoning effort
- **Ollama host** — shows the configured value and the address actually dialled
Choices persist to `.copilot-settings.json` (gitignored), so a model picked while
preparing survives a restart. **Reset to defaults** discards that file and falls
back to `.env`. Settings never contain the API key.
Note on reasoning models: reasoning tokens count against the output budget on most
providers, so a reasoning model on a small budget can spend the lot thinking and
return nothing. Keep effort **low** and the budget generous. Ollama thinking
models are sent `think: false` for the same reason.
## 8. Troubleshooting
**`Disconnected` in the header, charts empty.**
The API died. Restart with `npm run dev`; the dashboard reconnects on its own with
backoff and the server comes back pre-warmed. No page reload needed.
**Port already in use.**
Set `PORT` in `.env` for the API. For the dev server, change `server.port` in
`web/vite.config.js` — and its proxy target if you moved the API.
**Copilot returns nothing, or the panel shows a note about an empty answer.**
Expected behaviour, not a break: it fell back to the rule engine and still
answered. Usually a reasoning model exhausting its token budget. Raise max output
tokens or lower reasoning effort in the settings page, or switch to **Local**.
**Copilot says Ollama is unreachable.**
`ollama serve` is not running, or the host is wrong. The settings page shows the
address being dialled — check that first. `0.0.0.0` in your environment is normal
and handled.
**OEE reads 100% and there are no alarms.**
The pre-warm did not run. Check the startup log for `pre-warmed 30 simulated
minutes`. If it is absent, `PREWARM_SEC` has been overridden in
`server/ingest/simulatedSource.js`.
**Server exits immediately with a long message about MQTT or OPC-UA.**
`TELEMETRY_SOURCE` is set to a stub adapter. Set `TELEMETRY_SOURCE=simulated` in
`.env`. The adapters are documented stubs — see [architecture.md](architecture.md).
**3D view is choppy.**
Drop the clock to 20×. The model is fine; the render loop is competing with the
projector. The scene loads no external assets, so this is never a network issue.
**The line is in a strange state.**
Press **Reset**.
## 9. What is where
```
server/
sim/ line.js (part flow, WIP, blocking) · stations.js (physics)
faults.js (injectable profiles) · kpi.js (OEE)
analytics/ trend.js (least squares) · anomaly.js (frozen baselines)
alarms.js (thresholds, latching, predictions)
ingest/ source.js (the seam) · simulatedSource.js
mqttSource.js, opcuaSource.js (documented stubs)
ai/ copilot.js · context.js · rules.js · settings.js
web/src/
three/ Scene.jsx, Station.jsx, Conveyor.jsx
panels/ KPIs, station strip, charts, alarms, what-if, copilot, settings
scripts/ dev.mjs · simcheck.mjs · apicheck.mjs
docs/ demo-script.md · architecture.md · installation.md · the deck
```
Connecting real plant data is [architecture.md](architecture.md) — the ingest
seam, both adapter stubs, and the four things that bite when the data is real.