Skip to content

Latest commit

 

History

1,094 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation



APEX Circuit Logo

PostgreSQL      Grafana      Docker      Node.js      Python      TypeScript      Linux

Industrial Monitoring System (IMS)

High-Precision Manufacturing Telemetry & Statistical Process Control

Audience: Open-Source Community, System Evaluators, Deployment Engineers. Objective: The primary entry point to the IMS codebase, outlining capabilities, architecture, and deployment steps. Provenance: Architecture, versions and commands re-verified against the repository (main, after PR #22/#23) on 2026-09-26. Runtime evidence links carry their own capture dates.

APEX Circuit LDI NOC Banner

Typing SVG
Release License Docker Grafana Node-RED TimescaleDB
Tests K6 Synthetic Data


System Overview

IMS (Industrial Monitoring System) bridges the gap between high-precision manufacturing and enterprise IT. It is a telemetry monitoring platform built on Node-RED, TimescaleDB, and Grafana, integrating IT infrastructure metrics with OT (Operational Technology) data into a single, unified PostgreSQL-backed repository.

The Factory Floor Reality (OT): In advanced PCB manufacturing, Laser Direct Imaging (LDI) machines require zero-latency decision making. A shift in laser temperature or vacuum pressure can instantly cause registration errors, producing expensive scrap. Operators need immediate, color-coded Andon boards to stop the line when Statistical Process Control (SPC) limits (like Cpk) drop below acceptable thresholds.

The Convergence (IT/OT): IMS provides this visibility by blending traditional IT rigor with OT realities. It monitors infrastructure health (servers, network switches, ingestion latency) side-by-side with LDI machine telemetry, and its ingestion path is load-tested with K6 against a simulated fleet (100 servers by default, configurable). When an LDI alignment fails, engineers can instantly correlate it against network drops or server CPU spikes using the same single pane of glass.

The Architecture (IT): Under the hood, performance is driven by a stateful Node-RED pipeline managing async data ingestion and PgBouncer handling connection pooling. TimescaleDB performs the heavy lifting—computing rolling 3σ baselines (Z-Scores) and continuous aggregates on the fly, ensuring Grafana renders sub-second dashboards even when querying millions of historical telemetry rows.

NOC Overview
NOC Overview — Fleet Health Envelope
Engineering Drill-Down
Engineering Drill-Down — Per-Machine Diagnostics
Capacity Planning
Capacity Planning — Predictive Forecasting
LDI Manufacturing Command Center
LDI Manufacturing — Command Center
LDI Operator Andon Board
LDI Andon Board — Operator Floor View
LDI Engineering Analytics
LDI Engineering — Yield & SPC Analytics

Explore the Ecosystem: View the full 22-Dashboard Macro-to-Micro Architecture Guide for a deep dive into how IMS scales from C-Level business metrics down to sensor-level diagnostic data across 4 operational domains (Infrastructure, LDI, Drilling, VCP).



Core Capabilities

Telemetry Ingestion

Parallel Node-RED walkers utilizing sequential bulk SNMP polling and HTTP endpoints, persisting data to TimescaleDB via PgBouncer transaction pooling.

**Verified:** [nodered-ingestion-20260813.txt](docs/evidence/runtime/nodered-ingestion-20260813.txt)

Statistical Process Control

Real-time SPC metrics (Cpk) and rolling 3σ baselines (Z-Score anomaly detection) evaluated at the database level for early warning alerts.

Continuous Aggregation

1-minute, 15-minute, and 1-hour continuous rollups automatically calculated by TimescaleDB to maintain sub-second Grafana rendering times over large time ranges.

**Verified:** [cagg-policies-20260813.txt](docs/evidence/runtime/cagg-policies-20260813.txt)


Quick Start (Two Paths)

Note

Simulator Boundary: Both paths run the IMS stack locally using a built-in SNMP/HTTP data simulator (ims-snmpsim). They do not connect to real factory equipment or external network devices. The simulator generates realistic, bounded telemetry and alarm sequences for validation.

Choose your path based on your role and what you want to achieve:

Path A: The Evaluator Tour (UI & Workflow)

Designed for Managers, UI/UX Reviewers, and System Evaluators wanting to see the dashboards in action.

git clone https://github.com/PATTANAKORN025/IMS.git
cd IMS
cp .env.example .env   # then replace EVERY secret value before first start (see below)
make up                # build-flows + docker compose up -d (all 16 services, simulator included)
sleep 40 && make verify
# browse to http://localhost:3000 (nginx front door; Grafana has no published port)

Warning

.env.example values are public. Before any start outside a throw-away laptop, generate new values for every password, token and key in .env (POSTGRES_PASSWORD, GRAFANA_ADMIN_PASSWORD, GRAFANA_DB_PASSWORD, ALARM_API_DB_PASSWORD, INGEST_API_KEY, ALERT_WEBHOOK_TOKEN, NODE_RED_CREDENTIAL_SECRET, NODE_RED_ADMIN_PASSWORD_HASH, PGADMIN_DEFAULT_PASSWORD, GRAFANA_RENDERER_TOKEN). See SECURITY.md.

What to expect: A gentle simulation (~10-15 rows/min) allowing you to click through the LDI Manufacturing Command Center, view the Operator Andon Board, and see real-time Cpk capability charts. Verified: docker compose ps on 2026-08-13, archived in docs/evidence/runtime/compose-ps-20260813.txt.

Path B: The Performance Proving Ground (Stress Test)

Designed for SREs, DBAs, and Architects who want to verify the system's actual performance under extreme IT/OT loads.

git clone https://github.com/PATTANAKORN025/IMS.git
cd IMS
cp .env.example .env   # replace every secret value first
make up-prod           # base compose file + docker-compose.prod.yaml resource limits
make test-load         # k6 run tests/k6/pipeline-stress.js (needs k6 on PATH)

What to expect: K6 ramps simulated servers (default TARGET_SERVERS=100; set the environment variable to scale up) against the Node-RED ingestion endpoints, with thresholds pipeline_success rate > 95 % and e2e_duration p95 < 10 s. You can monitor ingestion latency and PgBouncer queue depths live on the IMS Meta-Monitoring dashboard.

Path C: Live Telemetry & API Integration (Code Examples)

Designed for Integration Engineers and Software Developers connecting factory equipment, third-party MES, or custom scripts.

1. Ingest LDI Machine Telemetry via HTTP POST

Push time-series manufacturing metrics directly to the Node-RED ingestion pipeline through the Nginx reverse proxy:

curl -X POST http://localhost:3000/ldi-telemetry \
  -H "Content-Type: application/json" \
  -H "X-API-Key: ${INGEST_API_KEY}" \
  -d '[{
    "time": "2026-09-28T04:00:00Z",
    "factory": "F1",
    "process": "LDI",
    "eqp_id": "LDI-01",
    "mo": "MO-001234",
    "fpn": "PN-5678",
    "layer_name": "L1",
    "resist_dosage": 45.5,
    "scale_x": 1.002,
    "scale_y": 0.998,
    "temperature": 24.5,
    "humidity": 45.0,
    "scan_speed": 120.0,
    "air_vacuum": -15.2,
    "thickness": 1.2,
    "board_no": 1,
    "total_board": 100,
    "total_time": 450.5,
    "state": true,
    "pe_1": 1.1,
    "je_1": 2.2,
    "log_id": "LOG-10001"
  }]'

2. Alarm Lifecycle Transitions (Acknowledge & Resolve)

Interact with the alarm-api service to change state on active factory alarms (public.ldi_alarm_lifecycle):

# Step 1: Acknowledge an active alarm (Transition OPEN -> ACKNOWLEDGED via Nginx proxy)
# Caller identity is resolved automatically from the Grafana session cookie
curl -X POST http://localhost:3000/alarm-api/alarms/ack \
  -H "Content-Type: application/json" \
  -H "Cookie: grafana_session=YOUR_SESSION_COOKIE" \
  -d '{
    "logdate_ms": 1790568000000,
    "logid": "LOG-10001"
  }'

# Step 2: Permanently resolve an alarm (Transition ACKNOWLEDGED -> RESOLVED)
curl -X POST http://localhost:3000/alarm-api/alarms/resolve \
  -H "Content-Type: application/json" \
  -H "Cookie: grafana_session=YOUR_SESSION_COOKIE" \
  -d '{
    "logdate_ms": 1790568000000,
    "logid": "LOG-10001",
    "resolution_note": "Replaced pneumatic filter; verified vacuum pressure within spec."
  }'

3. Generate Synthetic Stand-In Data (Drilling & VCP)

Populate the eap_backup database with realistic synthetic data to verify CNC drilling and VCP plating dashboards without proprietary factory records:

# Generate 7 days (168 hours) of synthetic operational data
node scripts/mock/eap-mock-data.js --hours=168 --apply

# Verify query resolution and panel row counts across drilling and VCP dashboards
node scripts/mock/verify-mock-dashboards.js --container=ims-timescaledb --psql-user=ims_admin

4. Query Continuous Aggregates (TimescaleDB SQL)

Sub-second analytical queries powered by Continuous Aggregates (CAGGs) over millions of historical telemetry rows:

-- Query 15-minute rollups for fleet machine performance
SELECT
  bucket AS "time",
  eqp_id AS machine_id,
  ROUND(avg_temperature::numeric, 2) AS temperature,
  ROUND(avg_scan_speed::numeric, 2) AS scan_speed
FROM public.ldi_data_15m
WHERE eqp_id = 'LDI-01'
  AND bucket > NOW() - INTERVAL '24 hours'
ORDER BY bucket ASC;
Known Limitations & Manual Configuration
  • The nginx front door listens on plain HTTP (${GRAFANA_PORT:-3000} on the host, port 80 in the container); add TLS termination for production.
  • Alertmanager delivery to LINE/Teams stays silent until LINE_CHANNEL_ACCESS_TOKEN, LINE_USER_ID and TEAMS_WEBHOOK_URL are set in .env.
  • pgadmin binds to loopback (127.0.0.1:5050) for security; the nginx front door (3000) is the only service published externally.
  • make targets use standardized portable scripts: on Windows, make verify and make backup run PowerShell natively, while Linux and macOS use POSIX shell.

Verification & Evidence

Architectural claims are backed by test scripts and dated evidence files. .github/workflows/ci.yml runs the same checks as node scripts/pre-commit.js plus gitleaks, compose validation and Prometheus linting whenever GitHub Actions is available for the repository. For load test results, visual regression evidence, and disaster recovery validations, refer to the Evidence Index.

Available Commands
Command Description
make help List all available Makefile targets
make doctor Check prerequisites (docker, compose, node)
make check Run full pre-commit verification gate (pre-commit.js)
make check-env Verify .env required keys and check against leaked secrets
make up Build flows, then start all 16 services (simulator included)
make up-prod Same, with the docker-compose.prod.yaml resource overlay
make down / make restart Stop the stack / restart node-red, grafana, alertmanager, prometheus
make logs Tail Node-RED logs
make verify Full system health check (containers, DB, pipeline, alerts)
make build-flows / make validate-flows Merge nodered_data/flows/*.json into flows.json / assert it is valid
make snapshot-flows / make deploy-flows Back up flows.json / POST the split flows to Node-RED
make test-unit All standalone unit test suites in tests/unit/
make test-load K6 pipeline stress test (TARGET_SERVERS, default 100)
make test-visual / make test-visual-ldi Playwright dashboard screenshot regression
make validate-dashboards Grep dashboard JSON for corrupted hex codes
make backup / make restore FILE=<path> Database dump / restore

The full pre-commit suite (all unit tests, the tests/lint/ linters, dashboard and flow JSON validation) runs with node scripts/pre-commit.js.


Architecture

%%{init: {"flowchart": {"nodeSpacing": 20, "rankSpacing": 32, "padding": 10, "wrappingWidth": 150, "curve": "basis"}, "sequence": {"wrap": true, "width": 170, "actorMargin": 36, "boxMargin": 8, "noteMargin": 8, "messageMargin": 30, "mirrorActors": false}, "state": {"padding": 6}, "theme": "base", "themeVariables": {"fontFamily": "Inter, Segoe UI, Helvetica, Arial, sans-serif", "fontSize": "14px", "primaryColor": "#334155", "primaryTextColor": "#ffffff", "primaryBorderColor": "#1e293b", "lineColor": "#64748b", "textColor": "#64748b", "secondaryColor": "#475569", "tertiaryColor": "#f1f5f9", "clusterBkg": "transparent", "clusterBorder": "#94a3b8", "titleColor": "#64748b", "edgeLabelBackground": "#475569", "nodeTextColor": "#ffffff", "noteBkgColor": "#fef3c7", "noteTextColor": "#1e293b", "noteBorderColor": "#d97706", "actorBkg": "#334155", "actorTextColor": "#ffffff", "actorBorder": "#1e293b", "actorLineColor": "#94a3b8", "signalColor": "#64748b", "signalTextColor": "#64748b", "labelBoxBkgColor": "#334155", "labelBoxBorderColor": "#1e293b", "labelTextColor": "#ffffff", "loopTextColor": "#64748b", "activationBkgColor": "#e2e8f0", "sequenceNumberColor": "#ffffff", "stateLabelColor": "#ffffff", "compositeBackground": "transparent", "transitionColor": "#64748b", "transitionLabelColor": "#64748b"}}}%%
flowchart TB
  accTitle: IMS system overview
  accDescr: Telemetry from SNMP devices, LDI machines and the plant EAP database flows through nginx and Node-RED into TimescaleDB, is shown in 22 Grafana dashboards, and alerts reach LINE and Teams through Node-RED.

  subgraph SRC["Sources"]
    SNMPDEV["Servers & switches<br/>SNMP v2c agents"]:::ext
    LDIM["LDI machines"]:::ext
    EAPSRC["Plant EAP database<br/>drilling & VCP"]:::ext
  end

  USERS["Operators & engineers<br/>browser"]:::actor
  PROXY["nginx :3000<br/>single front door"]:::ingress

  subgraph NR["Node-RED"]
    NR_LDI["ldi_ingestion.json<br/>API key · validate · staging"]:::flow
    NR_SIM["ldi_simulator.json<br/>ldi_alarm_simulator.json"]:::flow
    NR_SNMP["ingestion.json<br/>5 SNMP walkers → v9 parser"]:::flow
    NR_ALERT["alerting.json<br/>/alert-webhook"]:::flow
  end

  PGB["PgBouncer :5432<br/>SCRAM · transaction pool"]:::app
  TSDB[("TimescaleDB · ims<br/>hypertables + CAGGs")]:::store
  EAPDB[("TimescaleDB · eap_backup<br/>machine_event · vcp_*")]:::store

  subgraph APPS["Applications"]
    GRAF["Grafana 13<br/>22 dashboards"]:::viz
    ALARM["alarm-api<br/>ack / resolve"]:::app
    TWIN["factory-twin-3d"]:::app
  end

  subgraph MON["Monitoring & alerting"]
    PROM["Prometheus"]:::obs
    BBOX["Blackbox exporter"]:::obs
    AM["Alertmanager"]:::obs
  end
  NOTIFY["LINE · MS Teams"]:::notify

  SNMPDEV -->|"SNMP v2c · 30 s"| NR_SNMP
  LDIM -->|"POST /ldi-telemetry"| PROXY
  PROXY -->|"/ldi-telemetry · /inject"| NR_LDI
  NR_SIM -->|"127.0.0.1:1880"| NR_LDI
  NR_SNMP -->|"nodered_writer"| PGB
  NR_LDI -->|"nodered_writer"| PGB
  PGB --> TSDB
  EAPSRC -.->|"restored copy"| EAPDB
  USERS -->|"HTTP :3000"| PROXY
  PROXY --> GRAF
  PROXY -->|"auth_request"| ALARM
  PROXY -->|"auth_request"| TWIN
  TSDB --> GRAF
  EAPDB -->|"drilling-timescaledb"| GRAF
  ALARM --> PGB
  TWIN --> PGB
  NR_SNMP -->|"/metrics"| PROM
  BBOX --> PROM
  PROM --> AM
  AM -->|"Bearer token"| NR_ALERT
  GRAF -->|"alert rules · Bearer token"| NR_ALERT
  NR_ALERT --> NOTIFY

  subgraph LEGEND["Legend · arrows = data flow"]
    direction TB
    subgraph LEGEND_0[" "]
      direction LR
      LG_actor["Person"]:::actor ~~~ LG_ext["External system"]:::ext ~~~ LG_ingress["Ingress / gateway"]:::ingress ~~~ LG_flow["Node-RED flow"]:::flow ~~~ LG_app["IMS service"]:::app
    end
    subgraph LEGEND_1[" "]
      direction LR
      LG_store["Data store"]:::store ~~~ LG_viz["Grafana / UI"]:::viz ~~~ LG_obs["Monitoring"]:::obs ~~~ LG_notify["Notification"]:::notify
    end
    LEGEND_0 ~~~ LEGEND_1
  end
  NOTIFY ~~~ LEGEND
  style LEGEND fill:transparent,stroke:#94a3b8,stroke-dasharray:3 3
  style LEGEND_0 fill:transparent,stroke:transparent
  style LEGEND_1 fill:transparent,stroke:transparent
  classDef actor fill:#475569,stroke:#1e293b,color:#ffffff,stroke-width:1px
  classDef ext fill:#57534e,stroke:#292524,color:#ffffff,stroke-width:1px
  classDef ingress fill:#1d4ed8,stroke:#1e3a8a,color:#ffffff,stroke-width:1px
  classDef app fill:#0f766e,stroke:#134e4a,color:#ffffff,stroke-width:1px
  classDef flow fill:#0e7490,stroke:#164e63,color:#ffffff,stroke-width:1px
  classDef store fill:#b45309,stroke:#78350f,color:#ffffff,stroke-width:1px
  classDef viz fill:#4338ca,stroke:#312e81,color:#ffffff,stroke-width:1px
  classDef obs fill:#6d28d9,stroke:#4c1d95,color:#ffffff,stroke-width:1px
  classDef notify fill:#b91c1c,stroke:#7f1d1d,color:#ffffff,stroke-width:1px
  classDef future fill:#f8fafc,stroke:#94a3b8,color:#475569,stroke-width:1px,stroke-dasharray:4 3
Loading
Data Flow — Step by Step
  1. Collection — Every 30 seconds (Poll Fleet (30s)), Node-RED forks 4 walkers for network switches (CPU, Storage, Network, Temp) and 5 for servers (+LDI). The device registry is reloaded from public.devices every 5 minutes. LDI machines also push JSON to POST /ldi-telemetry through nginx, authenticated with INGEST_API_KEY.
  2. Walking — Sequential async bulk walks (session.subtree with maxRepetitions: 50). Single UDP socket eliminates switch-level packet drops. Circuit breaker trips after 2 failures with automatic HALF_OPEN probe.
  3. Parsing — sre_parser maintains per-device state in flow context (dev_state_<deviceId>), buffers rows in batch_buf_<deviceId>. Offline heartbeat (_walker: "offline") immediately zeros all metrics on device failure.
  4. Storage — Timer-gated independent flushing: each table type (sys/net/ldi) inserts only if its buffer has rows. Partial walker failures don't block unrelated data writes.
  5. Continuous Aggregation — TimescaleDB refresh policies run across the 7 continuous aggregates (ldi_data_1m, ldi_data_15m, ldi_data_1h, ldi_data_hourly, sys_hourly, net_hourly, ldi_hourly). Live retention (verified against running database, see docs/architecture/DATA_RETENTION.md): raw sys_metrics/net_metrics/ldi_metrics 30d, ldi_data 180d, ldi_data_1h and ldi_data_hourly 2yr (the 3 *_hourly continuous aggregates currently have no automated retention policy).
  6. Visualization — 22 dashboards in four Grafana folders: 01 Drilling (4: fleet overview, shift production, machine investigation, anomaly & root cause), 02 LDI (10: executive overview, operator Andon, factory digital twin, fleet command center, machine snapshot, engineering analytics & SPC, alarm console, alarm response MTTA/MTTR, alarm dictionary, data readiness), 03 Platform & NOC (5: NOC overview, engineering drill-down, AIOps capacity, ingestion latency, meta-monitoring), 04 VCP plating (3: fleet overview, operations console, real-time wall).
  7. Alerting — Prometheus scrapes /metrics, Alertmanager routes to LINE Messaging API + MS Teams with runbook links (real delivery requires operator-configured credentials, absent by design). Z-Score anomalies via Grafana SQL over TimescaleDB.
Dashboard Architecture

22 dashboards — 4 drilling, 10 LDI manufacturing, 5 infrastructure, 3 VCP plating (monitoring/grafana/dashboards/{drilling,manufacturing,infrastructure,vcp}/, one Grafana folder each — see Ownership for the domain boundary). The drilling and VCP dashboards read the eap_backup database through the drilling-timescaledb data source; without plant data, load synthetic data. Full table with panel counts and descriptions: Dashboard Inventory — auto-generated from the dashboard JSON itself (node scripts/generate-dashboard-inventory.js), CI-checked so it can't silently drift from the real dashboards the way a hand-typed table can.

Design System: Cyberpunk HUD — #030407 background, Tailwind palette (#22C55E Healthy, #F59E0B Warning, #EF4444 Critical, #00F2FE Info, #3B82F6 Accent — the approved tokens in GRAFANA_DESIGN_SYSTEM.md §2.1), Roboto Mono for stat values, glassmorphism panels, Grid-24 overlap-free layout.


NOC Wall-Display

Create a playlist in Dashboards → Playlists and start it from the playlist page; Grafana 13 plays every dashboard in kiosk mode. scripts/create-playlist.sh automates this but still calls Grafana's legacy id-based playlist API, so re-check it after each Grafana upgrade.

Mode URL parameters Use case
Kiosk ?kiosk Wall display — hides the navigation chrome
Kiosk + fit ?kiosk&autofitpanels Wall display — also scales panels to the screen height
Operator Andon /d/ims-ldi-operator-andon?kiosk Read-only floor board; interactive work belongs on the Alarm Console

Use kiosk: the older kiosk=tv TV mode is not one of Grafana 13's kiosk options (some dashboard links in this repo still carry it).


Tech Stack
Layer Technology Purpose
Orchestration Docker Compose 16-service stack (docker-compose.yaml) + production resource overlay
Collection Node-RED + net-snmp Sequential async bulk SNMP walks, 5-thread parallel walker
Database TimescaleDB 2.29 (PostgreSQL 16) + PgBouncer 1.25 Hypertables, CAGG rollups, native compression, retention policies
Visualization Grafana 13.1.2 + image renderer 22 dashboards (4 drilling + 10 LDI + 5 infrastructure + 3 VCP)
Alerting Prometheus + Alertmanager Metric scraping, inhibition rules, LINE Messaging API + MS Teams webhooks
Load Testing K6 Pipeline stress, thresholds success > 95 %, e2e p95 < 10 s
Services Node.js 22 (Express) alarm-api (acknowledge/resolve write path), factory-twin-3d (Floor 1 twin)
Front door nginx 1.31 Single published UI port; same-origin routing to Grafana, alarm-api, twin, Node-RED ingest
SLA Probing Blackbox Exporter HTTP/TCP/ICMP endpoint monitoring
Database Schema
  • devices — device registry, single source of truth for both SNMP-polled infra and LDI machines (device_type)
  • sys_metrics / net_metrics — infra telemetry (CPU/RAM/disk/temp, per-interface RX/TX), hypertables
  • ldi_metrics — legacy manufacturing throughput/PE/JE/humidity/power/vibration, hypertable
  • ldi_data / ldi_alarm_log — V2 normalized LDI telemetry + alarms, exact-event RCA join via related_log_id, hypertables
  • sys_hourly / net_hourly / ldi_hourly / ldi_data_1m / ldi_data_15m / ldi_data_1h / ldi_data_hourly — continuous aggregates
  • v_machine_spc_fleet / v_ldi_rca_recent_window / v_ldi_rca_truth_test — materialized views, refreshed every 60s

Exact column counts, the full view/CAGG list, and applied-migration count: Database Schema Inventory — auto-generated from information_schema + timescaledb_information.* (node scripts/generate-schema-inventory.js), CI-checked against the live database.

Project Structure
IMS/
├── docker-compose.yaml         # 16 services; docker-compose.prod.yaml adds resource limits
├── proxy/nginx.conf            # the single front door (Grafana, alarm-api, twin, LDI ingest)
├── monitoring/
│  ├── grafana/
│  │  ├── dashboards/{drilling,manufacturing,infrastructure,vcp}/  # 4 + 10 + 5 + 3 provisioned dashboards (source of truth)
│  │  ├── library-panels/        # shared library panel (Fleet Health Score)
│  │  └── provisioning/, grafana.ini
│  ├── prometheus/, alertmanager/, blackbox/, snmpsim/
├── nodered_data/
│  ├── flows/                  # 5 split flow files (source); flows.json is a build artifact
│  ├── lib/                    # circuit-breaker.js, parser.js, snmp-normalize.js, units.js
│  └── settings.js
├── postgres/init/              # first-boot bootstrap SQL + grafana password script
├── database/migrations/        # numbered forward-only migrations (max 093), applied by db-migrate
├── services/
│  ├── alarm-api/              # acknowledge/resolve write path (Express + pg)
│  └── factory-twin-3d/        # Floor 1 digital twin (Express, lib/*.js + public/ viewer)
├── tests/                      # unit/ + lint/ (no infrastructure), e2e/, smoke/, playwright/, k6/, ...
├── scripts/                    # build-flows.js, migrate-entrypoint.sh, verify-deployment.*, backup/restore, generators
├── assets/                     # README screenshots and banner
├── docs/                       # English docs: architecture/, operations/, user/, admin/, audit/, evidence/, ...
├── th/, zh-CN/                 # Thai and Simplified Chinese mirrors of the docs
└── .agents/skills/             # agent skills used by AI tooling

Documentation & Community

Executive & Business Strategy

Document Description
Business Value & ROI Executive summary, cost savings, MTTR reduction, and strategic impact
Platform Book (start here) Navigational hub for the entire documentation set, terminology glossary
Product Context Product purpose, target audience, and positioning

Manufacturing & LDI Intelligence

Document Description
Manufacturing Platform Plan Infra/manufacturing domain separation, validation/soak/DR rollout plan
Manufacturing Domain The LDI schema/dashboard pattern and onboarding flow
LDI SPC Guide Process capability (Cpk) methodology and formula
LDI RCA Guide Root-cause correlation (Lift/Confidence) methodology
LDI Validation Protocol 4-phase production sign-off procedure

Core Architecture & Security

Document Description
Architecture System context, ADRs, streaming architecture, CAGG strategy
Visual Architecture Mermaid C4 Model diagrams and sequence flows
Data Flow End-to-end pipeline diagrams, the real CAGG rollup chain
Database Schema Auto-generated table/column/view reference (CI-checked)
Security Model Trust boundaries, per-adapter authentication, and RBAC
Equipment Integration (EAP) SNMP, HTTP/JSON, and SECS/GEM adapter contracts
Ownership Domain boundaries enforced via CODEOWNERS
Design System Semantic color palette, typography, threshold contracts
Dashboard Inventory Auto-generated dashboard/panel-count table (CI-checked)
Synthetic Mock Data Stand-in data generator for drilling and VCP verification

Operations & SRE Playbooks

Document Description
User Manual Dashboard guide, metric reference, alert response playbooks
Admin Manual Container ops, device registration, migrations, backup/recovery
Operator SOP Standard Operating Procedures for factory floor / Level 1 NOC
Troubleshooting & Alarms Alarm code resolution and troubleshooting playbook
Incident Response Severity framework + real worked incident examples
Alarm Severity Guide The 4-tier severity taxonomy, ISA-18.2 scope
Backup & Restore Real dr-test.sh evidence, procedure, and caveats
DR Test Plan 3-drill disaster-recovery test plan
Data Retention Live retention/compression policy
Release Checklist What to verify before tagging a release
Troubleshooting Common issues, debugging commands, recovery procedures
Operations Runbook Day-to-day stack operations and recovery commands
Factory Twin Operator Guide How to read the Floor 1 twin and its evidence states
Production Readiness Release gate status and open risks

Community & Reference

Document Description
Video Onboarding Script Storyboard and guide for recording onboarding GIFs/Videos
Contributing Development workflow, branch naming, commit conventions
Code of Conduct Community standards and enforcement
Security Policy Vulnerability reporting
Changelog Release and merge history
Bug Report Report a bug or regression
Feature Request Suggest a new feature

Built with precision. Designed for uptime.

MIT License — 2026 IMS Contributors

About

Real-time SNMP monitoring for PCB manufacturing

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages