Errors have physical cost
A wrong answer can scrap a batch, miss a failing asset or misdirect a technician. Error analysis has to be scenario-specific, not an average score.
A demo shows that a system can work once. Validation shows how it behaves across real operating conditions, where it fails, and who is accountable when it does. This page describes the method NeoBram applies before an industrial AI system is trusted with real work. It is an educational description of our approach, not a certification claim.
Why industrial validation is different
A wrong answer can scrap a batch, miss a failing asset or misdirect a technician. Error analysis has to be scenario-specific, not an average score.
Products, shifts, seasons, sensors and documents change. A system validated on last quarter's data needs monitoring, not trust.
Quality, safety and regulatory owners keep authority. Validation produces the evidence they need to accept or reject the system.
Validation lifecycle
Intended use, decision owner, baseline metric, failure conditions and acceptance criteria are written down before any model work.
A held-out, representative test set measures the candidate system against the agreed metrics, including known edge cases.
The system runs on one real workflow with named reviewers, an approval step and a safe fallback.
Live behaviour, drift, override rates and user acceptance are measured against the offline results.
The client's process, quality and safety owners accept or reject the system against the written criteria.
Drift, incidents, overrides and business KPIs are watched, with rollback and retraining paths agreed in advance.
A gate can send the project backwards or stop it. A validation process that can only conclude “proceed” is not a validation process.
What we evaluate
Every engagement defines pass/fail thresholds per scenario before implementation. A system without written acceptance criteria cannot be validated, only demonstrated.
Held-out test sets that reflect real shifts, products, asset states and edge cases. More data does not fix an unrepresentative test set.
Once live, the same metrics are measured on production traffic. Divergence between offline and online results is investigated, not explained away.
Both error directions are counted separately, because their operational costs differ. Alarm fatigue and missed events are tracked as first-class metrics.
Retrieval systems are judged on whether answers trace to a controlled source and version, and whether the system declines when evidence is missing - not on fluency.
Catch rate and false-reject rate are measured against a client-approved defect library under production lighting, speed and variation.
Agents are tested for tool-use permission boundaries, unsafe action refusal, escalation behaviour and repeatability - not only task completion.
Every tool an agent can call is tested for least privilege: what it may read, what it may write, and what always requires human approval.
Input distributions, output quality and business KPIs are monitored after go-live, with agreed thresholds that trigger review or rollback.
Prompts, models, retrieval indexes and configuration are versioned so any change can be traced and reversed. Rollback is rehearsed, not assumed.
When confidence is low or the system is unavailable, the workflow falls back to the existing process. Responsible people can always override the system.
Inputs, outputs, versions, approvals and overrides are logged so that any decision the system supported can be reconstructed afterwards.
Model metrics are necessary but not sufficient. The pilot must move the agreed operational baseline - review time, escapes, downtime - or explain why not.
Who validates, who accepts, who operates, who monitors and who can stop the system is written down per engagement. NeoBram does not replace the client's quality, safety or regulatory authority.
Practical checklist
Related reading
A practical next step
If you are planning an industrial AI pilot, bring one workflow and its current baseline. We will help you define acceptance criteria and a validation plan before any build decision.
Discuss one industrial AI use case