VN Applied Labs
A harder test for AI that has to work on a real plant.
Environments, evaluations and deterministic verifiers built from synthetic plants, with abstention scored as part of the answer.

Real plants are the private exam. Synthetic plants are the classroom.
A model that answers a brownfield question has to reconcile a drawing, a scan and a site observation that disagree, and it has to know when to stop.
Every accountable engineering study we deliver keeps the path from physical state to evidence, constraint, decision and observed outcome. Applied Labs builds synthetic plants that reproduce that path without the customer's plant: a maintained reference manufacturer with two sites, task families with authored ground truth, and verifiers that grade a verdict, its cited evidence and its refusal to answer.
Customer geometry and site-identifying records never enter the classroom. The real sites tell us whether the synthetic ones are any good: a synthetic task family is calibrated against held-out real-site evidence before any training use is claimed for it.

What an AI builder receives.
Each deliverable is synthetic and authored by our engineers unless a separate rights grant says otherwise.
Industrial decision benchmark
Tasks on the synthetic reference manufacturer, scored on verdict correctness, cited evidence, discrepancy detection, evidence gaps and calibrated abstention.
The benchmark is synthetic only. It does not test real plants, process prediction or robot manipulation.
Environment packages
Task-family environments with the evidence, the ground truth and the evaluation interface for the agreed use.
Evaluation and training use are separate rights. Customer site data is not included by default.
Verifiers
Task-specific checks that inspect an agent's verdict, its evidence citations or a bounded plant artefact against authored ground truth, deterministically.
A verifier tests a defined task. It does not certify a model or authorise action on a plant.
Expert-reviewed trajectories
Engineer-authored rubrics and reviewed trajectories on synthetic decision cases, showing what the reasoning looks like when the agent has to work from physical evidence.
These are synthetic and written by our engineers, unless a rights grant permits derived customer use.
Calibration reports
The measured gap between a synthetic task family and the real decision pattern, on the held-out private exam.
A calibration report is required before any training-use claim. No accuracy result is implied here.
Sim-ready plant stages
A synthetic, task-scoped stage export for an agreed robotics or physical-AI evaluation, in an open format.
The stage is an environment artefact. We train no robot policies and claim no autonomous operation.
What the agent has to do.
- Reconcile conflicting evidence
The drawing, the scan and the site observation disagree. Which one is current, and which discrepancy matters for this decision.
- Answer a bounded feasibility question
Can the named equipment go in this bay with the required footprint, clearance, access and utilities? The answer must cite the evidence it rests on.
- Abstain when it should
Is the evidence sufficient? If not, the right answer is Not Demonstrated or Escalation Required, and the test scores it as such.
Physics task families follow the facility-physics build: solver-generated thermal and airflow cases with declared boundary conditions, scored on the prediction, the uncertainty declared and the refusal to answer outside the validated domain.
The physics families are planned. They follow the first facility-physics domain, and no result is implied for them here.
The boundaries of the test.
What do the tasks evaluate?
Reconciliation of conflicting records, bounded feasibility checks and appropriate abstention. A correct verdict must be supported by the right evidence, and recognising an unanswerable question is part of the score.
Can customer records enter a training set?
Only with negotiated permission for that use. Customer geometry and site-identifying records do not enter the synthetic classroom by default, and nothing trains a supplier's model except under a written mirror of the same terms.
How is synthetic performance checked?
Calibration compares a task family against held-out real-site evidence. A synthetic score on its own is not a claim about performance in a real plant, and we do not present it as one.
What comes next?
A constraint-based plant generator, so synthetic classrooms can be produced at volume, and physics task families with solver ground truth once the first facility-physics domain exists.
Tell us what the model has to do.
For frontier and physical-AI teams that need a test built from how an engineer actually decides on a plant.