Back to Blog

Data Collection Checklist Before Training an Equipment World Model

Not all sensor history is equal. Here is what to check before you feed data to a world model trainer.

Sophie Marchetti Data Collection

Abstract industrial data collection concept showing sensor streams feeding into a central model

A common question when a plant team decides to train a world model on their equipment is: "Do we have enough data?" The more useful question is: "Do we have the right data?" Plants with five years of historian records can still fail to have adequate training data if the right conditions were not met. Plants with eighteen months of history can have excellent training data if the data collection was thorough and the operating window was informative.

This is a practical guide to what we look for before committing to a training run. It is not exhaustive, and the priorities shift somewhat depending on the equipment type and the modeling objective. But these are the categories that have come up consistently across different plant types and historian configurations.

1. Tag selection and coverage verification

Start by identifying which tags are relevant to the modeling objective. For setpoint prediction, you need: the primary controlled variables (temperatures, pressures, flow rates, concentrations), the control variables (valve positions, setpoints, controller outputs), the key disturbance variables (feed composition proxies, ambient conditions, upstream conditions), and the equipment state indicators (speed, current, vibration if available).

For each candidate tag, verify three things before including it in the training set. First, that it is measured at the physical location you expect based on the P&ID. Historian tags are sometimes mislabeled or renamed after instrument replacements. Cross-check the tag name against the current P&ID and against recent calibration records if available. Second, that the measurement range is consistent with what you expect for that process variable. A pressure transmitter that shows values between 0 and 5 bar on a line that should run at 12 to 18 bar has probably been installed on a different line or has been misconfigured. Third, that the tag is being actively written to the historian at the expected scan rate and is not a legacy tag that stopped updating years ago.

Missing tags matter more than low-quality tags. A missing disturbance variable will cause the model to attribute the effects of that disturbance to the variables that are present, producing spurious correlations that degrade prediction accuracy for setpoint changes. If a key variable is not in the historian, the data collection plan should include adding it before training begins, not working around its absence.

2. Historical window selection

The historical window for training should cover a representative range of operating conditions, not just a long time period. Three months of data that includes several grade transitions, multiple disturbance events, and a range of production rates is generally more useful than twelve months of steady-state operation at a single grade and rate.

Check the operating logs for the candidate window to understand what events occurred. Identify the major maintenance events (shutdowns, equipment replacements, catalyst changes) and the boundaries between operating regimes (new product grades, feedstock changes, seasonal production variations). The goal is to select a window that is as diverse as possible in its representation of operating conditions while being reasonably homogeneous in terms of the physical configuration of the equipment. An equipment replacement in the middle of the window is a hard boundary: the data before and after the replacement should typically be treated as separate training contexts, not pooled.

For equipment with strong catalyst or fouling-related aging, pay attention to where you are in the current operating cycle. If you are near the end of a catalyst cycle and the model will be used at the beginning of the next one, the model trained on late-cycle data will have a different baseline than the equipment it will be predicting. Either include data from multiple cycle positions in the training window, or document clearly that the model is calibrated for a specific cycle age and will need retraining after each regeneration or replacement.

3. Control variability assessment

A world model learns the process response to control changes. If the process ran at essentially constant setpoints for the entire training window, the model has very limited information about how the process responds to interventions, and its predictions for setpoint changes will be unreliable. You need variability in the control variables during the training period.

Assess the control variability by looking at the distribution of each control variable across the training window. A setpoint that moved less than one standard deviation from its mean over the entire training period provides weak information about the process response to that setpoint. A setpoint that ranged across its full operating envelope, including several large step changes, provides strong information.

If the control variability is insufficient, there are a few options. Extending the historical window back further may capture periods with more variability. Requesting deliberate step tests on the live process is the most informative approach but has production costs. Accepting the limited variability and communicating clearly to the users of the model that predictions for large setpoint excursions are unreliable is the honest approach if the other options are not available.

4. Data resolution and sampling alignment

Check the scan rates for each tag. Different sensors in the historian may be configured at different scan intervals: fast-dynamics sensors (flow, pressure) at one-second intervals, slow-dynamics sensors (temperature, composition analyzers) at one-minute or five-minute intervals, and some batch quality measurements at even longer intervals. The world model training pipeline needs to handle this heterogeneity, and not all pipelines handle it equally well.

Check also for timestamp alignment. In some historian configurations, sensors are timestamped at the moment they are polled by the data acquisition system, not at the moment the physical measurement was made. If the polling cycle for a given scan rate takes several seconds to complete, sensors that are polled early in the cycle and sensors that are polled late will have timestamps that are offset from each other by up to the full cycle duration, even though their nominal scan rate is the same. For slow-dynamics processes this is usually not important. For fast processes where the response times are comparable to the polling offset, it can introduce apparent lead-lag relationships that are artifacts of the polling architecture, not process physics.

5. Gap handling and data continuity

Real historian data has gaps: communication failures, server maintenance, instrument problems. Check the continuity of each tag across the training window. A tag that has gaps longer than a few minutes scattered throughout the record is workable with appropriate interpolation or exclusion. A tag that has multi-hour or multi-day gaps is potentially a problem, particularly if those gaps coincide with operationally interesting periods (startups, unusual operating conditions) that are exactly the data you want for training.

For gaps in the control variable record, the appropriate treatment is generally to exclude the adjacent period from training rather than to interpolate through it. The reason is that you need to know what the controller was doing to interpret the process variable response. Interpolated control variable values during a communication gap may misrepresent a period when the operator was actually running in manual or had made a setpoint change that was not captured. Excluding those windows is conservative but correct.

6. Process state labeling

Before training, label the major operating modes present in the training data. At minimum: normal production, startup, shutdown, grade transition, and maintenance / off-line. The model should either be trained on normal production data only, with the other modes excluded, or should explicitly include mode as a conditioning variable so the model can be queried for predictions within a specific mode.

Training on all modes pooled without labeling will produce a model that represents the average of all modes, which typically does a poor job of predicting behavior in any specific mode. The startup trajectory and the grade transition trajectory have different dynamics from steady-state production, and including them untagged in the training data contaminates the steady-state response estimates.

The labeling does not need to be exhaustive. A simple binary label (in-spec production versus other) is sufficient for a first model. More granular labels for different grade campaigns, different feedstock batches, or different operator crews can be added incrementally as the model is refined. What matters is that the initial training run is not contaminated by obvious mode confounds that will undermine the model's steady-state prediction accuracy from the start.

None of these checks requires specialized data science expertise. They are engineering questions about data provenance, physical consistency, and operating history. The process engineer who worked on the unit for the last three years will be able to answer most of them in a conversation, and that conversation is where a world model project should start, before the first historian query is written.

Build a model of your machine

Start with one production line and 90 days of historian data. Live predictions in under 48 hours from first connection.