• Skip to main content
  • Skip to search
  • Skip to footer
Cadence Home
  • This search text may be transcribed, used, stored, or accessed by our third-party service providers per our Cookie Policy and Privacy Policy.

  1. Blogs
  2. Physical Systems Simulation (CAE)
  3. How Many Design Experiments Does a Machine Learning Predictor…
AnneMarie CFD
AnneMarie CFD

Community Member

Blog Activity
Options
  • Subscribe by email
  • More
  • Cancel
CDNS - RequestDemo

Discover what makes Cadence a Great Place to Work

Learn About

How Many Design Experiments Does a Machine Learning Predictor Actually Need?

1 Sep 2026 • 6 minute read

Machine learning prediction in CAE has a clear proposition: Train a predictor on solved design experiments, then evaluate new designs in seconds instead of hours.

The proposition is only as good as the training dataset, and the dataset is the expensive part. Every design experiment is a full, explicit crash run. So the first real engineering decision in an ML workflow is how many experiments to commit to, and it is usually made on instinct.

Four crash studies show the number of runs an ML predictor needs has nothing to do with model size or design variable count, and everything to do with how smooth the response is.

Four parametric crash and safety studies built on the ANSA Optimization tool

Damaged battery cells in a 60 km/h side impact against a rigid pillar

Figure 1: Damaged battery cells in a 60 km/h side impact against a rigid pillar. The internal short criterion triggers above a stress threshold, making the damaged cell count a discontinuous response, harder to predict than mass.

The Pattern That Is Not There

Dataset size does not track the number of design variables. The EV rocker study had the simplest parameterization in the group, just two plate thicknesses and a plate location, and it needed four times the experiments of the 25-run occupant safety study, which varied restraint system parameters, including dummy to seatbelt friction, seatbelt sensor time, and slipring position.

It does not track model size or load case severity either.

The 25-run occupant safety sled test. A THOR-50M ATD restrained by seatbelt and airbag.

Figure 2: The 25-run occupant safety sled test. A THOR-50M ATD parameterized on restraint variables, not geometry: slipring position, dummy to seatbelt friction, airbag venting, and sensor trigger times. Left, slipring position. Right, airbag interacting with the occupant.

What It Does Track

The EV rocker study isolates the real variable cleanly, because it trained predictors for two different responses from the same 100 runs and the same three design variables. Everything is controlled except the response.

Rocker mass returned a predictive power score of 0.98. Number of damaged battery cells returned 0.64. The learning curve in that study, plotting training dataset size against accuracy, showed mass converging on far fewer experiments than damaged cells.

Figure 3: Dataset size vs prediction accuracy. Mass reaches usable accuracy on a fraction of the runs the damaged cell count needs, from the same dataset and design variables.

The reason is visible in the physics. Mass is a smooth, near-monotonic function of plate thickness. The number of damaged cells is an integer count produced by a threshold, since an internal short is triggered when stress on a cell exceeds a defined value. A response that counts threshold crossings is discontinuous by construction. Small geometry changes either flip a cell over the threshold or they do not, and the predictor has to learn where all of those boundaries sit.

The front crash study points the same way from a different angle. Validated against the FE analysis, predicted intrusions came in between 0.9 and 12.6 percent error. Predicted accelerations at the seat fixation points were consistently worse, up to 16.8 percent. Intrusions are displacements. Accelerations are second derivatives, with far more high-frequency content. Same model, same dataset, same predictors, and the smoother response class predicts better.

Frontal offset impact in the 60-run front crash study

Figure 4: Frontal offset impact in the 60-run front crash study. Predicted intrusions fell within 0.9 to 12.6 percent of FE; seat fixation accelerations reached 16.8 percent.

The working rule: required dataset size and achievable accuracy track how smooth the response is, not how big the model is or how many design variables you have. Continuous, monotonic responses converge quickly and validate tightly. Threshold-triggered counts and oscillatory signals need substantially more data and plateau at higher error.

Rich Outputs Need Less Data Than You Would Think

There is a second finding worth planning around. In the front crash study, all 60 experiments were needed to train the key value predictors. The 2D curve predictors and the 3D field results predictors, covering displacement and plastic strain at three through-thickness integration points across every time step, were trained on 20.

Richer output from less data is unintuitive. The likely explanation is that each run contributes far more training data to a field predictor than to a scalar one. One run gives a single number for maximum toepan intrusion. The same run gives a spatially coherent field across the whole model at every time step.

The planning consequence: size your campaign for your least smooth response, not for your most demanding output format.

How to Size Before You Spend the Solver Budget

The tooling supports piloting rather than committing everything up front.

Inspect the data before you train anything. The correlation matrix and pair plots show which design variables actually move which responses. The predictive power score, which detects both linear and non-linear relationships on a 0 to 1 scale, flagged the difficulty gap between mass and damaged cells in the EV rocker study before a single predictor existed.

Correlation matrix for the EV rocker study

Figure 5: Correlation matrix for the EV rocker study. Mass is driven mainly by the two plate thicknesses, damaged cell count mostly by plate location, and thickness 2. Available before any predictor is trained.

Read the learning curve, not the accuracy figure. The relationship between dataset size and accuracy tells you whether another twenty runs will help or whether you have plateaued. That is the difference between spending solver budget and wasting it.

Gate every prediction on the predictor's own error bounds. Both validated studies confirmed predictions fell within the reported Mean Absolute Error and within the confidence bounds of each individual prediction. A prediction outside those bounds is a signal to solve, not to trust.

Use design variable importance maps to re-spend budget. Cutting low-importance variables lets you spend the same number of runs on better coverage of the ones that matter.

Design variable sensitivity for HIC_15

Figure 6: Design variable sensitivity for HIC_15. Ranking variables by influence tells you which to keep when spending a limited number of runs.

Why the Sizing Decision Matters

Between the 25-run and 100-run studies here sits a fourfold difference in explicit solver time and calendar days. That is often the difference between a study that fits inside a design cycle and one that does not.

Undersize the dataset and the predictor fails validation, which costs you the pilot and the team's confidence in the method. Oversize it and you have paid for runs the learning curve would have told you to skip.

Get it right and the payoff is concrete. In the sled front crash study, predictors used as response surface models drove a Simulated Annealing optimization through 500 iterations in a few minutes, because each iteration's injury criteria prediction took seconds rather than a full solve.

Read the full methodology: the EV rocker study documents dataset generation, predictor quality reporting, and a transfer learning result where the trained predictor was reused after the design changed.

GET THE WHITE PAPER


CDNS - RequestDemo

Have a question? Need more information?

Contact Us

© 2026 Cadence Design Systems, Inc. All Rights Reserved.

  • Terms of Use
  • Privacy
  • Cookie Policy
  • US Trademarks
  • Do Not Sell or Share My Personal Information