machine learning in engineering

Machine learning in engineering describes the use of learning models for technical development tasks. Typical applications include surrogate models for CFD and FEA, fast variant evaluation, optimization, quality prediction, sensitivity analysis and data-based model calibration. The value is especially high when classical simulations are too slow or many variants must be evaluated. Clean training data, physical understanding and validation against simulation or measurement are decisive.

active learning

Active learning is a machine learning approach in which the model actively requests the most informative new data points. In engineering, this can mean running additional simulations where model uncertainty is high or the engineering decision is especially important. This can generate data more efficiently than fixed experimental designs. Active learning is particularly interesting when individual simulations are expensive and every additional data point must be justified.

adaptive sampling

Adaptive sampling means that new simulation or test points are selected based on previous results. The model therefore does not only decide upfront, but learns during the process where additional data is most valuable. Points are often added in regions of high uncertainty, strong nonlinearity or close to an optimum. This can make surrogate models more accurate with fewer expensive CFD, FEA or test bench points.

AI-assisted product development

AI-assisted product development describes the use of AI throughout the development process from concept to validation. It can structure requirements, accelerate simulations, evaluate variants, analyse test data and prepare engineering decisions. In CAE processes, it is especially useful when large datasets from DoE, CFD, test benches or operation are available. The greatest value arises when AI is combined with clear engineering targets, domain knowledge and traceable validation.

AI-assisted simulation

AI-assisted simulation uses artificial intelligence to make simulation processes faster, more robust or easier to evaluate. Examples include automated meshing, parameter selection, surrogate models, result prediction, error detection and fast approximation of expensive CFD or FEA calculations. AI is especially valuable when many operating points, geometry variants or optimization loops must be investigated. However, AI must be checked against physical references because a fast model is not automatically a correct model.

artificial intelligence in engineering

Artificial intelligence in engineering includes methods that analyse technical data, identify patterns, automate models or support decisions. This includes machine learning, optimization, image and signal analysis, language models, knowledge systems and automated postprocessing. In product development, AI can shorten development time, make simulation data more useful and evaluate variants more systematically. AI does not replace engineering expertise, but extends the toolchain when data quality, model limitations and validation are well controlled.

automated component optimization

Automated component optimization connects parametric geometry, simulation, postprocessing and an optimization algorithm into one continuous workflow. This allows many component variants to be investigated and improved systematically. Typical objectives include lower pressure loss, better cooling, lower drag, lower mass or higher structural robustness. The workflow must generate robust geometries, detect failed simulations and validate proposed optima with high-quality reference simulations or measurement data.

automated simulation workflow

An automated simulation workflow links geometry generation, meshing, solver execution, postprocessing and data storage into a repeatable process. It reduces manual work and makes variant studies, DoE and optimization efficient. Robust interfaces, error handling, version control and traceable result data are important. Without automation, large parameter studies in CFD and CAE are often too slow and error-prone.

batch simulation

Batch simulation refers to automatically calculating many simulation cases sequentially or in parallel. Each case has defined inputs, boundary conditions and postprocessing steps. Batch simulation is important for DoE, map generation, sensitivity analyses and training data for machine learning. Stable runs, clear data structure and automatic detection of failed cases are decisive.

Bayesian optimization

Bayesian optimization uses a probabilistic surrogate model to select expensive function evaluations as efficiently as possible. It is especially interesting when every CFD, FEA or test bench run is time-consuming. The method balances searching promising regions and exploring uncertain regions. In engineering applications, Bayesian optimization can be combined with Gaussian processes, Kriging or other surrogate models.

Bayesian optimization vs. genetic algorithm

Bayesian optimization uses a probabilistic surrogate model to select new evaluation points as efficiently as possible. It is especially suitable when each simulation is expensive and only few real function evaluations are possible. A genetic algorithm works evolutionarily with population, selection, crossover and mutation. It is robust for complex, nonlinear, discrete or difficult-to-differentiate design spaces. Bayesian optimization uses model uncertainty to balance exploration and exploitation. Genetic algorithms often require more function evaluations, but are easy to apply to very different problem types. For CFD optimization with expensive runs, Bayesian optimization is often more efficient. For large, mixed or strongly constrained search spaces, a genetic algorithm can be more pragmatic, especially when surrogate models or parallelization are used.

boundary condition parameter

A boundary condition parameter describes a variable boundary condition of a simulation or data model. Examples include inlet pressure, mass flow rate, ambient temperature, vehicle speed, heat flux, rotational speed or turbulence level. In surrogate models, boundary condition parameters are important because a component can behave very differently depending on operating state. The valid range of these parameters must be clearly limited to avoid uncontrolled extrapolation.

CAD-CFD coupling

CAD-CFD coupling describes the connection between a parametric CAD model and CFD simulation. Changes to geometry parameters can automatically be transferred into new CFD variants. This is especially important for shape optimization, duct design, aerodynamics, cooling and automated variant studies. Stable CAD-CFD coupling requires clean parametrization, robust geometry preparation and reliable meshing.

CFD automation

CFD automation transfers recurring CFD processes into robust, scriptable or workflow-based procedures. Typical steps include CAD import, geometry cleanup, meshing, boundary conditions, solver start and automated postprocessing. It is decisive for DoE, optimization, surrogate modelling and fast design exploration. The challenge is keeping the process stable even when geometries change.

computation time reduction

Computation time reduction describes measures that make simulations or model evaluations faster. Examples include model simplification, surrogate models, reduced order models, parallelization, adaptive sampling or efficient meshing. The goal is not only faster computation, but the ability to evaluate more variants with sufficient accuracy. Computation time reduction is especially important when CFD or FEA is coupled with optimization and machine learning.

classification model

A classification model assigns data points to fixed classes or categories. In engineering, it can distinguish between pass and fail, stable and unstable, healthy and faulty, or surge and no surge. Classification is useful when the goal is not an exact numerical value but a technical decision. Clear classes, balanced training data and evaluation of misclassification are important.

coefficient of determination, R²

The coefficient of determination R² describes what fraction of the variance in a target variable is explained by a model. A value close to 1 indicates good agreement with reference data, while a value close to 0 indicates low explanatory power. In engineering, R² is useful for regression models and surrogate models, but it is not sufficient on its own for model assessment. A high R² can still hide poor extrapolation, local errors or physically wrong trends.

confidence interval

A confidence interval gives a range in which an estimated value lies with a defined statistical confidence. It is used to show not only a single prediction value, but also its uncertainty. In engineering models, a confidence interval can be given for pressure loss, temperature, efficiency or lifetime, for example. It is important to interpret a confidence interval statistically correctly and not assume that it automatically covers all systematic model errors.

constraint

A constraint limits the solution space of an optimization. Examples include maximum packaging space, minimum wall thickness, allowable temperature, maximum pressure loss, strength limit, manufacturing restriction or regulation. Constraints are important so that the algorithm does not find unrealistic or non-manufacturable designs. Good optimization therefore needs not only objectives, but also clear engineering limits.

cross-validation

Cross-validation is a method in which data is split multiple times into training and validation portions. This allows model quality to be assessed more robustly than with a single data split. Cross-validation is especially helpful for small datasets because all data points are used multiple times. In engineering projects, care must still be taken that similar variants or dependent measurements are not unintentionally mixed between training and validation.

data-driven modelling

Data-driven modelling builds a model mainly from available data rather than from explicitly formulated physical equations. The data can come from simulations, measurements, test benches, vehicle operation or design of experiments. Such models are often fast and can approximate complex relationships if the application range is well covered by training data. Caution is required outside the data range because data-driven models can produce physically wrong predictions there.

data-driven model vs. physics-based model

A data-driven model learns relationships directly from available data. The data can come from CFD, FEA, test benches, vehicle measurements, operation or DoE. A physics-based model, by contrast, uses conservation equations, material laws, fluid mechanics, thermodynamics or structural mechanics. Data-driven models are often fast and can approximate complex patterns when the data space is well covered. Physics-based models are usually more explainable and often more robust at new operating points if the relevant mechanisms are modelled correctly. The weakness of data-driven models is extrapolation outside the training data. The weakness of physics-based models is often higher modelling and computation effort. In product development, the best solution is often not either data-driven or physics-based, but a deliberate combination of both approaches.

data preprocessing

Data preprocessing includes all steps that make raw data usable for machine learning. This includes cleaning, scaling, normalization, outlier checks, unit alignment, handling missing values and combining different data sources. In engineering, preprocessing is especially important because simulation and measurement data often have different formats, noise levels or boundary conditions. Poor preprocessing often leads to poor models, even when the algorithm itself is powerful.

data quality

Data quality describes how reliable, consistent, complete and technically meaningful the data is. Poor data quality can result from measurement noise, faulty simulations, inconsistent boundary conditions, outliers, wrong units or incomplete metadata. For machine learning in engineering, data quality is often more important than the choice of algorithm. Only clean and representative data can produce robust engineering predictions.

decision tree

A decision tree is a model that splits data through successive if-then decisions. It is easy to understand and can be used for both regression and classification. In engineering, it is useful for making simple decision rules or influencing variables visible. Single decision trees tend to overfit, however, and are often less robust than methods such as random forest or gradient boosting.

deep learning

Deep learning refers to machine learning with deep neural networks. It is especially strong for large datasets, complex patterns and high-dimensional input data. In engineering, deep learning can be used for image processing, anomaly detection, field prediction, surrogate models and automated evaluation. Its technical value strongly depends on having enough representative data and clear validation metrics.

deep neural network

A deep neural network has multiple hidden layers and can therefore approximate very complex functions. It is especially suitable for large datasets, high-dimensional inputs, images, field data or complex surrogate models. In CAE contexts, deep networks can approximate flow fields, temperature fields or result metrics, for example. However, the effort for training, data preparation, regularization and interpretability is higher than for simple models.

design exploration

Design exploration describes the systematic investigation of many design variants. The goal is to identify relationships, trends, sensitivities and possible optima in the design space. DoE, automated simulation, surrogate models and visualization are often combined for this purpose. In engineering, design exploration helps make better decisions before expensive hardware is built or test bench time is used.

design of experiments (general/methodological definition)

Design of experiments, or DoE, is a structured planning method for simulations or experiments. The goal is to gain as much information as possible about influencing variables and target quantities with as few sample points as possible. In engineering, DoE is often used to build surrogate models, evaluate sensitivities or prepare optimizations. A good experimental design covers the relevant parameter space efficiently and avoids unnecessary computation or test bench time.

design of experiments (CAE application)

DoE is the abbreviation for design of experiments and describes the systematic selection of experiment or simulation points. In CAE processes, DoE is used to vary many parameters deliberately instead of selecting variants randomly or manually. This makes influencing variables, interactions and robust design regions visible. DoE is an important basis for surrogate models, response surfaces and automated optimization.

design of experiments vs. optimization

DoE stands for design of experiments and describes the systematic selection of sample points in the parameter space. Optimization describes the targeted search for better or optimal solutions. DoE is mainly exploratory: it reveals relationships, sensitivities, interactions and data coverage. Optimization is more goal-oriented: it tries to improve an objective function under constraints. A good DoE is often the basis for a reliable surrogate model. This surrogate model can then be used in optimization. Without DoE, an optimization may search in poorly understood regions or miss important interactions. In practice, the sequence is often: define design space, run DoE, build model, check sensitivities and then optimize.

design optimization

Design optimization improves an engineering design by deliberately changing design variables. Typical examples include duct geometries, cold plates, wing angles, intake runner lengths, component thicknesses or air guides. The goal is not only one better value, but a technically robust solution with good function, manufacturability and validation. CFD, FEA, DoE, surrogate models and optimization algorithms are often combined.

design space

The design space is the part of the parameter space containing technically allowable design variants. It considers limits from packaging, manufacturing, cost, material, regulations, function and safety. In optimization, good or optimal solutions are searched within the design space. A clearly defined design space prevents an algorithm from proposing computationally attractive but technically unrealistic variants.

design variable

A design variable is a modifiable parameter of a design or simulation model. Examples include channel width, wing angle, radiator area, intake runner length, valve timing or component thickness. Design variables are varied deliberately in DoE, optimization and surrogate modelling. They should be chosen so that they are technically controllable and have relevant influence on the target variables.

digital twin

A digital twin is a digital model of a real product, system or process that can be connected with real data. In engineering, it can combine simulation, measurement data, operational data and machine learning to predict condition, behaviour or future development. Typical applications include condition monitoring, virtual calibration, predictive maintenance and fast variant evaluation. A digital twin is reliable only when model limits, data quality and updating are well controlled.

digital twin vs. simulation model

A simulation model is a digital model that calculates a technical system under defined boundary conditions. A digital twin goes beyond this because it can be connected to real data from a specific product, system or process. A CFD, FEA or 1D model can therefore be part of a digital twin, but it is not automatically a digital twin. The digital twin can update states, evaluate operational data, consider ageing or predict future behaviour. A pure simulation model is usually more case-based and manually supplied with inputs. The digital twin requires data connection, model maintenance, plausibility checks and often automated evaluation. Its value is especially high for condition monitoring, virtual calibration, predictive maintenance and digital product development. In short: the simulation model calculates behaviour, while the digital twin connects model, real system and live data.

error metric

An error metric quantitatively describes how strongly model predictions deviate from reference values. Examples include mean absolute error, mean squared error, relative error or maximum error. The choice of error metric influences which model errors are considered critical. In engineering, the error metric should match the technical decision, such as maximum temperature deviation for safety questions or mean error for trend models.

experimental design

An experimental design defines which parameter combinations are investigated in a test or simulation series. It determines how well the parameter space is covered and what conclusions can be drawn later. A good experimental design considers target variables, parameter limits, interactions, effort and statistical significance. In simulation, an experimental design can structure CFD runs, 1D models, FEA calculations or test bench experiments.

extrapolation

Extrapolation means applying a model outside the data or parameter range on which it was built. This is especially risky in engineering because technical relationships outside the known range can become strongly nonlinear or physically different. Surrogate models and machine learning models are often much less reliable in extrapolation than within the training range. Extrapolated results should therefore always be checked using physical understanding, additional simulation or measurement data.

feature engineering

Feature engineering describes the deliberate creation, transformation and selection of input features for a machine learning model. In engineering, this can produce dimensionless numbers, ratios, combined geometry parameters or physically motivated surrogate quantities. A good feature can improve model quality more than a more complex algorithm. Feature engineering is especially valuable when engineering data is limited and physical prior knowledge can be used.

feature extraction

Feature extraction describes deriving meaningful features from raw data. From a flow field, for example, pressure loss, mass flow rate, vortex metrics, maximum temperature or integrated forces can be extracted. The goal is to condense complex data so that a model can learn relevant engineering relationships more easily. Good feature extraction combines data processing with technical domain knowledge.

fractional factorial design

A fractional factorial design investigates only a selected subset of all possible parameter combinations. This significantly reduces effort compared with a full factorial design. The trade-off is that some interactions can no longer be fully separated or can only be evaluated approximately. In engineering, a fractional factorial design is useful when many parameters must be considered but computation time or test bench time is limited.

full factorial design

A full factorial design investigates all combinations of the selected parameter levels. This allows main effects and interactions to be analysed very completely. The disadvantage is that the number of experiments or simulations grows strongly with each additional variable. Full factorial designs are therefore mainly suitable for few parameters or inexpensive calculations.

Gaussian process

A Gaussian process is a probabilistic regression model that provides predictions together with uncertainty. This is especially useful in engineering because expensive simulations can be added selectively where uncertainty is high. Gaussian processes are often used for surrogate models, Bayesian optimization and small to medium-sized datasets. For very large datasets or very high dimensionality, however, they can become computationally expensive.

Gaussian process vs. neural network

A Gaussian process is a probabilistic regression model that provides both a prediction and an uncertainty estimate. This is especially useful for expensive simulations because additional CFD or FEA points can be placed deliberately in uncertain regions. A neural network is more flexible and can learn very complex nonlinear relationships with many input variables. However, it usually requires more data and careful regularization. Gaussian processes are strong for small to medium-sized datasets and smooth response surfaces. Neural networks are strong for large datasets, field data, images or high-dimensional relationships. A Gaussian process is often easier to use for Bayesian optimization because uncertainty is directly part of the model. A neural network can be more powerful, but is often less transparent and must be validated particularly carefully.

generalization

Generalization describes a model’s ability to respond reliably to new, previously unseen data. A model generalizes well when it has not merely memorized training points, but captured the underlying relationship. In engineering, generalization is decisive because models often need to evaluate new geometries, operating points or boundary conditions. Good generalization comes from representative data, suitable model complexity and proper validation.

genetic algorithm

A genetic algorithm is an evolutionary optimization method using principles such as selection, crossover and mutation. It is suitable for nonlinear, discrete or difficult-to-differentiate problems. In engineering, it can be used for shape optimization, variant selection, parameter optimization or multi-objective optimization. The disadvantage is that many function evaluations may be required, so surrogate models or automated simulations are often helpful.

geometry parameter

A geometry parameter describes a variable geometric property of a component or system. Examples include length, diameter, angle, radius, gap size, cross-section, curvature or position. In parametric CAE processes, geometry parameters are used to automatically generate and evaluate variants. For machine learning, they are important inputs when a model should predict the effect of design changes.

global optimization

Global optimization searches for the best solution within the entire defined design space. It is important when many local optima exist or the objective function is strongly nonlinear. The effort is usually higher than for local optimization because more regions must be investigated. In engineering projects, global optimization is often used in early phases to identify good solution regions.

global optimization vs. local optimization

Global optimization searches the entire defined design space for a good or best solution. It is important when the objective function is strongly nonlinear or contains several local optima. Local optimization, by contrast, starts from a specific initial point and improves the solution in its neighbourhood. It is usually faster, but can get stuck in a local optimum. Global optimization often requires more function evaluations and therefore more computation time. Local optimization is useful for fine tuning once a promising design region is known. In engineering projects, global search is often performed first and then refined locally. This combination is especially useful when expensive simulations, surrogate models and engineering constraints come together.

gradient boosting

Gradient boosting is a method in which many weak models are gradually combined into a strong model. Each new tree corrects errors made by the previous models. In engineering, gradient boosting can provide very accurate predictions for tabular simulation and measurement data. It is powerful, but requires careful hyperparameter selection and validation to avoid overfitting.

hybrid modelling

Hybrid modelling combines physics-based models with data-driven methods. A typical example is a CFD or 1D model that is accelerated, corrected or supplemented by a machine learning model. This can combine physical plausibility with high computational speed. Hybrid models are especially interesting for surrogate models, digital twins, control, optimization and simulations with many variants.

hybrid model vs. purely data-driven model

A hybrid model combines data-driven methods with physical prior knowledge or physics-based submodels. A purely data-driven model mainly uses data without explicitly embedding physics into the model structure. Hybrid models can include CFD-derived quantities, conservation equations, dimensionless numbers or physically motivated features. This often makes predictions more plausible, especially when data is limited. Purely data-driven models can be very powerful, but require good data coverage and careful validation. For new geometries, boundary conditions or operating ranges, hybrid models are often more robust. The effort is higher, however, because domain knowledge, model architecture and data pipeline must be combined properly. For engineering applications with safety or functional relevance, a hybrid approach is often more trustworthy than a pure black-box model.

influence analysis

Influence analysis evaluates which input variables, boundary conditions or design variables change an engineering target quantity most strongly. It is closely related to sensitivity analysis, but is often used more generally for cause-effect evaluations. In engineering, it can show whether pressure loss depends more strongly on channel cross-section, roughness, volume flow or temperature, for example. A good influence analysis supports robust engineering decisions and avoids optimizing unimportant parameters.

input parameter

An input parameter is a quantity passed into a model as an input. In engineering models, these can be geometry parameters, boundary conditions, material data, operating points or control variables. Input parameters define the case for which the model calculates a prediction. For surrogate models, all relevant input parameters must be clearly defined with meaningful value ranges.

interpolation

Interpolation means that a model calculates predictions between known data points. It is usually more reliable than extrapolation when the parameter space is well covered and the relationships between sampling points are sufficiently smooth. In CAE and machine learning, interpolation is often used to evaluate maps, response surfaces or surrogate models. Validation remains important even for interpolation because strong nonlinearities or poorly distributed sampling points can cause local errors.

interpolation vs. extrapolation

Interpolation means that a model predicts between known data points. Extrapolation means that a model is applied outside the known data or parameter range. Interpolation is usually much safer when the parameter space is well covered. Extrapolation is riskier because the model has no direct data basis there. In engineering applications, new physical effects, nonlinearities or limiting cases can occur outside the known range. Surrogate models and neural networks in particular can produce convincing but wrong results during extrapolation. Interpolation still requires validation if sampling points are poorly distributed or effects are strongly nonlinear. For engineering decisions, extrapolated results should always be checked using CFD, measurement data or physical plausibility.

Kriging

Kriging is an interpolation and regression method often used in engineering surrogate models. It is closely related to Gaussian processes and can generate smooth response surfaces from limited sampling points. In engineering, Kriging is commonly used for DoE, optimization, sensitivity analysis and surrogate models of expensive simulations. Its advantage is good accuracy with few sampling points, while its disadvantages are modelling effort and sensitivity to data quality.

Latin hypercube sampling

Latin hypercube sampling is a sampling method for efficiently covering a multidimensional parameter space. Each parameter range is divided into intervals so that each interval is represented as evenly as possible. This often achieves better coverage with fewer points than purely random sampling. In engineering, Latin hypercube sampling is commonly used for DoE, surrogate models and sensitivity analyses.

Latin hypercube sampling vs. full factorial design

Latin hypercube sampling is a space-filling sampling method for multidimensional parameter ranges. It aims to cover each parameter range evenly without calculating all possible combinations. A full factorial design, by contrast, calculates all combinations of predefined parameter levels. This provides very complete information about main effects and interactions. The disadvantage of full factorial design is that the number of runs grows rapidly with many parameters. Latin hypercube sampling is therefore more efficient when many continuous design variables are considered. However, it does not provide the same direct and complete structure as a full factorial design. For expensive CFD or FEA runs, Latin hypercube sampling is often the better starting strategy, while full factorial designs are more suitable for few parameters or inexpensive models.

linear regression

Linear regression describes a target quantity as a linear combination of input variables. It is simple, transparent and easy to interpret. In engineering, it is suitable for first trend analyses, simple correlations and robust baseline models. Its limitation appears for strongly nonlinear relationships, interactions or complex physical effects.

local optimization

Local optimization improves a solution starting from one initial point and its neighbourhood. It is often faster than global optimization and can be very efficient when the starting point is already good. The disadvantage is that it can get stuck in a local optimum. In engineering, local optimization is often used for fine tuning after global search or engineering experience has identified a good design region.

machine learning vs. classical optimization

Machine learning and classical optimization solve different tasks, but are often combined in engineering. Machine learning learns relationships from data and can predict target quantities quickly. Classical optimization systematically searches for a better solution in a defined design space. An optimization can work directly with CFD or FEA, but then requires many expensive simulations. Machine learning can be used as a surrogate model to accelerate these evaluations massively. Classical optimization requires an objective function, constraints and design variables. Machine learning requires training data, validation and a clear validity range. In modern CAE workflows, machine learning often provides the fast evaluation model, while optimization provides the search strategy.

mean absolute error

The mean absolute error, or MAE, is the average of the absolute deviations between prediction and reference value. It is easy to interpret because it has the same unit as the target variable. MAE treats all errors linearly and is therefore less sensitive to individual outliers than mean squared error. In engineering models, MAE is useful when typical deviations should be evaluated understandably.

mean squared error

The mean squared error, or MSE, is the average of the squared deviations between prediction and reference value. Large errors are penalized more strongly than small errors. This is useful when outliers or large engineering deviations are especially critical. MSE is less directly interpretable than MAE because it is given in the squared unit of the target variable.

measurement data

Measurement data is data obtained from real measurements on test benches, vehicles, sensors or components. It contains real physical effects, but can also include noise, measurement errors, calibration errors and unknown boundary conditions. In engineering, measurement data is especially important for validating simulations and machine learning models. Good use of measurement data requires clean sensors, documented boundary conditions and traceable data preprocessing.

metamodel

A metamodel is a model built from the results of other models or experiments. It describes the relationship between parameters and results at a higher abstraction level. In CAE processes, a metamodel is often generated from a DoE with many CFD or FEA runs. It is used to analyse relationships faster, identify sensitivities and perform optimization more efficiently.

model accuracy

Model accuracy describes how well a model represents the relevant engineering relationships. It is not determined only by low error values, but also by physical plausibility, robustness and validity range. In engineering, model accuracy must always be checked against reference data such as measurements, high-fidelity simulations or known limiting cases. A model with good statistics can still be technically poor if it gives wrong trends outside the training range.

model order reduction

Model order reduction describes methods that reduce complex simulation models to a smaller and faster model form. The goal is to preserve the most important dynamic or physical properties. It is used when detailed models are too slow for optimization, control, real-time applications or large variant studies. The quality of model order reduction must always be checked against the original model or measurement data.

multi-objective optimization

Multi-objective optimization considers several target quantities at the same time. In engineering tasks, objectives often conflict, such as low pressure loss, high heat transfer, low mass, low cost and good manufacturability. Instead of one single best solution, a set of good compromises is often obtained. These compromises are often shown using a Pareto front so engineers can select a technically meaningful solution.

neural network

A neural network is a machine learning model made of connected computational units called neurons. It can learn complex nonlinear relationships between inputs and outputs. In engineering, it is used for surrogate models, map replacement, image evaluation, sensor signals, condition monitoring and optimization. Neural networks can be very powerful, but require sufficient data, good scaling and careful validation.

nonlinear regression

Nonlinear regression describes relationships that cannot be captured by a simple linear function. In engineering, this is often the normal case, for example for flow resistance, heat transfer, material behaviour or efficiency maps. Nonlinear models can be much more accurate, but require more data and more careful validation. It is important that the chosen model form remains technically plausible and does not merely fit training data.

objective function

The objective function describes what an optimization algorithm should improve. It can minimize pressure loss, maximize heat transfer, reduce drag coefficient or increase efficiency, for example. With multiple targets, objective functions can be weighted or formulated as a multi-objective optimization. A poorly defined objective function often leads to computationally good but technically unusable solutions.

optimization

Optimization describes the systematic search for a better solution within a defined design space. In engineering, target quantities can include drag, pressure loss, maximum temperature, mass, efficiency or cost. Optimization always requires clear constraints and technically meaningful parameter limits. An algorithm only finds useful solutions when objective function, constraints and model accuracy are properly defined.

optimization algorithm

An optimization algorithm is a method that searches the design space for better solutions. Depending on the task, it can be gradient-based, evolutionary, stochastic, Bayesian or heuristic. The choice of algorithm depends on computation time, number of variables, nonlinearity, noise and constraints. In CAE projects, the algorithm alone is often not decisive, but its coupling with simulation, automation and validation.

output variable

An output variable is the result calculated or predicted by a model. In engineering, examples include drag, pressure loss, temperature, efficiency, stress, mass flow rate or component lifetime. Output variables can be scalar metrics or complete fields. A clear definition of the output variable is important so that training, validation and engineering interpretation are unambiguous.

overfitting

Overfitting occurs when a model reproduces training data very accurately but performs poorly on new data. The model then learns noise, random effects or specific details of the training data instead of the general relationship. In surrogate models, overfitting can lead to apparently high model accuracy but poor engineering predictions. Countermeasures include validation data, cross-validation, regularization, simpler models and better data coverage.

overfitting vs. underfitting

Overfitting occurs when a model reproduces training data too precisely and learns noise or random effects. It looks very good in training, but fails on new data. Underfitting occurs when a model is too simple and cannot capture the actual relationship. In that case, both training error and test error are high. In engineering, overfitting is dangerous because a surrogate model can appear highly accurate but evaluate new designs incorrectly. Underfitting is also problematic because important nonlinearities or interactions are lost. Validation data, cross-validation, regularization and better data coverage help against overfitting. Better features, more suitable models or a more physically meaningful model structure help against underfitting.

parameter space

The parameter space includes all possible combinations of a model’s input parameters. It defines the range in which a surrogate model, DoE or optimization model is intended to be valid. The more parameters are included, the more effort is required to cover the space sufficiently. In engineering, a meaningful limitation of the parameter space is important to keep models accurate, efficient and technically useful.

parameter study

A parameter study deliberately investigates the influence of one or more parameters on an engineering result. It can be performed manually, automatically, simulation-based or measurement-based. In CAE projects, it is often used to identify trends, compare variants or check modelling assumptions. A good parameter study has clear target quantities, meaningful parameter limits and traceable evaluation.

parameter study vs. design of experiments

A parameter study investigates how one or more parameters influence a target quantity. Parameters are often varied manually or sequentially, for example one factor at a time. Design of experiments is more systematic and selects sample points statistically or space-filling. This allows DoE to cover the parameter space more efficiently and reveal interactions between parameters more effectively. A simple parameter study is easy to understand and useful for initial engineering questions. However, it quickly becomes inefficient when many input variables or nonlinear relationships are involved. DoE is better suited when surrogate models, sensitivity analyses or optimizations are to be built. In engineering, both complement each other: parameter studies explain individual effects, while DoE creates a reliable data basis for multidimensional questions.

parametric CAD model

A parametric CAD model describes geometry through variable parameters rather than a fixed single geometry. Examples include angles, radii, lengths, cross-sections, positions or gap sizes. For automated simulation and optimization, a parametric CAD model is the basis for generating variants in a controlled way. It is important that parametrization remains robust and does not create faulty geometries when parameters change.

parametric simulation

Parametric simulation systematically investigates the influence of variable model parameters on simulation results. It can vary geometry parameters, boundary conditions, material values or operating states. In engineering, it is used to identify trends, sensitivities, robust regions and optimization potential. Parametric simulation becomes especially powerful when combined with automation, DoE and automated postprocessing.

Pareto front

A Pareto front shows solutions where one objective cannot be improved without worsening at least one other objective. It is especially useful for trade-offs such as downforce vs. drag, cooling performance vs. pressure loss or mass vs. stiffness. The Pareto front makes visible which designs are technically efficient and which are clearly dominated. For decisions, it is often more valuable than a single optimization value.

Pareto front vs. objective function

An objective function describes which quantity an optimization should improve. It can minimize pressure loss, reduce temperature, increase efficiency or reduce drag, for example. A Pareto front arises in multi-objective optimization when several target quantities are considered simultaneously. It shows solutions where one objective cannot be improved without worsening another. The objective function is therefore the mathematical formulation of the optimization goal. The Pareto front is the result of an optimization with trade-offs. A single weighted objective function can hide trade-offs if weights are chosen arbitrarily. A Pareto front makes these trade-offs visible and supports engineering decisions.

physics-based modelling

Physics-based modelling describes models based on physical equations, conservation laws and known material or fluid laws. Examples include CFD, FEA, 1D engine simulation and thermal networks. The advantage is better extrapolation capability when the relevant physical mechanisms are represented correctly. The disadvantage is often higher modelling effort, longer computation time and necessary assumptions about boundary conditions, material data and model parameters.

physics-informed neural network

A physics-informed neural network, or PINN, combines neural networks with physical equations or constraints. The model learns not only from data, but is additionally guided by conservation equations, boundary conditions, initial conditions or known physical/material laws. This can produce more physically plausible solutions than a purely data-driven network when data is limited. In engineering applications, PINNs are especially interesting for fluid simulation, heat transfer, inverse problems and model reduction, but they are not automatically faster or simpler than classical simulation.

polynomial regression

Polynomial regression extends linear regression with powers and interactions of input variables. This allows curved relationships to be described more easily. It is well suited for manageable response surfaces with moderate nonlinearity. With high polynomial order or many parameters, however, it can become unstable and strongly overfit.

prediction accuracy

Prediction accuracy describes how close model predictions are to reference values. It is often evaluated using error metrics such as MAE, MSE or relative errors. In engineering applications, the decisive question is whether the accuracy is sufficient for the technical decision. A prediction does not need to be perfect, but it must be reliable enough within the relevant design space.

random forest

Random forest is a machine learning method made of many decision trees. It can robustly represent nonlinear relationships and interactions. In engineering, it is suitable for regression, classification, feature importance and fast surrogate models. Its disadvantage is that random forests do not extrapolate smoothly and are often only partly reliable outside the training range.

random forest vs. neural network

Random forest consists of many decision trees and is particularly robust for tabular engineering data. It can capture nonlinear relationships and interactions without requiring extensive preprocessing. A neural network is more flexible and can learn more complex functions, fields or high-dimensional inputs. However, it usually requires more data, more hyperparameter tuning and careful scaling. Random forest is often quick to build and provides useful information about feature importance. Neural networks can achieve better accuracy with large datasets and complex patterns. Random forest usually extrapolates poorly outside the training range and does so in a rather piecewise manner. Neural networks also extrapolate uncertainly, but often appear smoother and therefore potentially more misleading.

reduced order model

A reduced order model, or ROM, is a simplified model with significantly reduced mathematical complexity. It aims to reproduce the essential behaviour of a detailed model with much shorter computation time. In engineering, ROMs are used for fluid flow, structures, thermal systems, control and digital twins. The key is that the reduction preserves the effects relevant to the engineering question and does not merely save computation time.

regression model

A regression model predicts continuous target quantities. In engineering, these can include drag, pressure loss, temperature, efficiency, stress or lifetime. Regression models learn from data how input variables relate to a numerical target quantity. They are especially suitable for surrogates, map replacement, trend analysis and fast predictions in optimization loops.

response surface

A response surface is a mathematical surface that describes the relationship between input parameters and a target quantity. It is often built from design of experiments or simulation data. In product development, it helps identify trends, optima and sensitivities quickly. Response surfaces are easy to understand, but can become inaccurate for strongly nonlinear or high-dimensional relationships.

response surface model

A response surface model is a surrogate model that describes an engineering target quantity as a function of multiple input parameters. It is often used for DoE evaluation, optimization and variant comparison. Examples include drag over vehicle ride height and yaw angle or pressure loss over volume flow and geometry parameters. Model quality strongly depends on the selected sampling points and the chosen approximation function.

sampling strategy

A sampling strategy describes how points in the parameter space are selected. It influences whether a model covers the design space uniformly, finds local optima or improves uncertain regions selectively. Examples include random sampling, Latin hypercube sampling, full factorial designs, adaptive sampling and active learning. A good sampling strategy saves computation time and improves the quality of surrogate models and optimizations.

sensitivity analysis

Sensitivity analysis investigates how strongly input parameters influence a model’s target variables. It shows which parameters are important and which have little effect. In engineering projects, it helps prioritize design variables, reduce model complexity and better understand technical causes. Sensitivity analysis is especially valuable in combination with DoE, surrogate models and optimization.

sensitivity analysis vs. optimization

Sensitivity analysis investigates how strongly input parameters influence a target quantity. Optimization, by contrast, deliberately searches for better parameter combinations. Sensitivity analysis answers the question: what matters? Optimization answers the question: which variant is better or optimal? Sensitivity analysis helps prioritize design variables and remove unimportant parameters. This makes subsequent optimizations more efficient and robust. Optimization without sensitivity understanding can waste computation time in unimportant regions. Sensitivity analysis alone, however, does not provide an optimal solution. In good CAE processes, understanding the influencing variables comes first, followed by targeted optimization.

shape optimization

Shape optimization changes the outer or inner shape of a component without completely changing its basic topology. In CFD projects, it can optimize radii, cross-sections, inlet shapes, diffuser angles, guide vanes or cooling channel paths. The advantage is that the resulting geometries often remain closer to the existing design. Shape optimization is especially useful when function should be improved without completely rebuilding packaging, manufacturing or interfaces.

simulation automation

Simulation automation describes the automatic execution of recurring simulation steps. This includes model setup, boundary conditions, mesh generation, solver run, monitoring, result extraction and report generation. It is especially valuable when many variants or operating points must be evaluated systematically. Good automation saves time and also improves reproducibility and quality assurance.

simulation data

Simulation data is results generated from numerical models such as CFD, FEA, 1D simulation or system simulation. It can include scalar metrics, time histories, field data, force coefficients, temperatures, pressure losses or stresses. For machine learning, simulation data is attractive because it can be generated systematically and automatically over many variants. Its quality depends strongly on modelling assumptions, boundary conditions, meshing, solver quality and validation.

single-objective optimization vs. multi-objective optimization

Single-objective optimization improves one objective function, such as minimum pressure loss or minimum drag coefficient. Multi-objective optimization considers several objectives at the same time, often with conflicts between them. Typical trade-offs are cooling performance vs. pressure loss, downforce vs. drag or mass vs. stiffness. Single-objective optimization usually provides one clear computational solution if constraints are fulfilled. Multi-objective optimization, by contrast, provides a set of compromise solutions, often shown as a Pareto front. In real engineering projects, multi-objective optimization is often more realistic because technical systems rarely have only one goal. The disadvantage is more demanding evaluation and decision-making. The advantage is transparency: it shows which compromises are technically possible and which designs are clearly dominated.

support vector machine

A support vector machine is a method for classification and regression. It finds a decision boundary or regression function that describes data points with good separation or error control. With kernel functions, nonlinear relationships can also be captured. In engineering, it is useful for small to medium-sized datasets, but can become computationally expensive for large datasets.

surrogate model(general/technical definition)

A surrogate model is a fast replacement model for a slower, more expensive or more complex reference model. In engineering, it often replaces CFD, FEA, 1D simulation or test bench data in variant studies and optimization. The surrogate learns the relationship between input variables and target quantities from training data. Its validity range must be clearly defined because a surrogate model can become unreliable outside the trained data domain.

surrogate model (terminology / application examples)

An Ersatzmodell is the German term for a surrogate model. It represents technical behaviour in a simplified and fast way without rerunning the full original simulation every time. Typical applications include aerodynamic maps, pressure loss models, heat exchanger maps, component optimization and system simulation. A good surrogate model is fast, sufficiently accurate and transparent regarding training data, error and validity range.

surrogate model vs. CFD simulation

A surrogate model is a fast replacement model learned from existing CFD simulations, measurement data or combined data sources. A CFD simulation, by contrast, calculates the flow directly using physical equations, mesh, boundary conditions and numerical solvers. CFD provides local flow fields, pressure distributions, temperature fields and vortex structures, but is computationally expensive. The surrogate model provides results in seconds or milliseconds, but requires high-quality training data beforehand. It is especially useful for design exploration, optimization, sensitivity analysis and fast variant evaluation. Its disadvantage is the limited validity range: outside the trained parameter space, a surrogate model can produce wrong trends. CFD is the more reliable method for new geometries and detailed physics, while the surrogate model accelerates known problem classes. In good engineering workflows, the surrogate proposes candidates that are then validated with CFD.

surrogate model vs. reduced order model

A surrogate model usually represents the relationship between input parameters and target quantities from data. A reduced order model, or ROM, reduces the mathematical order of a detailed physical model. Surrogate models often provide scalar quantities such as pressure loss, drag coefficient, maximum temperature or efficiency. ROMs can also represent dynamic states, fields or system behaviour in reduced form. A surrogate model can be strongly data-driven and does not need to be derived directly from the original equations. A ROM often has a closer connection to the original simulation model, for example through modes, basis functions or reduced state spaces. Both approaches reduce computation time, but follow different modelling philosophies. In engineering, a surrogate model is often the pragmatic solution for fast scalar predictions, while a ROM is more suitable for dynamic system simulation and field-related model reduction.

target variable

A target variable is the quantity that a machine learning model is intended to learn or optimize. In supervised learning, it is the reference value against which the model prediction is compared. In CAE projects, target variables can include drag coefficient, maximum cell temperature, pressure loss, downforce or NOx emission. The target variable must be technically meaningful, measurable or calculable and consistently defined across all data points.

test bench data

Test bench data is measurement data generated under controlled conditions on a test bench. Examples include engine data, turbocharger maps, pump curves, radiator maps, material data or component characteristics. It is often more repeatable than road test data, but does not always represent all real installation conditions. For model calibration, validation and surrogate models, test bench data is especially valuable when boundary conditions and measurement uncertainty are well documented.

test data

Test data is used to evaluate model quality on previously unseen data. It should not be used for training or parameter selection so that the evaluation remains as independent as possible. In engineering, test data shows whether a surrogate model or regression model is reliable for new geometries, operating points or boundary conditions. A clear separation between training data, validation data and test data prevents overly optimistic results.

test dataset

A test dataset is the separate dataset used to finally evaluate a completed model. It should contain representative cases that the model has not seen during training or validation. In CAE and engineering applications, the test dataset can consist of additional CFD runs, measurements or deliberately withheld DoE points. It provides the most important evidence of whether a model works reliably outside the training data.

topology optimization

Topology optimization searches for a favourable material distribution within a defined packaging space. It is often used for structural components, but can also be relevant for flow channels, cooling structures and additive manufacturing. The result often shows organic-looking geometries that must then be cleaned up and converted into manufacturable designs. Realistic load cases, boundary conditions and manufacturing constraints are essential.

training data

Training data is the data used by a machine learning model to learn relationships. In engineering, it often comes from CFD simulations, FEA runs, test bench data, measurement series, DoE studies or operational data. The quality of the training data largely determines whether the model can make robust predictions later. Representative operating points, clean target values, consistent units and good coverage of the parameter space are important.

training dataset

A training dataset is the structured collection of training data for a machine learning model. It contains input parameters, target variables and often additional metadata such as geometry version, simulation setup, boundary conditions or measurement uncertainty. For engineering models, not only the number of data points matters, but also their distribution in the parameter space. A good training dataset is traceable, cleaned and physically plausible.

training data vs. test data

Training data is used so that a machine learning model can learn its relationships. Test data is used only after model development to independently check prediction quality. The key difference is that test data should not be known to the model during training or model selection. Otherwise, the assessment becomes too optimistic and true generalization remains unclear. In engineering projects, training data can come from CFD runs, measurement data, test bench data or DoE variants. Test data should contain representative cases that are technically relevant, not merely statistically convenient. A clean separation prevents similar variants from unintentionally appearing in both datasets. For surrogate models, a reliable test dataset is essential before the model is used in optimization or design decisions.

training effort

Training effort describes the effort required to build and train a machine learning model. It includes data generation, data cleaning, feature engineering, model selection, hyperparameter tuning, computation time and validation. In engineering, data generation is often the largest effort because high-quality CFD, FEA or test bench data is expensive. A good project therefore evaluates early whether the expected model value justifies the training effort.

uncertainty quantification

Uncertainty quantification describes the systematic assessment of how certain or uncertain a model prediction is. It considers input scatter, measurement uncertainty, model error, numerical error and limited training data, for example. In engineering, it is important because technical decisions depend not only on the mean prediction, but also on the risk of wrong predictions. Especially for surrogate models, Bayesian optimization and safety-critical designs, uncertainty should be stated explicitly.

underfitting

Underfitting occurs when a model is too simple to represent the underlying relationship. It then performs poorly on both training data and new data. In engineering, this can happen when a linear approach is used for strongly nonlinear flow, thermal or structural relationships. Countermeasures include better features, a more suitable model, more relevant input variables or a more physically meaningful model form.

validation data

Validation data is used during model development to select model variants, hyperparameters or training stopping points. It is therefore used for model selection, not for final independent evaluation. In engineering projects, validation data helps detect overfitting early and assess model quality in the relevant parameter range. The final statement about transferability should still be based on separate test data.

virtual product development

Virtual product development uses simulation, digital models, automation and data analysis to develop products before physical prototypes are built. It reduces development time, hardware cost and engineering risk when models are built reliably early on. Important elements include CAD, CFD, FEA, 1D simulation, DoE, optimization, surrogate models and test data correlation. The greatest value arises when virtual methods do not work in isolation, but are systematically coupled with testing, design engineering and decision processes.

virtual test bench

A virtual test bench is a simulation-based environment in which components or systems are tested computationally. It can prepare, supplement or partially replace real test bench experiments. In a virtual test bench, many variants, load cases and boundary conditions can be investigated faster than with hardware. Its reliability depends on how well models, maps, controls, boundary conditions and validation match the real application.

virtual test bench vs. physical test bench

A virtual test bench is a simulation-based environment in which components or systems are tested computationally. A physical test bench measures real hardware under controlled conditions. The virtual test bench is fast, variant-capable and can support early development decisions before prototypes are available. The physical test bench provides real measurement data including effects that are often only simplified in models. These include manufacturing scatter, ageing, leakage, sensors, actuator behaviour and real environmental conditions. The virtual test bench strongly depends on model quality, maps, boundary conditions and validation. The physical test bench is more expensive and slower, but indispensable for validation and calibration. The best approach is to use the virtual test bench for pre-design and variant reduction and the physical test bench for validation and final release.

XGBoost

XGBoost is a widely used and efficient implementation of gradient boosting. It is often used for structured tabular data and frequently achieves high accuracy with relatively short training time. In engineering, it is suitable for surrogate models, fault classification, quality prediction and fast variant evaluation. As with all ML models, data quality, feature selection and validation are more important than the algorithm name itself.