Simulation engineer
Your work stops disappearing into folders.
You keep running the simulations you already run. The results become a company asset that others can find, trust, and build on. No new tools to learn at the bench.
Simr Data Platform · Patent pending
Your CAE and test data holds enormous value. Most of it is trapped in files no model can use. The Simr Data Platform captures, structures, and versions that data. It connects every dataset back to its source with a digital thread.
The problem
Engineering teams already generate the data that Physics AI needs. It is not in a form any model can use. Here is what stands in the way.
01
Results live on HPC scratch, file shares, and laptops. There is no single queryable record of what was run.
02
A result file rarely carries its mesh, solver, boundary conditions, and material model. Without that context it cannot train a reliable model.
03
You cannot trace a model output back to the run and inputs that produced it. That blocks trust, validation, and audit.
04
Change geometry or setup and old and new results get mixed. Reproducibility breaks. Comparisons stop being meaningful.
05
Every solver writes differently. Data scientists spend most of their time on extraction and cleanup, not on modeling.
06
You want surrogate models and copilots. You have no clean, labeled, ML-ready dataset to train them on yet.
How it works
Capture
Hooks sit on your simulation runs and test benches. Data is captured as it is produced, not months later by hand.
Structure
Raw outputs become structured records with a shared engineering vocabulary. You can search and filter across projects.
Version
Every dataset is versioned. Change a setup and the prior state stays intact and comparable.
Connect
Each dataset, input, and output links back to its origin run and test. Lineage is native, not bolted on.
Who it is for
Simulation engineer
You keep running the simulations you already run. The results become a company asset that others can find, trust, and build on. No new tools to learn at the bench.
Data scientist
You get training sets that are structured, labeled, and traceable to source. Provenance comes with the data, so you can validate a model and defend its inputs.
our faqs
What is the Simr Data Platform?
The Simr Data Platform is engineering data infrastructure for simulation (CAE) and physical test data. It converts proprietary simulation and test results into structured, versioned, queryable datasets that stay usable across engineering workflows and AI systems. Its core is a patent-pending Semantic Data Layer that turns locked, proprietary output into an open, standardized data layer.
What is the Semantic Data Layer?
The Semantic Data Layer is the foundation of the Simr Data Platform, and it is patent pending. It converts proprietary simulation and test output into an open, standardized form and attaches engineering meaning: what each result is, the inputs that produced it, and the context that says when it is valid. Because the data carries its meaning, it becomes queryable and comparable across runs, disciplines, and tools, and ready for both engineers and AI to consume.
What problem does it solve?
Engineering teams generate enormous amounts of simulation and test data, but very little of it gets reused. Most results are reviewed once, then set aside. Several distinct problems keep that data from being reused or used for AI:Proprietary formats. Solver output is written in binary formats built to be viewed one run at a time. Pulling thousands of runs into a single dataset is slow, manual, and requires specialized tools.Scattered across file shares. Results sit on HPC scratch, shared drives, and laptops, organized by folder and file names. That breaks the moment a naming convention slips, and it cannot aggregate across projects.Missing context. A result rarely travels with the inputs that produced it (geometry, mesh, boundary conditions, solver version). Without that context, it cannot be trusted, compared, or reused.No versioning. When geometry or setup changes, old and new results blur together, so reproducibility and comparison break down.Not dataset-shaped. None of it is structured as the consistent, labeled tables a machine learning model can train on.The Simr Data Platform addresses each of these. It converts proprietary output into an open, standardized, versioned data layer, keeps every result connected to its context, and turns the collection into datasets that engineers and AI systems can reuse across cycles, teams, and tools.
Who is the Simr Data Platform for?
It is for engineering teams that build complex physical products, including automotive, aerospace, semiconductor, robotics, and energy. Within those teams it serves simulation engineers, data engineers, data scientists, and engineering leaders.
What is the difference between the Simr Data Platform and the Simr Compute Platform?
The Compute Platform runs simulations in your own cloud (Simr's SimOps product). The Data Platform makes the resulting simulation and test data structured, reusable, and AI-ready. You can use either one on its own, or both together.
Is the technology patented?
The core Semantic Data Layer is patent pending.
How does the Simr Data Platform work?
It converts proprietary simulation and test output into an open, standardized data layer. Engineers then curate and validate that data, including 3D inspection, to confirm what is correct and relevant. The validated data is versioned, tracked across design cycles, and served to downstream tools and AI through an API.
How does the Simr Data Platform turn raw simulation files into a dataset?
Engineers build the pipeline themselves, in a drag-and-drop interface, with no programming required. A pipeline usually extracts data from the source, attaches engineering context, normalizes it, generates features, and assembles a dataset, but you compose and adapt those steps to fit your data and your goal. The source is a CAE run or a physical test. The output is a structured, labeled table that links each run's design inputs and conditions to its selected results. (A feature here is an engineering variable or result chosen as a column for analysis or model training.) The original solver files stay intact. The platform builds the dataset from them.
Do I have to build the pipelines myself, or can Simr help?
You can do either. The drag-and-drop interface lets your engineers build and own their pipelines. Simr also offers a white glove service, working directly with your engineers on pipeline definitions, interactive visualizations, automated reports, and AI model training. Teams often start with hands-on help from Simr and take on more themselves over time.
What does a prepared dataset look like?
Each engineering run becomes a row that links its inputs to its outcomes. For example, 500 crash simulations could become a table where each row holds the design inputs (such as rail thickness, material, and geometry), the operating condition (such as a 56 km/h frontal impact), the selected results (such as intrusion, peak acceleration, and energy absorption), and a label (such as pass or fail). The same pattern applies across disciplines and to physical test data.
How does it make simulation data usable for machine learning?
It captures each run's inputs and results together and standardizes them into consistent datasets. This produces ML-ready data, which means data that is structured, queryable, versioned, and traceable to its source. Models can train on it without manual extraction and cleanup.
What kind of data does Physics AI need to train?
Physics AI learns from the relationship between engineering inputs and outcomes, not from raw files. It needs many runs described consistently, with each run's inputs and outputs linked, its context preserved, and versioning so a change in geometry or setup does not corrupt the dataset. The Simr Data Platform assembles that structured relationship automatically.
How much simulation data do I need to start?
You can start with a single simulation result file. With one file, you develop and test your data pipeline, and the Simr Data Platform then runs that pipeline automatically on new files as they arrive. So you are not blocked waiting to accumulate a large archive. For training a model, coverage matters more than raw count. Your runs should span the geometries, loads, and conditions you care about, including edge and failure cases. Many teams also already own physical test, warranty, and teardown data that the platform can incorporate, which reduces how much new data you need to generate.
Can it combine simulation data with physical test data?
Yes. It can link a simulation prediction to the matching physical test measurement for the same product configuration, and record the difference between them. That combined record is used to train surrogate models, calibrate simulations against measured reality, detect anomalies, and build digital twins.
What is a digital thread, and does the platform provide one?
A digital thread is the connected lineage that links every dataset, input, and output back to the run or physical test that produced it, along with its geometry, solver, settings, and version. The Simr Data Platform maintains this thread natively, so any result or AI output can be traced to its source.
How is data provenance captured?
Provenance is captured automatically. As actions are taken on your data, the platform records them, and those records connect into the digital thread that traces every result back to its origin. Tracked actions include which application ran, which job produced a result, which inputs went in, which workflow executed, which design iteration it belongs to, which data pipeline processed the data, and which version of the pipeline definition was used. It can also capture actions from the tools it integrates with, so provenance is recorded across your workflow, not only inside the platform.
Do engineers stay in control of the data?
Yes. Engineers assign meaning, curate the relevant data, and validate correctness before anything moves downstream. Automation and AI handle repetitive work such as scanning, comparison, and anomaly detection, while core physics and engineering decisions remain with your engineers.
Can it track engineering KPIs across design iterations?
Yes. The platform links requirements to KPIs to domain metrics to simulation results, and tracks them across iterations. This shows how each change affects performance and surfaces regressions early.
Which simulation solvers does it support?
It works across CAE solvers rather than a single vendor's tools, and it supports both implicit and explicit solvers. Simr builds and maintains the parsers and semantics, so results become structured data the platform understands.
Which engineering disciplines does it support?
It supports a broad range of engineering disciplines, spanning the simulation types teams run on the Simr Compute Platform. These include structural and stress analysis (FEA), durability and fatigue, crash, drop, and impact, fluid dynamics (CFD), thermal and heat transfer, acoustics, noise, vibration, and harshness (NVH), electromagnetics, and radio frequency (RF), as well as coupled multiphysics problems. It handles both steady state and transient analysis, and both implicit and explicit solvers.
Which file formats and data sources can it ingest?
It ingests common simulation file formats and output from a range of test equipment. Because Simr builds and maintains the parsers, support for new formats is added by Simr, not by your team.
Does it work with Physics AI tools like Ansys SimAI, Altair PhysicsAI, and NVIDIA?
Yes. The Simr Data Platform prepares and serves data to whatever training tools you choose. It is neutral and sits upstream of the model, so it complements these tools rather than replacing them.
Does it replace my solvers or my SPDM system?
No. The Simr Data Platform works alongside your existing solvers, SPDM, and IT systems. It converts their output into an open, reusable data layer.
Where is the Simr Data Platform deployed?
It is deployed inside your own infrastructure, either in your cloud account or on-premise.
Does my proprietary data leave my environment?
No. Your simulation and test data stays inside your own infrastructure and remains under your control.
Does Simr train AI models on my data?
No. Your simulation and test data stays inside your own environment, and Simr does not move it out or train shared models on it. There is no shared foundation model for proprietary engineering data, so training happens on your own data, and any models built there remain yours.
Is my data secure?
Your data stays within your own security perimeter, under your existing controls.
How is the Simr Data Platform different from SPDM tools like Ansys Minerva, Siemens Teamcenter, and ESTECO VOLTA?
SPDM manages simulation artifacts and processes. It answers the question "where is the simulation and everything associated with it." The Simr Data Platform answers a different question: "what can we learn from this data, and how do we turn it into a reusable dataset." One way to place the layers: PLM manages the product, SPDM manages the simulation, and the Simr Data Platform operationalizes the engineering data for analytics and AI. It works alongside SPDM, not as a replacement.
How is it different from building this ourselves on a data lake?
A data lake stores files but does not interpret them. The Simr Data Platform adds the parsing, shared semantics, engineer-led curation, and versioning that make results comparable and trustworthy across solvers and disciplines. Simr builds and maintains that layer, including support for new formats.
How do I get started with the Simr Data Platform?
Contact Simr to talk with an engineer. Bring one real dataset problem from your team, and Simr will show how it maps to the platform and what a first dataset looks like.