logo
SimReady Library
EgoSuite
RoboFinals
Lightwheel-Platform Enterprise
Customers
About
logo

The First Newton-Native Benchmark:
Running the Full Evaluation Stack

on Newton

RoboFinals runs the full evaluation stack on Newton and releases the first Newton-native benchmark, opening with 22 household, hospital and factory tasks.

Training a manipulation policy keeps getting faster. Deciding whether it is ready for real hardware deployment has not changed. Each team builds its own environment, writes its own success criteria, and ends up with benchmark scores that only hold up inside that lab. Nothing transfers, nothing compares, and the field falls back on watching recordings and running the robot again to find out what a number actually meant.

The problem is a score is only as solid as the physics that produced it, and no engine has yet earned the right to be that shared ground truth. That is what is pulling the robotics field toward Newton. But a new physics engine starts out empty: no assets tuned to its solvers, no standard tasks, no scoring protocol. It can produce a beautiful demo. It still cannot tell you whether a policy is good enough.

Today Lightwheel announces that RoboFinals, our robot policy evaluation platform built on the open-source Isaac Lab-Arena, now runs the complete evaluation stack on Newton, an open GPU-accelerated physics engine, and includes the first fully Newton-native evaluation benchmark. Every layer, including assets, solver, robots, teleoperation demonstration data, training and evaluation, is built natively on Newton and verified end to end. The initial suite covers 22 household, hospital and factory tasks.

Newton was initiated by NVIDIA, Google DeepMind, and Disney Research. Lightwheel serves on its Technical Steering Committee, setting asset standards for the engine and leading deformable-solver development. Getting a benchmark onto a new engine is not a matter of pointing an existing test suite at a new backend; every layer has to be built and verified in place.

Rebuilding the Evaluation Stack on Newton

Integrating a simulator is not hard. Getting a simulator into a complete evaluation loop is. The hard part is what happens between layers. If an asset's mass or friction is wrong, the physics built on top of it is wrong. If the physics is wrong, a policy trained on it learns the wrong thing. If the policy is scored against tasks that were never verified against real contact behavior, the score means nothing. Every layer has to carry forward what the last one got right, or the final number is fiction. Below is how each layer, from assets to evaluation, was built and verified on Newton.

Assets and environments. We built simulation-ready assets and task environments directly against the latest Newton schema, spanning rigid bodies, articulated objects, and deformables such as cables. Physical parameters come from real-world measurement and calibration rather than hand-tuning, which is what anchors everything downstream. An asset that looks right and an asset that behaves right are two different things. Every asset in the benchmark was validated through interaction inside the engine—grasped, opened, plugged, and manipulated—before it entered the task suite.

Physics and solvers. NVIDIA Warp and Newton provide fast, accurate rigid-body simulation and interaction. We extended that foundation where our task suite demanded more. We developed our own solvers for deformables including cables and cloth, and built modules for multi-physics coupling and complex collision handling, achieving stable, penetration-free interaction between grippers and deformable objects. Combined with measurement-calibrated assets, this lets us reproduce complex real-world physical behavior more faithfully and narrow the sim-to-real gap.

Robots. We adapted and validated four embodiments in Newton: X7S, Dexmate, H2 plus, and G1. These span different morphologies and actuation schemes, covering kinematics, actuation, and contact behavior in each case. A policy evaluated on this stack controls a robot that responds faithfully inside the engine, and the robot layer is not hardcoded to one platform.

Teleoperation. We built a full teleoperation link into Newton-native environments, so human operators collect demonstrations directly inside the engine rather than importing trajectories recorded elsewhere. Data collected in the same environment the policy is evaluated in removes a whole class of distribution mismatch.

Data. Using that pipeline, we collected teleoperation datasets across the full task suite, with hundreds of demonstrations per task. Every episode passes a standardized quality check before it counts as training data.

Evaluation. RoboFinals runs evaluation in Isaac Lab-Arena. Every task ships with defined success criteria, an evaluation protocol, and a scoring script. Results are comparable across policies and reproducible across runs.

22 Contact-Rich Tasks Inside the Benchmark

The initial suite covers 22 tasks across three scenario families where embodied AI meets real demand.

In factory settings: connector insertion and alignment, removing defective material from a line, picking parts in irregular poses, cable plugging with port confirmation, and tray picking.

In household settings: opening cabinet doors and drawers, retrieving items from a refrigerator, loading a dishwasher, operating a microwave door and buttons, and tidying a countertop.

In hospital settings: organizing surgical instruments in a sterile room, seating hinged instruments such as clamps and scissors into tray slots and stringers.

These are contact-rich, long-horizon manipulation tasks. We chose them because they stress exactly the physics that next-generation engines are built to get right: friction, articulation constraints, deformation, and sustained contact.

Trained on Newton. Evaluated on Newton

Building the evaluation layers is one thing. Proving they work together is another. We ran the complete RoboFinals pipeline on Newton: collect teleoperation data in native environments, train policy models on that data in Isaac Lab, then evaluate those models against the standardized RoboFinals benchmark tasks. The assets hold up under real manipulation, the environments reproduce identical behavior across runs, the datasets support real training, and the evaluation can be rerun.

Worth being precise about what this establishes. Closing the pipeline proves the stack is internally consistent and that physical meaning survives every seam. Training and evaluation both happen in simulation, so this does not by itself prove agreement with hardware. What grounds the stack against reality is upstream: assets built from physical measurement and calibration, and solvers developed to reproduce real contact and deformation behavior. Real-world correlation is the next thing we intend to publish.

From Physics Engine to Evaluation Stack

A physics engine provides the laws of motion, solved fast and accurately on the GPU. It can tell you how an object will move. It cannot tell you whether a policy is good enough. Between those two things sit assets, solvers, embodiments, data, and scoring protocols, and every seam that has to hold between them.

RoboFinals is that layer. Newton is no longer only a physics backend that can run robot tasks. It is a place where policies can be judged, and where results obtained on it share a coordinate system for the first time. From “it runs” to “it can be evaluated” is the line an engine crosses on its way into real robot development.

Lightwheel
Insights from the frontier of Physical AI
Contact Us
Product
SimReady Library
EgoSuite
RoboFinals
Lightwheel-Platform Enterprise
About
Blogs
Careers
Contact Us
Customers
Copyright © 2026 Lightwheel Inc. All rights reserved.