

RoboChallenge, Lightwheel, and NVIDIA Collaborate to Bridge Simulation and Reality for Embodied AI Evaluation

Embodied AI Needs a New
Evaluation Infrastructure
Embodied AI is entering a new stage of development. With the rapid advancement of Vision-Language-Action (VLA) models, World Models, and robot foundation models, robots are becoming increasingly capable of perceiving, understanding, and executing complex tasks. The field is accelerating its transition from laboratory research to real-world deployment, ushering in the era of Physical AI.
However, as model capabilities continue to improve, a new challenge has emerged as a critical bottleneck for the industry: how can we evaluate embodied intelligence in a way that is scalable, reproducible, and comparable?
Unlike the large language model ecosystem, which benefits from widely recognized benchmarks such as ImageNet, GLUE, and MMLU, robotics has long suffered from fragmented evaluation environments. Different research teams rely on different robot platforms, task definitions, and experimental setups, making it difficult to establish a common standard even when evaluating similar capabilities. While real-robot testing provides the most reliable assessment, it is constrained by hardware costs, laboratory resources, and operational complexity, limiting its scalability.
As more developers build foundation models for robots operating in the physical world, the need for a unified evaluation infrastructure has never been more urgent.
Discovering Industry Needs Through
Large-Scale Real-Robot Evaluation
Over the past two years, RoboChallenge has grown into one of the world's largest real-robot evaluation platforms. Since the launch of the Table 30 Benchmark, RoboChallenge has provided standardized real-world evaluation environments for robotics companies, research institutions, and embodied AI foundation model developers.
In parallel, RoboChallenge has partnered with leading international conferences such as CVPR and ICRA to organize embodied AI competitions, completing more than 9,000 real-robot evaluations.
These large-scale deployments have revealed an increasingly clear trend: developers not only need real-robot testing to validate final performance, but also require a standardized simulation benchmark that supports rapid experimentation, continuous iteration, and large-scale development throughout the model lifecycle.
The future of embodied AI evaluation is not about replacing real robots with simulation—it is about creating a closed-loop system in which simulation and real-world evaluation reinforce each other.
From Table 30 to a Simulation
Benchmark
Based on this shared vision, RoboChallenge, Lightwheel, and NVIDIA held in-depth discussions during NVIDIA GTC on the future of embodied AI evaluation and jointly launched a new collaboration.
Built upon the open-source NVIDIA Isaac Lab platform, the project will be jointly developed by Lightwheel and Dexmal to create a simulation counterpart of the RoboChallenge Table 30 Benchmark, establishing a benchmark system that closely mirrors real-robot environments.
Unlike conventional simulation benchmarks that simply recreate tasks in virtual environments, this initiative aims to establish a one-to-one correspondence between simulation and physical robot evaluation. Task definitions, evaluation metrics, success criteria, and scoring methodologies will be aligned as closely as possible with the RoboChallenge real-robot platform, enabling meaningful comparison between simulation results and real-world performance.
This alignment will allow developers to study Sim-to-Real transfer under a unified benchmark and evaluate how effectively models generalize from simulation to physical deployment.
Lightwheel plays a critical role in this collaboration by transforming real-world RoboChallenge tasks—including objects, physical interactions, friction, deformation, operational procedures, and evaluation logic—into interactive, reproducible, and benchmark-ready simulation assets.
Powered by its proprietary "Solve–Measure–Generate" full-stack simulation technology, Lightwheel combines accurate physical parameter measurement with SimReady world generation capabilities to significantly improve the physical consistency between simulation and reality. This provides a more trustworthy foundation for Sim-to-Real research and cross-model performance evaluation.
Establishing the RoboChallenge
Simulation Working Group
To advance this vision, RoboChallenge will establish a new Simulation Working Group, jointly led by Lightwheel and RoboChallenge.
The working group will be responsible for developing RoboChallenge's simulation benchmark ecosystem and designing future simulation and real-world benchmarks.
It will serve as a collaborative platform connecting researchers, developers, and industry partners to jointly build a more open, standardized, and reproducible evaluation framework while continuously evolving benchmark specifications, evaluation protocols, and developer toolchains.
Why Isaac Lab?
The choice of NVIDIA Isaac Lab as the technical foundation for this collaboration is deliberate.
As one of the fastest-growing open-source robotics simulation platforms, Isaac Lab has become an essential infrastructure for the global Physical AI community. Built on NVIDIA Omniverse and Isaac Sim, it provides high-fidelity simulation capabilities together with a unified framework for robot learning, reinforcement learning, and Vision-Language-Action (VLA) model training.
Equally important, its open ecosystem enables researchers worldwide to easily access, reproduce, and extend benchmark environments, fostering continuous community-driven innovation.
Meanwhile, Lightwheel has built extensive expertise in robotics simulation and Physical AI infrastructure, including simulation environment development, benchmark engineering, robot development toolchains, simulation world generation, real-world physics alignment, large-scale data generation, and industrial-grade evaluation systems.
By combining Lightwheel's engineering capabilities with RoboChallenge's experience in large-scale real-robot evaluation, the collaboration will efficiently achieve high-fidelity mapping from physical tasks to simulation environments.
Building RoboChallenge's First
Reproducible Simulation Benchmark
The first phase of the project will focus on the Table 30 Benchmark.
The teams plan to implement a complete simulation version of Table 30 within Isaac Lab while ensuring that evaluation logic, success rate calculations, and task execution workflows remain fully consistent with the real-robot benchmark.
This will give RoboChallenge its first fully reproducible simulation benchmark, allowing developers to participate in benchmark development, algorithm validation, and Sim-to-Real research without requiring access to physical robots.
By lowering the barrier to participation, the project will significantly expand RoboChallenge's global developer ecosystem.
Introducing a Simulation-First
Competition Framework
Building on this foundation, RoboChallenge also plans to evolve its competition format toward a simulation-first paradigm.
Future RoboChallenge competitions will first release tasks and evaluation environments through Isaac Lab. Participants will develop and evaluate their models within a standardized simulation environment, while the system automatically performs large-scale evaluation and generates public leaderboards.
Top-performing teams will then advance to RoboChallenge's real-robot platform for final validation and championship evaluation.
By combining large-scale simulation screening with real-world verification, this approach dramatically lowers participation barriers while maximizing the efficiency of valuable real-robot resources, enabling broader participation in embodied AI research and innovation.
Looking Ahead: Toward Industrial-
Scale Real-World Benchmarks
From a longer-term perspective, Table 30 is only the beginning.
RoboChallenge, Lightwheel, and NVIDIA are jointly exploring the next generation of industrial-scale benchmark suites—the Real Series.
The initiative aims to define standardized tasks across real industrial scenarios, including robotic manipulation, assembly, tool use, logistics, quality inspection, and human-robot collaboration.
Each task will have corresponding implementations in both Isaac Lab and the RoboChallenge real-robot platform, establishing a unified mapping between simulation and physical environments.
Ultimately, researchers will be able to develop and benchmark models in simulation while validating them on real robots, measuring embodied intelligence under a single, consistent evaluation framework.
Building the Next Generation
Infrastructure for Physical AI
Throughout the history of artificial intelligence, infrastructure has been the catalyst for innovation.
ImageNet accelerated the progress of computer vision. Standardized benchmarks transformed the development of large language models. Today, embodied AI is entering its own infrastructure era.
By combining RoboChallenge's real-robot evaluation capabilities, Lightwheel's simulation engineering expertise, and the open ecosystem of NVIDIA Isaac Lab, the three partners aim to build the next generation of Physical AI evaluation infrastructure.
Their vision is to enable developers to build in simulation, validate in the real world, and measure progress under a unified standard.
Because the future of embodied AI depends not only on more powerful models, but also on evaluation infrastructure that is open, reproducible, scalable, and trusted by the global community.