World Humanoid Robot Games
(Credit: Unsplash)

AI Keeps Getting Smarter—And Now Scientists Want to Know If It Can Build Robots

Artificial intelligence (AI) has rapidly become part of mainstream culture, with people around the world now using AI systems every day. Researchers at Harvard University and the Georgia Institute of Technology have designed a new benchmark to determine whether AI can tackle a considerably different challenge: robot engineering.

From writing code to creating AI agents and designing software, AI systems have demonstrated increasingly sophisticated capabilities in the digital world. Robots, however, introduce a physical, three-dimensional challenge. Researchers are now testing how far AI coding agents can go in meeting the demands of robotics engineering.

Now, a research team led by Na Li, a professor at Harvard’s John A. Paulson School of Engineering and Applied Sciences, and Bo Dai, an assistant professor at Georgia Tech, is tackling the question with a new open-source benchmark called RLE-Bench. The benchmark measures whether general-purpose AI agents can complete tasks associated with building and engineering robotic systems.

RLE-Bench evaluates whether an AI agent can understand how software interacts with sensors, motors, mechanical components, and the laws of physics. It includes 48 tasks across four areas, ranging from developing control and perception algorithms to designing robot hardware and interfaces.

Some challenges involve creating algorithms that allow robots to move, teaching robotic systems to perform tasks, interpreting sensor data, and designing hardware components and interfaces.

“Benchmarks are important because they give the field a common measuring stick, allowing us to compare systems under the same conditions and identify where important capability gaps remain,” Li said in a statement. “Robotics already has many valuable benchmarks, but many focus primarily on learning and evaluating control policies. RLE-Bench takes a broader view. Can an AI system perform the range of engineering tasks required to make a robotic system actually work?”

The tasks take place in simulation, where AI agents can create code, run their solutions, review the results, and make decisions to improve their work. Importantly, however, the challenges are designed around real-world engineering constraints.

An AI system, for example, might produce a design that appears successful in simulation but fails once factors such as weight, torque, balance, or geometry are taken into consideration.

Li notes that the current benchmark is only the beginning.

“RLE-Bench 1.0 is only a starting point,” he says. “We hope researchers and robotics practitioners will contribute tasks based on the engineering problems they encounter in their own work. Our goal for RLE-Bench 2.0 is to capture a much broader picture of what an AI system needs to do to function as a robot learning engineer.”

In one example, an AI agent must design a mobile base capable of supporting multiple robotic arms that need to reach objects at different heights on a shelf. While the AI-generated design might appear to work digitally, its real-world counterpart could be unstable and tip over.

China’s Robot Olympics 

China seems to be advancing quickly with regard to both robots and AI. In August 2026, Beijing hosted the second World Humanoid Robot Games, an event that put robotic systems through a variety of physical challenges.

The games featured more than 2,000 robots from 666 teams across 16 countries, with events involving potential real-world applications such as household tasks, industrial operations, and emergency rescue.

RLE-Bench and competitions like the World Humanoid Robot Games evaluate very different aspects of robotics and therefore cannot be directly compared. RLE-Bench asks whether general-purpose AI agents can perform the engineering work required to design and develop robotic systems, while physical competitions provide developers with opportunities to test how robots perform under real-world conditions.

Together, these approaches illustrate two sides of a growing challenge in robotics: developing AI systems capable of reasoning about physical machines and determining whether the resulting robotic technologies can actually perform reliably outside a simulated environment.

Chrissy Newton is a PR professional and the founder of VOCAB Communications. She currently appears on The Discovery Channel and Max and hosts the Rebelliously Curious podcast, which can be found on YouTube and on all audio podcast streaming platforms. Follow her on X: @ChrissyNewton, Instagram: @BeingChrissyNewton, and chrissynewton.com. To contact Chrissy with a story, please email chrissy @ thedebrief.org.