People are using Super Mario to benchmark AI now

Thought Pokémon was a tough benchmark for AI? One group of researchers argues that Super Mario Bros. is even tougher.

Hao AI Lab, a research org at the University of California San Diego, on Friday threw AI into live Super Mario Bros. games. Anthropic’s Claude 3.7 performed the best, followed by Claude 3.5. Google’s Gemini 1.5 Pro and OpenAI’s GPT-4o struggled.

It wasn’t quite the same version of Super Mario Bros. as the original 1985 release, to be clear. The game ran in an emulator and integrated with a framework, GamingAgent, to give the AIs control over Mario.

Super Mario Bros. AI benchmark — **Image Credits:**Hao Lab

GamingAgent, which Hao developed in-house, fed the AI basic instructions, like, “If an obstacle or enemy is near, move/jump left to dodge” and in-game screenshots. The AI then generated inputs in the form of Python code to control Mario.

Still, Hao says that the game forced each model to “learn” to plan complex maneuvers and develop gameplay strategies. Interestingly, the lab found that so-called reasoning models like OpenAI’s o1, which “think” through problems step by step to arrive at solutions, performed worse than “non-reasoning” models, despite being generally stronger on most benchmarks.

One of the main reasons reasoning models have trouble playing real-time games like this is that they take a while — seconds, usually — to decide on actions, according to the researchers. In Super Mario Bros., timing is everything. A second can mean the difference between a jump safely cleared and a plummet to the death.

Games have been used to benchmark AI for decades. But some experts have questioned the wisdom of drawing connections between AI’s gaming skills and technological advancement. Unlike the real world, games tend to be abstract and relatively simple, and they provide a theoretically infinite amount of data to train AI.

The recent flashy gaming benchmarks point to what Andrej Karpathy, a research scientist and founding member at OpenAI, called an “evaluation crisis.”

“I don’t really know what [AI] metrics to look at right now,” he wrote in a post on X. “TLDR my reaction is I don’t really know how good these models are right now.”

At least we can watch AI play Mario.

S4 Capital downgrades sales outlook as tariff issues hinder financial outlook

Oil slips on rising OPEC+ output, despite Canadian supply concerns

Australia raises minimum wages by 3.5%

International tourist spending in Europe seen up 11% this year, report says

Could the euro replace the dollar as global reserve currency? Its not getting any lesslikely

S4 Capital downgrades sales outlook as tariff issues hinder financial outlook

Oil slips on rising OPEC+ output, despite Canadian supply concerns

Australia raises minimum wages by 3.5%

International tourist spending in Europe seen up 11% this year, report says

Could the euro replace the dollar as global reserve currency? Its not getting any lesslikely

People are using Super Mario to benchmark AI now

Share

Revival of UVB-76: Cold War Ghosts in Modern Warfare

How to watch Apples WWDC 2025 keynote

Profitable African fintech PalmPay is in talks to raise as much as $100M

North America takes the bulk of AI VC investments, despite tough political environment

iOS 19: All the rumored changes Apple could be bringing to its new operating system

Popular

Microsofts new AI agent can control software and robots

Waymo and Toyota are dating. If they get serious, a new autonomous vehicle could be created.

Elon Musk and Donald Trump are smack talking each other into their own digital echo chambers

RFK Jr. promptly cancels vaccine advisory meeting, pulls flu shot campaign

Anthropic CEO says spies are after $100M AI secrets in a few lines of code

Discord appoints former Activision Blizzard exec Humam Sakhnini as CEO

Related Articles

Elon Musk and Donald Trump are smack talking each other into their own digital echo chambers

Revival of UVB-76: Cold War Ghosts in Modern Warfare

How to watch Apples WWDC 2025 keynote

Profitable African fintech PalmPay is in talks to raise as much as $100M

North America takes the bulk of AI VC investments, despite tough political environment

iOS 19: All the rumored changes Apple could be bringing to its new operating system

Attacks on the Three Facets of My Identity

Data breach at newspaper giant Lee Enterprises affects 40,000 people

About Us

Popular Category

Editor Picks

Airlines entrusted less flight paths as worldwide dispute zones broaden

2 kids amongst 11 eliminated in stampede throughout IPL title events in India: How the catastrophe unfolded

People are using Super Mario to benchmark AI now

Share

Related posts:

Popular

Related Articles

About Us

Popular Category

Editor Picks