Skip to content

Get On Scoop

  • Advertise with us!
Advertise
  • Home
  • Tech Scoop
  • Anthropic used Pokémon to benchmark its newest AI model
  • Tech Scoop

Anthropic used Pokémon to benchmark its newest AI model

getonscoop 02/24/2025
Anthropic used Pokémon to benchmark its newest AI model


Anthropic used Pokémon to benchmark its newest AI model. Yes, really.

In a blog post published Monday, Anthropic said that it tested its latest model, Claude 3.7 Sonnet, on the Game Boy classic Pokémon Red. The company equipped the model with basic memory, screen pixel input, and function calls to press buttons and navigate around the screen, allowing it to play Pokémon continuously.

A unique feature of Claude 3.7 Sonnet is its ability to engage in “extended thinking.” Like OpenAI’s o3-mini and DeepSeek’s R1, Claude 3.7 Sonnet can “reason” through challenging problems by applying more computing — and taking more time.

That came in handy in Pokémon Red, apparently.

Compared to a previous version of Claude, Claude 3.0 Sonnet, which failed to leave the house in Pallet Town where the story begins, Claude 3.7 Sonnet successfully battled three Pokémon gym leaders and won their badges. 

Anthropic Pokemon Red
Image Credits:Anthropic

Now, it’s not clear how much computing was required for Claude 3.7 Sonnet to reach those milestones — and how long each took. Anthropic only said that the model performed 35,000 actions to reach the last gym leader, Surge.

It surely won’t be long before some enterprising developer finds out.

Pokémon Red is more of a toy benchmark than anything. However, there is a long history of games being used for AI benchmarking purposes. In the past few months alone, a number of new apps and platforms have cropped up to test models’ game-playing abilities on titles ranging from Street Fighter to Pictionary.



Source link

Continue Reading

Previous: Chelsea are over-reliant on Cole Pamer – Enzo Maresca
Next: 2025 Is a Year Full of Meteor Showers: The Next One Arrives This Week

Related Stories

B&H gaming monitor sale: Save up to $500 Samsung Odyssey G95C 49-inch Curved Ultrawide Gaming Monitor
  • Tech Scoop

B&H gaming monitor sale: Save up to $500

03/10/2025
Manus probably isn’t China’s second ‘DeepSeek moment’ Manus
  • Tech Scoop

Manus probably isn’t China’s second ‘DeepSeek moment’

03/09/2025
5 underrated movies on Netflix you need to watch in March 2025 Nicolas Cage puts his fists together in The Unbearable Weight of Massive Talent.
  • Tech Scoop

5 underrated movies on Netflix you need to watch in March 2025

03/09/2025

Live Scoop

What the Scoop?

Categories

  • Current Events
  • Food Scoop
  • News Scoop
  • Tech Scoop

You may have missed

Oatmeal Protein Cookies – WellPlated.com Oatmeal Protein Cookies – WellPlated.com
  • Food Scoop

Oatmeal Protein Cookies – WellPlated.com

04/28/2025
Weekly Meal Plan 4.27.25 – WellPlated.com Weekly Meal Plan 4.27.25 – WellPlated.com
  • Food Scoop

Weekly Meal Plan 4.27.25 – WellPlated.com

04/28/2025
Easy Refried Beans – Mel’s Kitchen Cafe Easy Refried Beans - Mel's Kitchen Cafe
  • Food Scoop

Easy Refried Beans – Mel’s Kitchen Cafe

04/28/2025
Easy Spicy Mayo Recipe Easy Spicy Mayo Recipe
  • Food Scoop

Easy Spicy Mayo Recipe

04/26/2025

Terms & Services | Privacy Policy

  • Partners
  • Press
  • About
  • Useful
Copyright © All rights reserved. | DarkNews by AF themes.