Post

Log inSign up

Post

Log inSign up

Mikel Bober-Irizar on X: "You've seen some of the puzzles o3 failed, but have you seen the attempts? Yesterday, @OpenAI's o3 dramatically beat the SOTA at @arcprize. But there were 34 tasks that even it couldn't solve with 16 hours of thinking. I've compiled and analyzed all of o3's mistakes below 🧵"

@mikb0b
Mikel Bober-Irizar
@mikb0b
You've seen some of the puzzles o3 failed, but have you seen the attempts? Yesterday, @OpenAI's o3 dramatically beat the SOTA at @arcprize. But there were 34 tasks that even it couldn't solve with 16 hours of thinking. I've compiled and analyzed all of o3's mistakes below 🧵
12 tasks from the ARC-AGI dataset. Each task consists of a number of 2-D reasoning tasks where you are given a handful of input/output pairs and a test example from which you must guess the final output correctly. In all 12 of these highlighted cases, o3 gets the answer wrong.
12:09 AM · Dec 22, 2024·
281.9K
Views
34

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email

Relevant people

Avatar
Mikel Bober-Irizar@mikb0bFollow
24 // Kaggle Competitions Grandmaster & ML/AI Researcher. Building video games @iconicgamesio, machine reasoning @Cambridge_CL, bioscience @ForecomAI.

Trending now

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    You've seen some of the puzzles o3 failed, but have you seen the attempts? Yesterday, @OpenAI's o3 dramatically beat the SOTA at @arcprize. But there were 34 tasks that even it couldn't solve with 16 hours of thinking. I've compiled and analyzed all of o3's mistakes below 🧵
    12 tasks from the ARC-AGI dataset. Each task consists of a number of 2-D reasoning tasks where you are given a handful of input/output pairs and a test example from which you must guess the final output correctly. In all 12 of these highlighted cases, o3 gets the answer wrong.
    12:09 AM · Dec 22, 2024·
    281.9K
    Views
    34
  • @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    Dec 22, 2024
    Here's the post with all the examples and analysis: anokas.substack.com/p/o3-and-arc-a… You might have seen this task today on x dot com as a failure! I actually think o3's answer is as valid as the ground truth here
    8
    @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    Dec 22, 2024
    In several cases, we see o3 struggle to output an aligned grid at all. It seems that problems which require outputting repeated rows can make o3 struggle to keep track.
    6
    @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    Dec 22, 2024
    And on a couple occasions - the model appears to give up on the 2nd attempt, in this case outputting a single black pixel? It's unclear if OpenAI feeds in the previous attempt and tells the model it was wrong, or if something else led to this behaviour.
    3
    @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    Dec 22, 2024
    The full post goes through all these examples, and I'd love to hear your thoughts and theories:
    o3 and ARC-AGI: The unsolved tasks
    From anokas.substack.com
    7
  • @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    Dec 24, 2024
    For a deeper analysis of why o3 did so much better than previous models, and the caveats there might be in that evaluation, check out this thread!
    @mikb0b
    Mikel Bober-Irizar
    @mikb0b
    Dec 24, 2024
    Why do pre-o3 LLMs struggle with generalization tasks like @arcprize? It's not what you might think. OpenAI o3 shattered the ARC-AGI benchmark. But the hardest puzzles didn’t stump it because of reasoning, and this has implications for the benchmark as a whole. Analysis below🧵
    Tokenization labelled on the ARC-AGI prompt used by OpenAI: Find the common rule that maps an input grid to an output grid, given the examples below. Example 1: Input: 0 0 0 ...
  • @MLStreetTalk
    Machine Learning Street Talk
    @MLStreetTalk
    Dec 22, 2024
    Excellent analysis, thanks. If anything this just makes me more impressed with o3 - in many cases where it failed-it was close, or we could imagine how it could be improved quite easily
    1