AI model scores 29.5% on ARC-AGI-1 at $0.00070
A handful of colored grids sits between the machine and its answer: extend a line, copy a shape, find the hidden rule. Pathway’s 150-million-parameter BDH-CQ model solved 29.5% of the ARC-AGI-1 puzzles it faced with two attempts per task, at a reported cost of $0.00070 per task.
ARC-AGI-1 is a benchmark built around visual pattern recognition. The model receives a few completed examples on colored grids and must infer the rule for a new one. Unlike familiar chatbots that can work through a problem by generating a visible chain of written words, BDH-CQ updates an internal memory step by step and reasons in an internal workspace before producing its answer.
That smaller footprint is the point. Parameters are the internal values an AI adjusts while learning patterns; some of today’s largest models use hundreds of billions of them, while BDH-CQ uses 150 million. Pathway’s team says the model sets a new cost-efficiency record on ARC-AGI-1, and an independent team from Bielik AI and New York University separately confirmed the reported results.
The boundary is visible in the same test. BDH-CQ handled basic visual operations, including copying shapes and extending lines, but struggled with tasks involving larger sets of objects or several steps of visual reasoning. The work is presented in a 2026 arXiv preprint, and the researchers say they plan to train larger versions on mathematical, linguistic and visual problems.
The practical consequence is narrower and clearer than a wholesale replacement for large language models: if the approach scales, some small, repetitive reasoning jobs could require less computation and cost less to run. For now, that remains a possibility grounded in one benchmark.
Comments
Loading the thread…
Sign in to leave a comment. Sign in