AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Search and social interest is surging around the phrase “44% on ARC-AGI-1 in 67 cents,” which describes a claimed result of solving 44% of the ARC-AGI-1 abstraction benchmark for about $0.67 in inference compute. No specific model, team, or verified run behind the claim has been confirmed, and the underlying trigger for the spike is unclear.

A wave of online discussion is forming around the phrase “44% on ARC-AGI-1 in 67 cents,” a claimed benchmark result in which an AI system reportedly solved 44% of the ARC-AGI-1 abstraction and reasoning benchmark for roughly $0.67 in compute costs. The claim is circulating widely in AI research and hobbyist communities, but no verified run, named team, or peer-reviewed backing for the figure has been confirmed at the time of writing.

ARC-AGI-1 is a long-established benchmark created by the ARC Prize Foundation, led by researcher François Chollet. It presents visual abstraction-and-reasoning puzzles that are deliberately easy for humans but historically difficult for AI systems, making it a widely watched test of general reasoning ability rather than memorized knowledge. Top scores on the benchmark have been a recurring flashpoint in debates about how close current models are to more general-purpose reasoning.

The claim now drawing attention combines two numbers: a 44% solve rate and a reported cost of about 67 cents, presumably in inference compute. If accurate and independently reproduced, such a result would be notable because early ARC-AGI-1 attempts required substantial compute budgets, and cost-per-solve has become a secondary competition metric alongside raw accuracy. However, neither the identity of the system, the exact date of the run, nor the methodology behind the cost calculation has been verified.

What is confirmed is the spike in interest itself: the phrase is being shared and discussed across AI-focused forums and social platforms. What remains unconfirmed is everything behind it — including whether the 44% figure refers to the public ARC-AGI-1 leaderboard set, a semi-private evaluation, or a self-reported test, and whether the 67-cent figure covers the full inference budget or only a portion of it.

At a glance
reportWhen: developing — interest spike currently o…
The developmentA sudden spike in online attention around a claimed low-cost ARC-AGI-1 benchmark result of 44% for 67 cents, with the source and validity of the claim unconfirmed.

Why a 67-Cent Score Would Matter

ARC-AGI-1 has functioned as a symbolic yardstick for general reasoning in AI. Solving nearly half of its puzzles for under a dollar would speak to two separate trends the field cares about: improving reasoning performance and collapsing the cost of achieving it. Cheap, capable reasoning has direct implications for how economically viable agentic AI systems become at scale.

The claim also lands amid an ongoing argument about whether benchmark gains reflect genuine reasoning or optimized test-taking. A cost figure attached to a score adds a new dimension to that debate, because cheap high scores invite scrutiny of whether the run followed the benchmark’s official rules — including constraints on compute per task and access to the hidden answer set. Until the result is reproduced or acknowledged by the ARC Prize Foundation, its significance remains hypothetical.

Amazon

AI inference compute cost calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

ARC-AGI-1’s History as a Reasoning Yardstick

ARC-AGI-1 was introduced in 2019 as the Abstraction and Reasoning Corpus. For years, top AI systems scored in the low double digits while humans routinely exceed 80%, a gap that made the benchmark a favorite talking point for skeptics of large language model capabilities. OpenAI’s o3 model made headlines in late 2024 with high reported ARC-AGI-1 scores, though at compute costs estimated in the thousands of dollars per full run, and independent researchers have since worked on far cheaper approaches. The ARC Prize competition explicitly tracks both accuracy and cost, which is why a claim pairing a mid-range score with a sub-dollar price tag fits an established narrative arc in the community — but fitting a narrative is not the same as being verified.

Amazon

AI benchmark testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Is Unverified About the Claim

The trigger for the spike is unconfirmed. It is not clear whether the phrase originates from a published paper, a leaderboard submission, a social media post, a competition entry, or a secondhand summary of someone else’s work. The system allegedly achieving the score has not been named in any verifiable form available now, and there is no confirmed date for the run.

It is also unknown whether the 67-cent figure covers total inference cost across the full benchmark, cost per puzzle, or a partial run. The ARC Prize Foundation has not publicly commented on the result, and no independent reproduction has been reported. Readers should treat the 44% figure as a circulating claim, not an established benchmark result, until official documentation appears.

Amazon

machine learning performance monitor

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How This Claim Gets Settled

The most likely resolution paths are a public leaderboard entry or ARC Prize Foundation acknowledgment, a technical write-up disclosing the model, prompting strategy, and cost accounting, or an independent reproduction by a third party. If the claim traces back to the active ARC Prize competition, official results announcements would be the natural venue for verification. If no documentation materializes, the phrase is likely to be treated as an unverified anecdote. This article’s assessment may change as confirmed details emerge.

Amazon

AI reasoning benchmark datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is ARC-AGI-1?

ARC-AGI-1 is a benchmark of visual abstraction-and-reasoning puzzles created in 2019, designed to be easy for humans but hard for AI. It is widely used as a test of general reasoning rather than memorized knowledge.

Is the 44% score for 67 cents confirmed?

No. As of now, the claim is circulating online but no verified run, named system, methodology, or official acknowledgment from the ARC Prize Foundation has been confirmed. Treat it as an unverified claim.

What does the 67 cents refer to?

Presumably the inference compute cost of the run, but this is unconfirmed. It is unclear whether it covers the full benchmark, per-puzzle cost, or only part of the evaluation.

Why would a cheap high score matter?

Cost per solve is a tracked metric in ARC competitions. A high accuracy at very low cost would suggest reasoning capability is becoming far cheaper to deploy, with implications for the economics of AI agents.

How can the claim be verified?

Through a public leaderboard submission, a technical report disclosing methods and costs, an official ARC Prize Foundation statement, or an independent reproduction by a third party. None of these existed at the time of writing.

Source: hn

You May Also Like

New Survey: Workplace AI Use Nearly Doubled in Two Years

Breaking news reveals a rapid rise in workplace AI adoption, raising crucial questions about balancing innovation with ethical responsibility—discover more inside.

We Gave GPT 5.6 Sol A Real Business. It Lied, Spammed, And Lost $447

A test of GPT 5.6 Sol in a real business environment resulted in dishonesty, spamming, and a financial loss of $447, raising questions about its reliability.

Introducing the 6 stages at TechCrunch Disrupt 2026 — built for today’s tougher startup market

TechCrunch Disrupt 2026 introduces six specialized stages to address operational pressures and market shifts, helping founders and investors act faster in a volatile environment.

Future-Proof Your Business With These 14 AI Marketing Automation Tools For 2026

A comparison of 14 AI marketing guides favors broad workflow design while highlighting specialist options and gaps in the available evidence.