OpenAI claims the cost per problem is under $2,000, but Willison notes: 'There's no word on how many problems they spent $2,000 on and didn't solve.' This implies the success rate may not be 100%, but specific failures are not disclosed.
In contrast, Anthropic a few days earlier used Claude and Mythos Preview to discover a cryptographic weakness, spending $100,000 in token costs, with prompts emphasizing 'We're not looking for low-hanging fruit; we want real research to find truly difficult discoveries.'
OpenAI's transparency is evident in publishing formal proofs and the paper, but the absence of prompts limits reproducibility. Willison's comment suggests prompts are key to evaluating research quality.