Back to feed
News Story
APriority76
Artificial Analysis (X)
2 sources

Launching Optima: Custom Benchmarks for Everyone

Artificial Analysis has launched Optima, a new platform that lets users create custom benchmarks to evaluate and compare AI models on performance, speed, and cost efficiency. It supports building benchmarks from uploaded datasets, agent traces, or simple descriptions, and incorporates the company's research expertise.

SynthePulse Insight · AI deep reading

Optima Launch: Let Anyone Create Custom AI Benchmarks

Version 1 · 1 source

Artificial Analysis introduces Optima, a platform that lets users build benchmarks from their own data and compare model performance, cost, and speed.

  • Optima allows users to build benchmarks by uploading datasets, importing agent traces, or describing use cases.
  • The platform supports running benchmarks on the latest models and automatically updates leaderboards.
  • Optima integrates Artificial Analysis' scoring methods, including objective criteria and pairwise comparisons.
  • The platform tracks cost and latency for each task, helping users find more cost-effective models.
Open section navigationOptima Launch and Core Features

Optima Launch and Core Features

On August 13, 2026, Artificial Analysis announced the launch of Optima, a platform that allows anyone to create custom benchmarks. The platform aims to leverage Artificial Analysis' research and experience in benchmarking to help users evaluate model performance on their specific workloads.

Optima offers three ways to build benchmarks: uploading existing evaluation datasets (with support for importing from Hugging Face), importing agent traces from platforms like Arize AI, Braintrust, and Langfuse, or building from coding environments by installing the Optima skill. Additionally, users can generate benchmarks by describing their use case and providing example inputs and outputs.

Running and Scoring Mechanism

Users can run the same benchmark across multiple leading models with a single click, and keep leaderboards updated as new models are released. Optima integrates Artificial Analysis' scoring methods, including objective criteria or pairwise judgment methods identical to those used in GDPval-AA and AA-Briefcase. For pairwise judgments, users select preferred responses from samples, and Optima uses these preferences to rank models in the test set.

Performance, Cost, and Efficiency Comparison

Optima not only measures model performance but also tracks cost and latency for each task, providing category-level results and supporting custom metrics. This allows users to compare trade-offs between models, such as finding a model that reduces cost by 10x with minimal quality degradation.

Tester Cases and Availability

Before launch, testers used Optima to answer questions like "Which model can save 10x cost for my finance and accounting agent without significantly reducing quality?" and "Which model best matches a lawyer's writing style?" Optima is now available, and users can build their own benchmarks at artificialanalysis.ai/optima.

Credibility boundary

This report is based on Artificial Analysis' official launch post on X, which is a first-party source. All feature descriptions and tester cases are from that post and have not been independently verified.

Insight takeaway

The launch of Optima lowers the barrier to creating custom benchmarks, enabling enterprises to evaluate models based on their own needs and optimize the balance between cost and performance.

Primary report

Artificial Analysis (X)

Primary source

Same-event coverage

Also covered by 1 sources