Optima Lets Users Create Custom Benchmarks for Their Own Workloads
Optima Lets Users Create Custom Benchmarks for Their Own Workloads August 13, 2026 Artificial Analysis has launched Optima, a platform for creating custom benchmarks from a user...
By Software Development Team
Optima Lets Users Create Custom Benchmarks for Their Own Workloads
August 13, 2026
Artificial Analysis has launched Optima, a platform for creating custom benchmarks from a user’s own data and workflows. Users can run the resulting benchmark across leading models and compare quality, cost per task, and time per task.
Introducing Optima
Building and operating model benchmarks can be difficult, particularly when evaluations need to reflect a specific workload. Optima packages Artificial Analysis’ research and benchmarking infrastructure into a platform for testing models against custom use cases.
The platform is designed to help users identify the best model for a task, or find an alternative to an existing setup with lower cost or faster execution while maintaining a similar level of performance.
How Optima works
Optima applies the benchmarking methods and infrastructure used by Artificial Analysis across several stages of the evaluation process.
Create benchmarks from existing data and workflows
Users can create an Optima benchmark in several ways:
- Upload an evaluation dataset from their own files or from Hugging Face.
- Import agent traces from platforms such as Arize, Braintrust, and Langfuse.
- Install the Optima skill to build a benchmark using context from a coding environment and previous sessions.
- Describe a use case and provide example inputs and outputs, allowing Optima to build the benchmark.
Test the latest models
A single benchmark can be run across leading models, with leaderboards updated when new models are released.
Apply Artificial Analysis grading methods
Optima can evaluate responses against objective rubric criteria. It also supports the pairwise judging method used for Artificial Analysis benchmarks, including GDPval-AA and AA-Briefcase.
For pairwise judging, users select their preferred responses from a sample. Optima then uses those preferences to rank models across the test set.
Compare performance, cost, and time
Optima tracks more than model quality. Cost per Task and Time per Task are measured alongside benchmark scores, with category-level results and support for custom metrics. This allows users to examine tradeoffs between models for a particular use case.
Examples from pre-release testing
Before the launch, a group of pre-release testers used Optima to create benchmarks for several tasks, including:
- Finding a model that could reduce costs by 10x without a meaningful decline in quality for a finance and accounting agent.
- Identifying the model whose writing style best matched that of lawyers for a legal agent.
- Determining which model could best identify different elements in a custom image dataset.
Availability
Optima became available on August 13, 2026. Users can create benchmarks from their own data, agent traces, coding environments, or described use cases, then evaluate models using quality, cost, time, and custom metrics.