Show HN: Artificial Analysis tool to create custom benchmarks for any use case
Artificial Analysis Optima Build your own custom benchmark Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom…
Artificial Analysis Optima Build your own custom benchmark Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency Try OptimaHow it works Contract Review Benchmark Build agent drafting tasks Benchmark how well models review our supplier contracts — flag risks, extract terms. Working Reading your example contracts Drafting grading rubrics Writing task 18 of 24 Contract Review Benchmark 24 tasks · rubrics attached · 4 categories Why build your benchmark with Optima Results specific to your use case Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released. Cut cost and time by over 10x Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models Built on Artificial Analysis grading expertise Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head How it works 1 Give Optima your context Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you. 2 Choose your evaluation type Task style Objective Rubric judge Subjective Optima Q&A The model answers questions that have known correct answers Document input The model answers questions about your uploaded files Agentic The model completes tasks and produces deliverables, using tools along the way Interaction Simulate a real conversation with different types of users Coming soon Q&AObjective grading Example task prompt Which HS tariff code applies to lithium-ion e-bike batteries? How it's graded The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion. 3 Run across models Or bring your own agent Your own agent can compete in the same run, over HTTP. 4 Grade and decide Contract Review Benchmark24 tasks · 4 models ModelScoreCost / TaskTime / Task Claude Fable 560$0.1874s GPT-5.6 Sol59$0.1152s Kimi K357$0.0461s Gemini 3.6 Flash50$0.0229s Pricing Optima pricing is based on token usage. Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top. Rubric grading is $0.25 per criterion, per model. Pairwise grading is $0.75 per match. At the start of benchmark creation, each benchmark run and each grading pass, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for what was actually used, regardless of the estimate. Find the best model for your work Bring your own tasks or describe your use case, and get graded results with the costs attached. Try Optima