Benchmark Explorer: does the ranking survive your weights?
Every score in the Startrise LLM benchmark blends 25% browser gates, 45% a blind three-judge panel and 30% a blind human rating. Change that blend, drop briefs you don't care about, and watch 15 models, including the August releases Grok 4.6, Qwen3.8-Max and Muse Spark 1.2, re-rank live from the real per-cell scores.
Re-weight the blend and watch the table move
Sliders are relative: we normalise them to 100%. Where a cell has no blind human rating, its human share is redistributed across gates and judges, exactly as the study does. A cell that failed a gate stays capped at 40. Like the study, each cell is rounded to one decimal before averaging.
Filter by brief
Leaderboard
Yellow: August additions. Δ = places moved vs the study default. Tap a model to compare it.
Compare two models, brief by brief
Cost vs score: what a point of quality costs
Tap a dot to read its numbers.
Cost = generating all 12 briefs at list rates, log scale; it doesn't change with the brief filter. Source: COST.md, STUDY.md Addenda A–C.
Share this view
Your weights, briefs and comparison go into the link. Copy it, or comment on any model row, brief or dot: comments stick to the model, not the pixel, so they survive re-ranking.
Sources and method
- Per-cell technical, judge and human scores: sr-llm-benchmark results/2026-07-26-0159, 2026-08-09-2106, 2026-08-12-0114, 2026-08-12-2359 scores.json. Embedded unchanged; human ratings stored 0–10, shown ×10.
- Blend, redistribution and the 40 gate cap follow STUDY.md §1 and config/scoring.json. At 25 / 45 / 30 with all briefs, this page reproduces every published overall score.
- Suite cost: COST.md (July field) and STUDY.md Addenda A–C (Muse $0.62, Qwen3.8-Max $2.92, Grok 4.6 $1.63).
- N=1 per cell, one human reviewer. Small gaps are noise under any weighting.