Startrise Labs · Explorer
Interactive · real data

Benchmark Explorer: does the ranking survive your weights?

Every score in the Startrise LLM benchmark blends 25% browser gates, 45% a blind three-judge panel and 30% a blind human rating. Change that blend, drop briefs you don't care about, and watch 15 models, including the August releases Grok 4.6, Qwen3.8-Max and Muse Spark 1.2, re-rank live from the real per-cell scores.

Re-weight the blend and watch the table move

–Leader under these weights
–Best August entry
–Models that changed rank vs the study's 25 / 45 / 30

Sliders are relative: we normalise them to 100%. Where a cell has no blind human rating, its human share is redistributed across gates and judges, exactly as the study does. A cell that failed a gate stays capped at 40. Like the study, each cell is rounded to one decimal before averaging.

Filter by brief

Leaderboard

Yellow: August additions. Δ = places moved vs the study default. Tap a model to compare it.

    Compare two models, brief by brief

    Model AModel B* no blind human rating

    Cost vs score: what a point of quality costs

    Suite cost versus re-weighted score for 15 models

    Tap a dot to read its numbers.

    Cost = generating all 12 briefs at list rates, log scale; it doesn't change with the brief filter. Source: COST.md, STUDY.md Addenda A–C.

    Share this view

    Your weights, briefs and comparison go into the link. Copy it, or comment on any model row, brief or dot: comments stick to the model, not the pixel, so they survive re-ranking.

    Sources and method