How Startrise Labs runs LLM studies
Every Startrise Labs study scores a model with three instruments that fail in different ways: automated browser gates, a blind three-judge LLM panel and a blind human rating. Then we publish everything that's wrong with the measurement, right next to the numbers.
Every score is a blend of three instruments that fail differently
Gates catch dead pages, judges read the source, a human catches what both miss. When a cell has no human rating yet, its 30% is redistributed proportionally and the row is flagged as un-reviewed, so a provisional number never passes as a settled one.
Six stages, same harness for every model
Briefs + one output contract
Plain-markdown briefs joined with a shared contract: one self-contained HTML file, no placeholders.
Single-shot generation
Identical system prompt, no follow-up turn, streamed. Recovered output is logged as a contract violation.
Automated gates
Headed Chromium at 1440×900, run serially so pages don't fight over the GPU.
Blind judge panel
Screenshots, full source and gate readings go to three judges who never learn who built it.
Blind human review
One reviewer rates each cell with model names hidden.
Blend and publish caveats
Weights applied, gate failures capped, and every known measurement defect written up next to the numbers.
Gates measure behaviour in a real browser, not code style
| Gate | What it checks |
|---|---|
self_contained | No local or non-CDN resources; contract kept |
loads | Opens from file:// without error |
console_clean | Zero uncaught exceptions |
renders | Luminance spread: a blank page scores near zero |
fps | Frames the page itself schedules over 3 seconds |
interaction | Pixel change after scripted scroll, pointer and keys |
economy | Lines and bytes, proportionate to the brief |
identity | Self-pitch brief only: does the page name its author correctly |
Only the checks a brief declares are counted and weights renormalise over them, so a static landing page is never marked down for 0 fps.
Three judges, three labs, three deliberately different lenses
Craft
A design director: composition, type, colour, motion, point of view.
Engineering
A senior frontend engineer reading source: what was built versus faked.
Brief
The client: was every explicit requirement met, completely.
Each judge scores craft, technique, adherence and originality 0–10. We take the median per axis, not the mean, so one outlier can't drag a result, and we keep the spread: a 4-point disagreement on any axis flags the cell as a split panel for human eyes.
What we publish next to every number
Sample size
Runs are single-shot today. Our scoring spec calls for N=3 minimum, N=5 preferred, with medians and IQRs. Until then, small gaps are noise.
Conflicts of interest
When a judge model is also in the field, we say so and measure its self-bias in a separate audit.
Contamination
Published briefs can enter training data. Models released after publication get a disclosure; a private holdout set is required.
Our own bugs
When a gate is wrong we fix it, recompute every cell and document what moved.
Config is not the model
Reasoning models spend output budget thinking. A too-low max_tokens measures our config, so we give headroom and say so.
Next: Startrise Local Code
A sister study for small local coding models on Apple Silicon. The question: which models give the most useful work per second per GB inside a 32GB Mac envelope, with weights on an external SSD. The primary score is unit tests passing after a single-shot repair. Results aren't in, so this page shows none.
Sources
- sr-llm-benchmark/STUDY.md §1 (method), §7–8 (caveats)
- sr-llm-benchmark/config/scoring.json (weights, gates, axes)
- sr-local-code-bench/STUDY.md and config/scoring.json (pilot protocol)