Startrise Labs
Method · no results

How Startrise Labs runs LLM studies

Every Startrise Labs study scores a model with three instruments that fail in different ways: automated browser gates, a blind three-judge LLM panel and a blind human rating. Then we publish everything that's wrong with the measurement, right next to the numbers.

Every score is a blend of three instruments that fail differently

0.25Automated gates: Playwright drives every page in a real browser
0.45Judge panel: median of 3 LLM judges from 3 labs, blind to authorship
0.30Blind human rating: model names hidden until a task is fully scored

Gates catch dead pages, judges read the source, a human catches what both miss. When a cell has no human rating yet, its 30% is redistributed proportionally and the row is flagged as un-reviewed, so a provisional number never passes as a settled one.

Six stages, same harness for every model

  1. Briefs + one output contract

    Plain-markdown briefs joined with a shared contract: one self-contained HTML file, no placeholders.

  2. Single-shot generation

    Identical system prompt, no follow-up turn, streamed. Recovered output is logged as a contract violation.

  3. Automated gates

    Headed Chromium at 1440×900, run serially so pages don't fight over the GPU.

  4. Blind judge panel

    Screenshots, full source and gate readings go to three judges who never learn who built it.

  5. Blind human review

    One reviewer rates each cell with model names hidden.

  6. Blend and publish caveats

    Weights applied, gate failures capped, and every known measurement defect written up next to the numbers.

Gates measure behaviour in a real browser, not code style

GateWhat it checks
self_containedNo local or non-CDN resources; contract kept
loadsOpens from file:// without error
console_cleanZero uncaught exceptions
rendersLuminance spread: a blank page scores near zero
fpsFrames the page itself schedules over 3 seconds
interactionPixel change after scripted scroll, pointer and keys
economyLines and bytes, proportionate to the brief
identitySelf-pitch brief only: does the page name its author correctly

Only the checks a brief declares are counted and weights renormalise over them, so a static landing page is never marked down for 0 fps.

Three judges, three labs, three deliberately different lenses

Craft

A design director: composition, type, colour, motion, point of view.

Engineering

A senior frontend engineer reading source: what was built versus faked.

Brief

The client: was every explicit requirement met, completely.

Each judge scores craft, technique, adherence and originality 0–10. We take the median per axis, not the mean, so one outlier can't drag a result, and we keep the spread: a 4-point disagreement on any axis flags the cell as a split panel for human eyes.

What we publish next to every number

Sample size

Runs are single-shot today. Our scoring spec calls for N=3 minimum, N=5 preferred, with medians and IQRs. Until then, small gaps are noise.

Conflicts of interest

When a judge model is also in the field, we say so and measure its self-bias in a separate audit.

Contamination

Published briefs can enter training data. Models released after publication get a disclosure; a private holdout set is required.

Our own bugs

When a gate is wrong we fix it, recompute every cell and document what moved.

Config is not the model

Reasoning models spend output budget thinking. A too-low max_tokens measures our config, so we give headroom and say so.

Study in progress · no results yet

Next: Startrise Local Code

A sister study for small local coding models on Apple Silicon. The question: which models give the most useful work per second per GB inside a 32GB Mac envelope, with weights on an external SSD. The primary score is unit tests passing after a single-shot repair. Results aren't in, so this page shows none.

Sources