Default Index
Turn frontend taste arguments into an inspectable benchmark
Ownership
Solo Designer & Researcher
Team
Independent
Primary proof
108

Overview
Default Index is a technical investigation into recurring design patterns in model-generated frontends and the instructions that may reduce them. It replaces screenshot-level taste claims with a registered protocol, artifact capture, rule-based detectors, uncertainty, and blind review. The current public release is a deterministic calibration corpus: 108 artifacts prove that the measurement system works end to end, but they are not presented as findings about frontier models.
The Challenge
Critiques of generated interfaces often collapse into a familiar but weak claim: everything looks the same. Without fixed briefs, preserved source, comparable render conditions, explicit detectors, and human review, that claim cannot distinguish a model default from a prompt effect, framework convention, or evaluator preference. The investigation needed a way to make design-pattern evidence inspectable before spending compute on provider comparisons.
Constraints
- •Fixture calibration must remain visibly separate from claims about real models or providers.
- •Every aggregate needs a path back to source, desktop and mobile renders, detector evidence, and review state.
- •Design-pattern detectors are proxies with uncertainty, not objective measurements of interface quality.
- •Provider runs and their cost remain gated until the protocol and smoke tests justify them.
Decision Log
Problem
Starting with live model runs would spend compute before the analysis pipeline had been calibrated.
Decision
Created deterministic fixture profiles and ran every condition through the complete generation-to-analysis system first.
Tradeoff
The first release validates the apparatus rather than answering the headline model question.
Impact
Pipeline failures and misleading detectors can be corrected before expensive evidence is collected.
Problem
A pattern count without retained artifacts is impossible to challenge.
Decision
Keep source, responsive renders, detector evidence, uncertainty, and blind-review state behind every aggregate.
Tradeoff
The corpus is heavier and the interface must support drill-down as well as summary.
Impact
A reader can move from a chart to the exact artifact that produced it.
Problem
One interface could not serve overview, detector debugging, and artifact comparison equally well.
Decision
Built three complementary surfaces: Observatory for the release view, Lab for pattern analysis, and Corpus for artifact inspection.
Tradeoff
The product has a steeper information architecture than a single dashboard.
Impact
Each mode answers a distinct question without discarding the shared evidence model.
Approach
1. Register the comparison
Six briefs, three fixture profiles, and three instruction conditions define the calibration matrix. Stable identifiers and manifests keep conditions comparable instead of relying on hand-picked screenshots.
2. Capture the full artifact
Each run retains source plus desktop and mobile renders. Playwright standardizes capture so responsive behavior can be inspected alongside the code that produced it.
3. Detect patterns with uncertainty
Static and visual detectors record their evidence rather than emitting unexplained labels. Ambiguous cases can enter a blind-review queue instead of being forced into a confident aggregate.
4. Separate calibration from claims
The product labels the current corpus as synthetic fixture evidence throughout the method and interface. Real provider findings remain a future, separately gated phase.
Outcome
The calibration release contains 108 deterministic artifacts across six briefs, three fixture profiles, and three instruction conditions. It demonstrates the complete measurement, rendering, detection, aggregation, and review pipeline while keeping the central frontier-model question explicitly unanswered.
Proof points
108
Calibration artifacts
6
Registered briefs
3
Evidence surfaces
Learnings
- →A benchmark is a product: its state labels, drill-down paths, and uncertainty language shape what readers believe.
- →Calibration can be a meaningful release when it proves the method and refuses the headline claim.
- →Design-pattern criticism becomes more useful when every disagreement has an artifact to inspect.
- →Security and design research share a discipline: preserve evidence, expose limits, and keep the human judgment visible.
Anti-Patterns Avoided
- ×Treating synthetic fixtures as evidence about current frontier models.
- ×Publishing pattern percentages without the artifacts and detector evidence behind them.
- ×Using visual taste as if it were a context-free quality metric.
- ×Spending provider compute before validating the protocol and failure modes.
Next Iterations
- →Run a minimal provider smoke test only after compute and credentials are explicitly approved.
- →Calibrate detector thresholds against blinded human review.
- →Publish real-model comparisons as a new evidence layer, not a rewrite of the fixture release.
Get In Touch
If you want to talk about similar work, email me.
Contact is the simplest place to start.
Next project
SkillScan