Multimodal Vision Agent Evaluation: UI/UX & Screenshot-to-Code
Screenshot-to-code systems should be evaluated as complete pipelines: image understanding, generated code, browser rendering, responsive behavior, accessibility, interaction, and repair loops. This reviewed protocol intentionally publishes no model leaderboard because the earlier draft's model names and scores were not supported by reproducible evidence.
1. Build a Representative Fixture Set
Use designs you have permission to evaluate and include multiple page shapes, densities, and breakpoints. Keep a hidden test subset so prompt tuning does not target every screenshot. Each fixture should include:
- Reference images at desktop and mobile sizes with exact viewport and device-pixel ratio.
- Required text, controls, states, assets, interactions, and accessibility expectations.
- Allowed framework, component library, fonts, icons, and asset licenses.
- Acceptance tolerances for layout, color, typography, overflow, and interaction.
- Known ambiguity notes where a screenshot cannot reveal behavior or responsive intent.
2. Score Multiple Quality Layers
| Layer | Example checks | Evidence | Common false positive |
|---|---|---|---|
| Visual | Geometry, spacing, color, type scale, asset placement | Rendered screenshot plus region-level diff | High global similarity hiding a broken button |
| Structural | Semantic landmarks, headings, labels, reusable components | DOM and accessibility-tree assertions | Pixel-perfect div soup |
| Behavioral | Keyboard use, focus, validation, loading, empty and error states | Browser interaction trace and assertions | Static screenshot passing while controls fail |
| Responsive | Breakpoints, wrapping, overflow, content order | Screenshots and DOM checks at declared viewports | Desktop-only optimization |
| Operational | Build success, console health, dependencies, bundle and repair cost | Build logs, lockfile diff, browser logs, total calls | Ignoring manual cleanup after generation |
3. Control the Generation Environment
Pin the model snapshot or public identifier, API version, prompt, system instructions, framework, dependency versions, browser, viewport, device scale, and maximum repair turns. Start every candidate from the same repository state. If an agent can inspect the rendered page and iterate, count all model calls, tool calls, elapsed time, and human edits.
Use the current provider documentation to construct image requests; request formats and model identifiers change. Store only images you are permitted to send, strip private metadata when appropriate, and understand provider retention terms before uploading customer designs.
4. Visual Metrics Need Human and DOM Context
SSIM or pixel differences can help locate mismatches, but animations, antialiasing, fonts, browser rendering, and large backgrounds can distort the score. Compare meaningful regions, mask approved dynamic areas, and pair image metrics with computed-style and DOM assertions. A reviewer should inspect the first viewport, long-page sections, interactive states, and mobile layout while blinded to the candidate name.
5. Interaction and Accessibility Gate
- Every visible control has a keyboard path and a programmatic name.
- Focus is visible and moves predictably through menus, dialogs, and forms.
- Headings and landmarks describe the page rather than merely matching font size.
- Text and controls remain usable at zoom and narrow widths.
- Loading, empty, validation, error, and success states are tested.
- Console errors, failed assets, hydration issues, and horizontal overflow fail the run.
6. Report Results Reproducibly
Publish per-fixture scores, median and worst-case results, confidence intervals where repeated runs exist, failure screenshots, and the complete configuration. Separate zero-shot generation from repaired output. Do not label one system the winner if it used more iterations, a different component library, private design metadata, or human correction.
7. Budget the Repair Loop
Many systems render, inspect, and revise their own output. Set a fixed maximum number of repair turns and preserve every intermediate screenshot and diff. Report zero-shot quality, final quality, total calls, elapsed time, and human edits separately. Otherwise a system that needed repeated expensive repair can appear equal to one that produced an acceptable first result.
8. Primary References
- W3C WCAG 2.2 quick reference
- Playwright visual comparisons
- OSWorld multimodal computer-use benchmark
9. Related Benchmark Matrix & Tools
What multimodal evaluation costs
Vision inputs are tokenised as images, so a screenshot-heavy evaluation consumes far more input than the equivalent text benchmark and the input line dominates. Below: one evaluated screenshot at 20K input tokens including the image, 8K cached, 4K output. Of the 20,000 input tokens, 8,000 are billed at the cache-read rate and 12,000 at full input rate.
| Model | Cost per evaluation | Monthly at 5,000 evaluations |
|---|---|---|
| Claude Sonnet 5 | $0.066 | $328 |
| Gemini 3.1 Pro | $0.074 | $368 |
| GPT-5.6 Sol | $0.131 | $656 |
Output is a third of the bill here, unusually high for an evaluation workload, because image-to-code tasks emit a lot. Cheaper image handling beats cheaper text: reducing resolution or cropping to the region of interest saves more than switching model.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.