Eval Results
eval-results
Pass rates per test case across two runs, with the regressions surfaced first and a sample size honest enough to say when a delta means nothing.
More from @scrimui
Render an AI reply token by token like ChatGPT's typing effect — blinking cursor, stop button, and no layout jump as the text grows.
What to show when a generation fails — a plain-English reason, a retry button, and a countdown for rate limits.
The loading state before the first token — bouncing dots, a blinking caret, or a labeled status line while the model thinks.
Upload a document, watch fields fill in, then review the flagged ones — per-field confidence, corrections that keep the original, export earned.
Switches for the tools a model may use — web search, code execution, file access — with unmistakable on and off states.
The result card for AI-generated images, audio and video — queue and stage states, variants, download and regenerate, and a safety block that isn't an error.
A voice-first conversation — live waveform states, a recording input, a spoken transcript and a typed fallback.
Render a React component from a tool result instead of text — a skeleton in the widget's own shape, and a text fallback when the client has no renderer.