Lab
Experiments
Questions about generative UI we are testing in the open. Each one becomes a week of posts on X and a write-up here when it is done.
-
Grade the effect, not the pixels
Two identical screens that do very different things. What does it take to tell them apart before a user finds out?
-
How many tools can an agent hold?
Agents get worse as you give them more tools. We're measuring where that happens and which few tools deserve a place.
-
What the agent sees
An overlay that shows a page the way an agent reads it. Try it on the home page
-
Interfaces that learn from their verdicts
Generate many, test all of them, keep the best, and train on the difference. Can a small model learn one company's taste?
-
The reversibility ladder
Undo, support can undo, costly to undo, can't be undone. How much confirmation should each step up require?
-
Answers that aren't paragraphs
When should an assistant reply with an interface instead of text? We're looking for the line, and for when it moves.