
Waitlist open
AI engineering moved to the systems around the model. Same model for everyone. A real environment. A hidden eval. Improve the loop until the number goes up.
Leave your email. We will send a seat when the playground is up.
How it works
The starter is already in the editor, with traces that fail. You do not burn a run just to watch it lose. Read those, then change the loop.
System prompt, every tool call, every result, the grader verdict. This is how you debug a harness.
Thirty hidden tasks. Same model. Pass is 85 percent. A careful loop lands higher, with tokens still in budget.
One email. We send the link when the playground is up.
FAQ
The software around the model: the loop, the tools, the context, retries, when to stop. Not the model.
If everyone picks a different model, you cannot tell who built the better loop. Same model, same environment, same hidden eval. The only variable is the harness.
The full transcript of one task: system prompt, every message, every tool call and result, the final answer, the grader verdict. You debug from this, not from the score.
When the eval environment is hosted. We will email you a seat.
Yes for the waitlist and the first problems. A cheap pinned model. Run is three tasks. Submit is capped so nobody burns the bill.