Agentic coding challengesProve you can drive the agent.
kodwai is a platform where developers solve real coding challenges on their own machine with their own AI coding agent (Claude Code, Cursor, or Codex).
- free for developers
- runs on your machine
- claude code · cursor · codex
Same agent.Different driver.
Same challenge, same agent. One developer pasted the spec and walked away. The other planned, tested and pushed back. The tests barely differ. The scores don't.
| bookshelf-rest-api | hands-off | hands-on |
|---|---|---|
| tests passed | 11/12 | 14/14 |
| traps caught | 1/3 | 3/3 |
| direction | 11/45 | 41/45 |
| total | 44 | 90 |
Two commands.One score.
No browser sandbox. Your machine, your editor, the agent you already pay for.
- CH.01in your terminal
Start
Pick a challenge and run one command. The CLI downloads the problem, starter files and tests, then starts the timer.
- CH.02in your editor
Solve
Work the way you normally do. Plan, prompt, test and push back when the agent gets it wrong.
- Claude Code
- Cursor
- Codex
- CH.03in your terminal
Submit
Send your code, commits, test runs and the agent transcript. You get a score out of 100 and the moments that moved it.
Scored on howyou direct the agent.
Passing tests is not enough. Direction weighs the most, because it is the part one lucky prompt can't fake.
How you steer, check and break down the work.
- Spec Precision
- Verification Rigor
- Decomposition
- Recovery
- Intent Fidelity
- Engagement
What shipped: tests and code quality, or the challenge's own rubric.
- Tests
- Code Quality
- Complexity
The hidden traps you caught, and how far you beat a solo AI.
- Edge-Case Coverage
- Lift over AI
Default weights shown. Challenges with their own rubric use 45 / 45 / 10.
Then your session, annotated.
Every scored run lists up to 6 moments that moved your score, chess style, each with its evidence. If you made mistakes, at least 2 of them stay on the list.
- !! brilliant
- ✓ good
- ?! dubious
- ?? miss
Real tasks.Hidden traps.
Realistic tasks with real tests and a few hidden traps: requirements that are easy to miss if you just paste the spec.
15 challenges · 3 easy · 3 medium · 9 hard
Bookshelf REST API
Junior Backend Engineer interview. Build a small REST API from scratch with CRUD, filters, validation, persistence, and a test suite that would catch a regression next sprint.
New here? Start with this one. It is the easy one.
A race you canwin this week.
A new league starts every Monday. Your first scored run puts you in a group of up to 30. The top 5 move up a division, the bottom 5 move down.
- Bronze
- Silver
- Gold
- Platinum
- Diamond
- Master
- 14 runs341.5
- 23 runs296
- 33 runs270
- 42 runs251.5
- 53 runs236
- 62 runs204
- 72 runs187.5
- 82 runs150
- 91 run131
- 101 run96
- 111 run88.5
- 121 run72
- 131 run60
- 141 run45
Leagues reset.Tiers stick.
Tiers follow your Direction Rating.
Your tier comes from your Direction Rating, an Elo-style number based on how well you direct the agent. Leagues start over every week, but your rating and tier carry over. Grandmaster starts at 1800.
- Bronze0 to 999
- Silver1000 to 1149
- Gold1150 to 1299
- Platinum1300 to 1449
- Diamond1450 to 1599
- Master1600 to 1799
- Grandmaster1800+
minimum direction rating for each tier · the grandmaster ring uses all three axis colors
Every scored run gets a share card with your total, your three axis scores and your best moment.
Agentic coding (also called AI-native coding, or vibe coding) is building software by directing an AI agent instead of typing every line. kodwai is a platform where developers solve real coding challenges on their own machine with their own AI coding agent (Claude Code, Cursor, or Codex) and get scored on how well they direct the agent, across three axes: Direction, Outcome, and Lift.
Claude Code, Cursor or Codex. The CLI asks which one. You solve on your own machine, and the CLI only collects from the challenge folder.
Yes for developers. Your first 3 submissions are scored on our key. After that you add your own Anthropic API key in Settings for unlimited runs. It is encrypted and only used to score your work.
Not really. Direction weighs the most, and a session with no spec, no checks and no redirects scores low there however green the tests are. The hidden traps catch the rest.
A new league starts every Monday at 00:00 UTC. Your first scored run puts you in a group of up to 30. Points are your best score on each challenge, times 1, 1.5 or 2 for easy, medium or hard. The top 5 move up a division and the bottom 5 move down.
Skip the puzzles.Build something real.
Pick a challenge, solve it with your agent and see your score. Your first 3 submissions are scored for free.
Start a challenge