March Madness AI Tournament
The Idea
I wanted to see what would happen if I had a few different AI models build me a March Madness bracket using nothing but stats, no seeding allowed. So I wrote a long, detailed prompt telling each model to act like an elite analyst and sports data scientist, ignore seed numbers completely, and instead weigh things like offensive and defensive efficiency, strength of schedule, recent momentum, injuries, and KenPom-style advanced metrics to fill out the entire bracket from the Round of 64 to the championship.
The Prompt
Here's the exact prompt I gave each model, word for word:
The Setup
I ran the exact same prompt through Claude, ChatGPT, and Gemini to see how differently they'd think. Each one had to explain its reasoning for every matchup, call out upsets, and name a Final Four and a national champion. The goal was to see whether stripping out seeding would lead to wildly different brackets or if the models would mostly converge on the same teams.
How They Did
Claude and Gemini landed on a lot of the same teams in the back half of the bracket, while ChatGPT drifted toward a few different picks in the earlier rounds. Claude ended up taking Arizona as its national champion, while Gemini went with Duke. It was interesting watching three models reach different conclusions from the exact same instructions and the exact same data.
Why I Think They Diverged
My guess is that ChatGPT's earlier-round differences came from weighting recent momentum and "fraud team" detection more heavily than the other two, which led it to fade a couple of teams that Claude and Gemini kept alive. The Arizona-versus-Duke split at the top is the more interesting one. Both models had access to the same efficiency-style reasoning, but Claude seemed to lean harder into Arizona's defensive efficiency and guard play down the stretch, which the prompt explicitly told it to prioritize in late rounds. Gemini, on the other hand, seemed to weight Duke's overall talent and coaching experience more, treating those as the safer, lower-variance pick for a championship game. Neither read is "wrong" given the prompt — it's just a good example of how two models can apply the same criteria and still land on different champions depending on which factors they quietly prioritize.
What I'd Do Differently
Next time I'd push each model harder on its reasoning for the deep tournament picks, since that's where the brackets diverged the most. It'd also be fun to score all three brackets against the real results once the tournament wraps up and see which model's "ignore the seeds" approach actually held up.