March Madness AI Tournament

April 2026 · 4 min read


The Idea

I wanted to see what would happen if I had a few different AI models build me a March Madness bracket using nothing but stats, no seeding allowed. So I wrote a long, detailed prompt telling each model to act like an elite analyst and sports data scientist, ignore seed numbers completely, and instead weigh things like offensive and defensive efficiency, strength of schedule, recent momentum, injuries, and KenPom-style advanced metrics to fill out the entire bracket from the Round of 64 to the championship.

The Prompt

Here's the exact prompt I gave each model, word for word:

You are an elite college basketball analyst, sports data scientist, and betting strategist. Your task is to generate the most accurate possible March Madness bracket. IMPORTANT: Completely IGNORE official seeding. Do NOT use seed numbers as a factor. Instead, evaluate each matchup using a weighted model that includes: CORE FACTORS: ∙ Team efficiency metrics (Offensive Efficiency, Defensive Efficiency, Net Rating) ∙ Strength of Schedule (SOS) ∙ Recent performance (last 10 games, momentum) ∙ Injuries (key players out, limited, or returning) ∙ Head-to-head matchups (if applicable) ∙ Coaching experience and tournament history ∙ Player talent (NBA prospects, star power, depth) ∙ Turnover rate, rebounding rate, shooting splits (3PT%, FT%, eFG%) ∙ KenPom-style advanced metrics (simulate if needed) CONTEXTUAL FACTORS: ∙ Game location and travel distance ∙ Crowd advantage / pseudo-home games ∙ Altitude or environmental effects if relevant ∙ Matchup styles (e.g., fast vs slow pace, zone vs man defense) ∙ Upset probability based on statistical variance (3-point reliance, etc.) ADVANCED THINKING: ∙ Identify "fraud" teams (overrated records due to weak schedules) ∙ Identify "dangerous" lower-ranked teams (underseeded by public perception) ∙ Adjust for volatility (teams dependent on 3PT shooting are higher risk) ∙ Factor in tournament-specific trends (guard play, experience, defense wins) OUTPUT FORMAT: 1. Provide the FULL bracket from Round of 64 to Championship 2. For each game, briefly explain WHY that team wins (1–2 sentences max) 3. Clearly identify: ∙ Upsets (even though seeding is ignored, identify unexpected outcomes) ∙ Final Four teams ∙ National Champion 4. Include a short summary of your overall strategy and biggest edge GOAL: Maximize bracket accuracy and expected value — this should feel like a sharp bettor's bracket, not a casual fan's. Think step-by-step, but output only the final polished bracket and reasoning. Also simulate each matchup probabilistically (at least conceptually), and prefer teams with higher consistency and lower variance in later rounds. Avoid picking too many high-variance teams deep into the tournament unless justified. Prioritize guard play, free throw shooting, and defensive efficiency in late rounds.

The Setup

I ran the exact same prompt through Claude, ChatGPT, and Gemini to see how differently they'd think. Each one had to explain its reasoning for every matchup, call out upsets, and name a Final Four and a national champion. The goal was to see whether stripping out seeding would lead to wildly different brackets or if the models would mostly converge on the same teams.

How They Did

Claude and Gemini landed on a lot of the same teams in the back half of the bracket, while ChatGPT drifted toward a few different picks in the earlier rounds. Claude ended up taking Arizona as its national champion, while Gemini went with Duke. It was interesting watching three models reach different conclusions from the exact same instructions and the exact same data.

Why I Think They Diverged

My guess is that ChatGPT's earlier-round differences came from weighting recent momentum and "fraud team" detection more heavily than the other two, which led it to fade a couple of teams that Claude and Gemini kept alive. The Arizona-versus-Duke split at the top is the more interesting one. Both models had access to the same efficiency-style reasoning, but Claude seemed to lean harder into Arizona's defensive efficiency and guard play down the stretch, which the prompt explicitly told it to prioritize in late rounds. Gemini, on the other hand, seemed to weight Duke's overall talent and coaching experience more, treating those as the safer, lower-variance pick for a championship game. Neither read is "wrong" given the prompt — it's just a good example of how two models can apply the same criteria and still land on different champions depending on which factors they quietly prioritize.

What I'd Do Differently

Next time I'd push each model harder on its reasoning for the deep tournament picks, since that's where the brackets diverged the most. It'd also be fun to score all three brackets against the real results once the tournament wraps up and see which model's "ignore the seeds" approach actually held up.