{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "0",
   "metadata": {},
   "source": [
    "# College hoops tiers\n",
    "\n",
    "**The brief:** a college basketball account wants an end-of-season Big Ten tier list, the kind fans argue about,\n",
    "but backed by a power rating instead of vibes, as a 1080 x 1080 post plus a 1200 x 675 cut. The rating is\n",
    "SportsDataverse's adjusted efficiency margin for men's college basketball through `sportsdataverse.mbb`, and\n",
    "`team_tiers` (sdvplot's port of sdvplotR's Tiermaker) draws the list."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1",
   "metadata": {},
   "outputs": [],
   "source": [
    "import tempfile\n",
    "from pathlib import Path\n",
    "\n",
    "import matplotlib.pyplot as plt\n",
    "import polars as pl\n",
    "import sportsdataverse.mbb as mbb\n",
    "from IPython.display import Image\n",
    "from PIL import Image as PILImage\n",
    "\n",
    "import sdvplot\n",
    "from sdvplot.matplotlib import team_tiers\n",
    "\n",
    "SEASON = 2026  # the 2025-26 season, named by the year it ends\n",
    "CONFERENCE = \"Big Ten Conference\"\n",
    "OUT = Path(tempfile.mkdtemp(prefix=\"sdvplot-recipe-\"))  # where the exports go; use your own folder"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "2",
   "metadata": {},
   "source": [
    "## 1. Get the data\n",
    "\n",
    "One row per team: adjusted offense, defense and their margin (points per 100 possessions better than an average\n",
    "Division I team, adjusted for opponents). The ratings carry ESPN team ids; sdvplot's team table supplies each\n",
    "team's conference and abbreviation for the same ids."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3",
   "metadata": {},
   "outputs": [],
   "source": [
    "ratings = mbb.load_mbb_ratings(SEASON)\n",
    "teams = sdvplot.teams(\"mbb\").select(\"team_id\", \"abbr\", \"short_name\", \"conference\")\n",
    "assert ratings.schema[\"team_id\"] == teams.schema[\"team_id\"]  # ESPN ids as strings on both sides\n",
    "league = (\n",
    "    ratings.join(teams, on=\"team_id\")\n",
    "    .filter(pl.col(\"conference\") == CONFERENCE)\n",
    "    .sort(\"adj_em\", descending=True)\n",
    "    .select(\"team_id\", \"abbr\", \"short_name\", \"rank\", \"adj_em\")\n",
    ")\n",
    "league"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4",
   "metadata": {},
   "source": [
    "## 2. The first draft\n",
    "\n",
    "Five tiers of equal size, straight into `team_tiers`."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "5",
   "metadata": {},
   "outputs": [],
   "source": [
    "equal = pl.int_range(pl.len()) * 5 // pl.len() + 1  # five tiers of (nearly) equal size\n",
    "draft = league.select(team=\"team_id\", tier_no=equal)\n",
    "fig = team_tiers(draft, \"mbb\")\n",
    "plt.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6",
   "metadata": {},
   "source": [
    "It looks like a tier list, but the tiers are wrong. Equal-sized groups put Michigan, the best team in the country,\n",
    "in the same tier as teams more than ten points worse per 100 possessions, and split neighbors who are barely a\n",
    "point apart. The default labels (\"Elite\" ... \"What are they doing?\") are jokes, not information.\n",
    "\n",
    "## 3. Let the ratings draw the lines\n",
    "\n",
    "Tiers should break where the ratings do. Sorting by rating and measuring the gap to the team above makes the\n",
    "widest gaps easy to find; cutting at the four widest gives five tiers whose members are close to each other."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "7",
   "metadata": {},
   "outputs": [],
   "source": [
    "league = league.with_columns(gap=pl.col(\"adj_em\").shift(1) - pl.col(\"adj_em\"))\n",
    "cut = league[\"gap\"].drop_nulls().sort(descending=True)[3]  # the fourth-widest gap\n",
    "league = league.with_columns(tier_no=(pl.col(\"gap\").fill_null(0) >= cut).cum_sum() + 1)\n",
    "league.filter(pl.col(\"gap\") >= cut).select(\"short_name\", \"adj_em\", \"gap\", \"tier_no\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "8",
   "metadata": {},
   "source": [
    "Each row above starts a new tier: Michigan stands alone, and the gaps at the top (6.6 points) and near the bottom\n",
    "(5.1) are the conference's real dividing lines.\n",
    "\n",
    "## 4. Labels that say something\n",
    "\n",
    "Each tier's label becomes the rating range of its members, so the list carries its own evidence, and the title states\n",
    "the finding. `tier_rank` keeps teams in rating order inside a tier. The subtitle explains the rating and the rule for\n",
    "the cuts; the caption names the data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "9",
   "metadata": {},
   "outputs": [],
   "source": [
    "tiers = (\n",
    "    league.group_by(\"tier_no\", maintain_order=True)\n",
    "    .agg(lo=pl.col(\"adj_em\").min(), hi=pl.col(\"adj_em\").max())\n",
    "    .sort(\"tier_no\")\n",
    ")\n",
    "tier_desc = {\n",
    "    row[\"tier_no\"]: f\"{row['hi']:+.1f}\" if row[\"lo\"] == row[\"hi\"] else f\"{row['hi']:+.1f} to {row['lo']:+.1f}\"\n",
    "    for row in tiers.iter_rows(named=True)\n",
    "}\n",
    "best = league.row(0, named=True)\n",
    "data = league.select(team=\"team_id\", tier_no=\"tier_no\", tier_rank=pl.int_range(1, pl.len() + 1).over(\"tier_no\"))\n",
    "\n",
    "\n",
    "def tier_list(height=0.12, alpha=0.8, theme=\"dark\"):\n",
    "    return team_tiers(\n",
    "        data,\n",
    "        \"mbb\",\n",
    "        title=f\"{best['short_name']} stood alone in the Big Ten\",\n",
    "        subtitle=f\"Tiers by adjusted efficiency margin, {SEASON - 1}-{SEASON % 100:02d}.\\n\"\n",
    "        \"A new tier starts at each of the four widest gaps.\",\n",
    "        caption=\"Margin: points per 100 possessions better than an average D-I team.\\n\"\n",
    "        \"Data: SportsDataverse adjusted ratings via sportsdataverse-py  |  sdvplot Tiermaker\",\n",
    "        tier_desc=tier_desc,\n",
    "        height=height,\n",
    "        alpha=alpha,\n",
    "        theme=theme,\n",
    "    )\n",
    "\n",
    "\n",
    "fig = tier_list()\n",
    "plt.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "10",
   "metadata": {},
   "source": [
    "## 5. Fix the contrast\n",
    "\n",
    "On the dark Tiermaker background, three logos almost vanish: Iowa's black hawk, Penn State's navy lion and Michigan\n",
    "State's dark green Spartan. Most college logos are drawn for a white page, so the fix is a light background:\n",
    "`theme=\"light\"` draws the tiers on white with dark labels and lines. The logos go to full opacity too (`alpha=1`);\n",
    "the default 0.8 softens them against the dark background but washes them out on white."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "11",
   "metadata": {},
   "outputs": [],
   "source": [
    "fig = tier_list(alpha=1, theme=\"light\")\n",
    "plt.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "12",
   "metadata": {},
   "source": [
    "## 6. Export at social sizes\n",
    "\n",
    "`team_tiers` returns an ordinary matplotlib figure with constrained layout, so `set_size_inches` re-lays it at each\n",
    "export size. The logos are a fraction of the panel's height, and the square's panel is taller but narrower than the\n",
    "wide cut's, so nine logos in one tier would collide there: the square gets a smaller `height`. Nine teams in a row\n",
    "is also why the wide 1200 x 675 cut is the better post here."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "13",
   "metadata": {
    "sdvplot_gallery": {
     "alt": "Tier list of the 2025-26 Big Ten men's basketball teams by adjusted efficiency margin, Michigan alone in the top tier, each tier labeled with its rating range",
     "title": "Big Ten men's basketball tier list"
    },
    "tags": [
     "gallery"
    ]
   },
   "outputs": [],
   "source": [
    "exports = {\"big_ten_tiers_1200x675.png\": ((8, 4.5), 0.12), \"big_ten_tiers_1080x1080.png\": ((7.2, 7.2), 0.075)}\n",
    "for name, (size, height) in exports.items():\n",
    "    fig = tier_list(height, alpha=1, theme=\"light\")\n",
    "    fig.set_size_inches(*size)\n",
    "    fig.savefig(OUT / name, dpi=150)\n",
    "    plt.close(fig)\n",
    "    print(name, PILImage.open(OUT / name).size)\n",
    "Image(OUT / \"big_ten_tiers_1200x675.png\", width=700)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "14",
   "metadata": {},
   "source": [
    "The 1080 x 1080 cut:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "15",
   "metadata": {},
   "outputs": [],
   "source": [
    "Image(OUT / \"big_ten_tiers_1080x1080.png\", width=540)"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python"
  },
  "sdvplot": {
   "description": "Turn a power rating into a Big Ten men's basketball tier list with team_tiers, tiers cut at the rating gaps, exported for social at 1080x1080 and 1200x675.",
   "label": "College hoops tiers",
   "position": 8
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
