Table of Contents
The {sportsdataverse} Python Module
sportsdataverse-py 
See CHANGELOG.md for details.
The goal of sportsdataverse-py is to provide the community with a Python package for working with sports data as a companion to the cfbfastR, hoopR, and wehoop R packages. It delivers tidy, polars-powered access to play-by-play, box score, schedule, roster, and betting data across the NBA, WNBA, NFL, MLB, NHL, college football, and men's and women's college basketball.
Recent releases add a near drop-in nflreadpy-parity NFL surface (with its own caching + config layer), a cross-league ESPN wrapper architecture spanning hundreds of endpoints, and dedicated parser modules that turn raw ESPN, NHL, and MLB Stats API payloads into clean DataFrames. Every loader returns a polars.DataFrame by default and a pandas.DataFrame on request via return_as_pandas=True. Beyond data aggregation and tidying, the package also supports benchmarking open-source expected points and win probability metrics for American football.
Installation
sportsdataverse-py can be installed via pip:
pip install sportsdataverse
# with full dependencies
pip install sportsdataverse[all]or from the repo (which may at times be more up to date):
git clone https://github.com/sportsdataverse/sportsdataverse-py
cd sportsdataverse-py
pip install -e .[all]Our Authors
Citations
To cite the sportsdataverse-py Python package in publications, use:
BibTex Citation
@misc{gilani_sdvpy_2021,
author = {Gilani, Saiem},
title = {sportsdataverse-py: The SportsDataverse's Python Package for Sports Data.},
url = {https://py.sportsdataverse.org},
season = {2021}
}