From the workshop bench
AskOski ML & Data Pipeline
AskOski helps 45,000+ Berkeley students plan their academic paths, and it needs fresh data and fresh models to do it. I built the ML and data infrastructure behind it: Apache Airflow pipelines that retrain models automatically from the latest student-information-system data, a BERT course recommender refactored with PyTorch Transformers to sub-250ms inference, and a migration to the SIS V5 API so student data keeps flowing reliably as the university’s systems evolve. I also rebuilt the Docker pipeline as multi-stage, cache-free installs, cutting the image size 80% from 25GB to 5GB.