M1 Artificial Intelligence · semester 7 · acquire, clean, mine
Data Acquisition, Processing and Mining for AI
One pipeline, built twice. The weekly labs teach each stage on data the staff supply; the paired project then runs the same stages end to end on a live API, and has to finish on a result somebody can read.
Two tracks, one pipeline
Labs
A stage a week
Week 1 works 200 000 New York taxi trips in pandas: load the Parquet, mask out the impossible rows, derive duration and pickup hour, group and plot. Later weeks land as they are taught.
Project
Four milestones, marked out of 20
A fixed pair registers one API and one precise slice of it, then carries that slice from a raw pull through cleaning and exploration to a mining result, a report and a data card.
Reading it
Source
The folder on GitHub
Every lab in its own folder, the project beside them, and the completed notebooks once the work is done.
README
The course README
The table of labs, the four project milestones and what each one holds, and the one conda environment the whole course runs in.
No data collected for the project is published here, and neither is anything the staff distributed — the handout, the data card template, the setup guide, the lab notebooks and the sample datasets all stay on disk, out of the repository. What is committed is my own work.