CS352: Machine Perception of Music and Audio
Northwestern University (Updated for: Summer-2026)
| Top | Calendar | Links | Readings | Example Projects |
Course Description
This course covers machine extraction of structure in audio files covering areas such as source separation (unmixing audio recordings into individual component sounds), sound object recognition (labeling sounds), melody tracking, beat tracking, and perceptual mapping of audio to machine-quantifiable measures.
This course is approved for the Breadth Interfaces & project requirement in the CS curriculum.
Prior programming experience sufficient to be able to do laboratory assignments in PYTHON, implementing algorithms and using libraries without being taught to do so (there is no language instruction on Python). Having taken EECS 211 and 214 would demonstrate this experience.
Course Textbook
Fundamentals of Music Processing
Time & Place
Lecture: Mon, Wed, 11am - 12:30pm CST in Technological Institute M164
Instructors & Office Hours
Annie Chu 1-2pm Mondays in Mudd 3202 or by appointment
Course Policies
Questions outside of class
Please use CampusWire for class-related questions.
Grading Policy
You will be graded on a 100 point scale (e.g. 93 to 100 = A, 90-92 = A-, 87-89 = B+, 83-86 = B, 80-82 = B-…and so on).
There are 4 core axes you will be graded on: assignments (45%), class participation (15%), midterm (10%), final project (30%)
- Assignments (individual) - jupyter notebook coding homework assignments. There are 4 assignments, I will drop your lowest assignment. This means you can skip any one assignment.
- Class participation - in-class paper reading activities and initial + mid check ins
- Final Project (groups of 2-3) - make/analyze/implement something you’re interested in!
Homework and reading assignments are solo assignments and must be your original work. There will occasionally be readings or exercises to prep for in-class work. Both readings and advanced topic lectures will be defined by student interest and finalized by end of Week 2. Topics of interest may include generative modeling for music/speech/SFX, audio-language models + understanding, corpus studies, ethics, etc.
AI policy
You are expected to write your own code and write up your own answers to questions (not ChatGPT or Gemini or Copilot). This is an optional class you are (presumably) taking because you’re interested.
Submitting assignments
Assignments must be submitted on the due date by the time specified on Canvas. If you are worried you can’t finish on time, upload a safety submission an hour early with what you have. I will grade the most recent item submitted before the deadline. Late submissions will not be graded.
Course Calendar
Bring headphones to class! Many class activities will involve listening to and/or recording sound.
(*)means subject to change based on student interest
| Week | Date | Topic | In Class Activity (updated per class) | Assignment Due | Points |
|---|---|---|---|---|---|
| 1 | Mon June 22 | Course intro, Recording Basics | Audio 101: Who Let the Dogs Out? Meeting Sign Up: LINK | ||
| 1 | Wed June 24 | Frequency & Pitch, Tuning Systems | Intake Assignment (due Friday June 26) | 5 (of 15) | |
| 2 | Mon June 29 | Amplitude & Loudness | |||
| 2 | Wed July 1 | Fourier Transforms & Spectrograms (on Zoom) | |||
| 3 | Mon July 6 | Convolution & Filtering | Convolution & FFT notebooks | HW 1 Audio Basics | 15 |
| 3 | Wed July 8 | Advanced Filtering: Source Separation w/ REPET | Paper Reading Activity | Read REPET paper before class | 5 (of 15) |
| 3 | Mon July 13 | MFCCs and Chromagrams | MFCC + Chroma notebooks | ||
| 4 | Wed July 15 | Self-Similarity | HW 2 Spectrograms, Masking | 15 | |
| 4 | Mon July 20 | Pitch Tracking + Midterm Review | Sign up for Midterm Meeting | ||
| 5 | Wed July 22 | Basic Classifiers (Sound Object Labeling) | HW 3 Infinite Jukebox | 15 | |
| 5 | Mon July 27 | Midterm & Final Project Overview | Midterm Meeting | 10 & 5 (of 15) | |
| 6 | Wed July 29 | Embeddings w/ Primer of Deep Learning & Autoencoders | Embeddings Notebook | ||
| 6 | Mon Aug 3 | Guest Lecture: Ethics in Music AI (Julia Barnett) + Project Proposals | |||
| 7 | Wed Aug 5 | Building Interactive Music Systems (HCI for Musicking) | HW 4 Using Embeddings (due Fri Aug 7) | 15 | |
| 7 | Mon Aug 10 | Workshopping Proposals | Project Proposal Due (EOD) | 5 (of 30) | |
| 8 | Wed Aug 12 | Generative Approaches to Controllable Sound Transformations + Project Workshop | |||
| 8 | Mon Aug 17 | Zoom meetings with project groups (no class: meetings by appointment) | Project Meeting 1 | 5 of (30) | |
| 9 | Wed Aug 19 | — async work on final projects – | |||
| 10 | Mon Aug 24 | Zoom meetings with project groups (no class: meetings by appointment) | Project Meeting 2 | 5 of (30) | |
| 10 | Wed Aug 26 | FINAL PROJECT PRESENTATIONS (on Zoom) | Deliverables (15 of 30) |
Course Reading
Fundamentals of Music Processing, Chapter 1
Fundamentals of Music Processing, Chapter 2 & Section 3.1
Fundamentals of Music Processing, Chapter 4
Fundamentals of Music Processing, Chapter 6
Fundamentals of Music Processing, Chapter 7
* REPET for Background/Foreground Separation in Audio
Chapter 4 of Machine Learning : This is Tom Mitchell’s book. Historical overview + explanation of backprop of error. It’s a good starting point for actually understanding deep nets.
Yin: a fundamental frequency estimator for speech and music - This is, perhaps, the most popular pitch tracker.
Crepe: A Convolutional Representation for Pitch Estimation - A deep learning pitch tracker that improves on Yin.
The dummy’s guide to MFCC - an easy, high-level read. Start with this.
From Frequency to Quefrency: A History of the Cepstrum - a historical analysis of the uses of cepstrums
Recovering sound sources from embedded repetition - This is a paper on how humans actually listen to and parse audio based on repetition. Read any time.
Helpful Links
Some CS352 Projects (Mentor: A.Chu)
Live Neural Harmonizer, Summer 2026
DIY Autotune, Summer 2026
Vowel-based Emotion Editing, Summer 2026
Benchmarking Neural vs Real Amps, Summer 2026
TimbreTune, Winter 2025
Other Places to get ideas
EECS 352 Final projects from 2017 and 2015
Facebook’s Universal Music Translation
A coursera corse on pitch tracking
Datasets
U of Iowa’s Music Instrument Samples Dataset
The SocialFX data set of word descriptors for audio
VocalSketch: thousands of vocal imitations of a large set of diverse sounds
Bach10: audio recordings of each part and the ensemble of ten pieces of four-part J.S. Bach chorales
Software
Python Utilities for Detection and Classification of Acoustic Scenes
Librosa audio and music processing in Python
Essentia: an open source music analysis toolkit includes a bunch of feature extractors and pre-trained models for extracting e.g. beats per minute, mood, genre, etc.
Yaafe - audio features extraction toolbox
Sonic Visualizer music viz software
Lily Pond, open source music notation software
SoundSlice guitar tab and notation website
| Top | Calendar | Links | Readings |