CSE 60772 - High Performance Distributed Systems

CSE 60772 - High Performance Distributed Systems

View the Project on GitHub

High performance distributed systems are used to execute scientific applications to enable discovery at very large scales. This class will explore the software and hardware principles needed to design high performance applications from bottom to top. We will explore low level programming of vector, multicore, and GPU systems, and then scale up to use large clusters through distributed programming models. A final project will tie it all together by evaluating and optimizing a scientific application at scale.

Instructors

Prof. Douglas Thain (dthain@nd.edu)
Office Hours: TBA
Office: 206B Cushing
Grad TA: Jin Zhou (jzhou24@nd.edu)
Office Hours: TBA
Office: 150B Fitzpatrick Hall (Student Commons)

Course Organization

Computing Resources

Course Schedule

Week Monday Wednesday Friday
Aug 24 Class Overview
Slides
The Whole Stack SIMD Parallelism
Reading: SIMD
A0: Warmup Due
Aug 31 Multicore Hardware
Reading: OpenMP
OpenMP OpenMP
A1: Vector Parallel Due
Sep 7 Perf Eval Perf Eval Short Demos 1
Sep 14 GPU Architecture
Reading: CUDA
CUDA Programming CUDA Programming
A2: Multicore Evaluation Due
Sep 21 CUDA Programming CUDA Programming Short Demos 2
Sep 28 CPU or GPU CPU and GPU Catch Up
A3: GPU Evaluation
Oct 5 Cluster Fundamentals HTCondor HTCondor
Project Proposal Due
Oct 12 Workflow Systems Makeflow and Work Queue Short Demos 3
A3: GPU Evaluation Due
Oct 19 Fall Break
Oct 26 TaskVine Parsl FaaS
Nov 2 Software Environments
A4: Workflow Evaluation Due
Pip, Conda, and Friends Containers
Nov 9 Storage Performance
Progress Report Due
Parallel Filesystem Wide Area Data
Nov 16 Supercomputing Week Short Demos 4
Nov 23 Beautiful Data Thanksgiving Break
Nov 30 Project Presentations
Dec 7 Project Presentations Project Presentations
Final Report Due
Final Exam, 7:30PM

Readings and Reference