Agenda

 

Agenda is subject to change. Times listed below are in Pacific.

Lesson materials

 

Tuesday, June 16 - Preparation Day (virtual)

TIme Topic Title Speaker(s)
9:00 am - 9:30 am 1.1. Welcome & Orientation Cindy Wong. Events Specialist
Mary Thomas, Computational Data Scientist & Director of the CIML Summer Institute 
9:30 am – 10:00 am 1.2 Accounts, Login, Environment, Running Jobs and Logging into Expanse User Portal Marty Kandes, Computational and Data Science Research Specialist

 

Tuesday, June 23 - HPC/Parallel Concepts (in person)

TIme Topic Title Speaker(s)
8:00 - 8:30 am Light Breakfast & Check-in
8:30 - 9:30 am 2.1 Welcome and Introductions Mai Nguyen, Lead for Data Analytics
9:30 am - 9:45 am Break  
9:45 am - 10:45 am

2.2 Parallel Computing Concepts
We will cover supercomputer architectures, the differences between threads and processes,

implementations  of parallelism (e.g., OpenMP and MPI), strong and weak scaling, limitations

on scalability (Amdahl’s and Gustafson’s Laws)  and benchmarking.

Andreas Goetz, Research Scientist & Principal Investigator
10:45 am - 11:45 am

2.3 Getting Started with Batch Job Scheduling
Batch job schedulers are used to manage and fairly distribute the shared resources of high-performance

computing (HPC) systems. Learning how to interact with them and compose your work into  batch

jobs is essential to  becoming an effective HPC user.

Marty Kandes, Computational and Data Science Research Specialist  
11:45 am - 1:00 pm Lunch Break  
1:00 pm - 2:15 pm

2.4 Data Management and File Systems
Managing data efficiently on a supercomputer is important from both users' and system's perspectives.

We will cover a few basic data management techniques and I/O best practices in the context

of the Expanse system at SDSC.

Marty Kandes, Computational and Data Science Research Specialist
2:15 pm - 3:45 pm 2.5 GPU Computing - Hardware architecture and software infrastructure
Brief overview of the massively parallel GPU architecture that enables large-scale deep learning
applications, access and use of GPUs on SDSC Expanse for ML applications
Andreas Goetz, Research Scientist & Principal Investigator
3:45 pm  - 4:00 pm Break  
4:00 pm - 5:30 pm

2.6 Software Containers for Scientific and High-Performance Computing
Singularity is an open-source container engine designed to bring operating system-level virtualization to scientific

and high-performance computing. With Singularity you can package complex computational workflows --- software

applications, libraries, and data --- in a simple, portable, and reproducible way, which can then be run almost anywhere.

Marty Kandes, Computational and Data Science Research Specialist
5:30 PM – 5:45 PM Q&A, Wrap-up  
6:00 PM - 7:30 PM

Reception

Kaleidoscope Rooftop (KA 1406)

 

Wednesday, June 24 - Deep Learning (in person)

TIme Topic Title Speaker(s)
8:00 am - 8:45 am Light Breakfast  
8:45 am - 10:15 am 3.1 Introduction to Neural Networks and Convolution Neural Networks
An overview of the main concepts of neural networks and feature discovery; the basic convolution neural network for digit recognition
Paul Rodriguez, Computational Data Scientist 
10:15 am - 10:30 am Break  
10:30 am - 11:30 am 3.2 Practical Guidelines for Training Deep Learning on HPC
Guildelines on running deep networks on Expanse, such as using notebooks, and batch jobs; also some discussion of multinode execution.
Paul Rodriguez, Computational Data Scientist 
11:30 am - 12:00 pm 3.3 Experiment Tracking
We will cover tools for tracking and organizing ML and DL experiments.
Mai Nguyen, Lead for Data Analytics 
12:00 pm - 1:30 pm Lunch
Group Photo
1:30 pm - 2:15 pm 3.4 Deep Learning Layers and Architectures
Overview of deep learning concepts, including layers, architectures, applications, and libraries.
Mai Nguyen, Lead for Data Analytics 
2:15 pm - 3:45 pm 3.5 Deep Learning Transfer Learning
Tutorial and hands-on exercises on the use of transfer learning and fine-tuning for efficient training of deep learning models.
Mai Nguyen, Lead for Data Analytics 
3:45 pm - 4:00 pm Break  
4:00 pm - 5:30 pm

3.6 Deep Learning – Special Connections and Transformers
The architecture of many networks use paths and connections in flexible ways; we will review gate, skip, and

residual connections and get some intuition about transformers.

Paul Rodriguez, Computational Data Scientist 

 

Thursday, June 25 - Scalable Machine Learning & Large Language Model (in person)

TIme Topic Title Speaker(s)

8:00 am - 8:30 am 

Light Breakfast  
8:30 am– 10:00 am

4.1 CONDA Environments and Jupyter Notebook on Expanse: Scalable & Reproducible Data Exploration and ML 

Set up reproducible and transferable software environments and scale up calculations to large datasets using parallel computing."

Marty Kandes, Computational and Data Science

Research Specialist

10:00 am – 10:15 am Break  
10:15 am - 10:52 am 4.2 Spark
Introduction to performing machine learning at scale, with hands-on exercises using Spark."
Mai Nguyen, Lead for Data Analytics
10:52 am - 11:30 am 4.3 Tools for Scaling
Overview of tools to distribute and parallelize processes."
Paul Rodriguez, Computational Data Scientist
11:30 am - 12:15 pm SDSC Data Center Tour (45-mins)  
12:15 pm - 1:45 pm Lunch  
1:45 pm -3:00 pm

4.4 LLM Overview
In this session we will present an introduction to Large Language Models and how they can be used to support

research This session is designed for people with a basic understanding of machine learning but no prior experience

with LLMs is required.

Mai Nguyen, Lead for Data Analytics 
Paul Rodriguez, Computational Data Scientist 

3:00 pm - 3:15 pm Break  
3:15 pm - 4:45 pm 4.4 LLM Overview -continued Mai Nguyen, Lead for Data Analytics 
Paul Rodriguez, Computational Data Scientist
4:45 pm - 5:15 pm NAIRR Introduction 

Mahidhar Tatineni, Director of User Services

5:15 pm - 5:30 pm Closing Remarks Mai Nguyen, Lead for Data Analytics