Sam Foreman
Sam Foreman
Computational Scientist @ Argonne National Laboratory. AI Group @ Leadership Computing Facility (ALCF)
Top Words
- training
- foundation
- models
- supercomputers
- aurora
- auroragpt
- model
- scientific
- scale
- building
All posts (24)
In ancient times1, back in ~ 2022โ2023, virtually all (production) PyTorch code was designed to run on NVIDIA GPUs. In April 2023, AMD announced day-zero support for PyTorch 2.0 within the ROCm 6.0 ecosystem, leveraging new features like TorchDynamo for performance gantt title AMD and Intel PyTorch Enablement...
Iโd like to try and post more this year. Ideally these would be less-polished, more-frequent updates on what Iโm thinking about / working on. Ongoing Projects AuroraGPT: Large Language Models for Scientific Applications on leadership-class supercomputers1. Additional details can be found in some of my recent talks:...
๐ง๐ปโ๐ป About Me ๐ก samforeman.me UIUC (2015): Engineering Physics + Applied Mathematics University of Iowa (2015โ2019): PhD. Physics1 ANL (2019โ2022): Postdoctoral Researcher ANL (2022โPresent): Assistant Computational Scientist Member of the AI/ML Group at ALCF Current Research: AuroraGPT: Foundation Models for...
๐ Simple Experiment to Compare Validation Loss Cool Down Comparison Note๐ W&B Report See W&B Report: Cooling Down Checkpoints for more details. โ๏ธ Cooling Down 256 Nodes of Aurora: Cooled down over last 10%: W&B Run: volcanic-blaze-4312 Explicit command: LR_DECAY_STYLE=constant \ OPT=ipex.fusedlamb \...
๐ง๐ปโ๐ป About Me ๐ก samforeman.me UIUC (2015): Engineering Physics + Applied Mathematics University of Iowa (2015โ2019): PhD. Physics1 ANL (2019โ2022): Postdoctoral Researcher ANL (2022โPresent): Assistant Computational Scientist Member of the AI/ML Group at ALCF Current Research: AuroraGPT: Foundation Models for...
๐ Distributed Training ๐ Scaling: Overview โ
Goal: Minimize: Cost (i.e. amount of time spent training) Maximize: Performance Note๐ Note See ๐ค Performance and Scalability for more details In this talk, we will explore the intricacies of training foundation models on supercomputers. We will discuss the architecture of...
๐ AERIS Figure 1: arXiv:2509.13523 ACM Gordon Bell Prize for Climate Modeling Finalist @ SCโ25 We demonstrate a significant advancement in AI weather and climate modeling with AERIS by efficient scaling of window-based transformer models. We have performed global medium-range forecasts with performance competitive...
Motivation When training on multiple data sources or domains, it is often desirable to smoothly interpolate between two distributions rather than switching abruptly. This ensures stable optimization and avoids sudden shifts in gradient statistics. We can achieve this with an annealing schedule that gradually shifts...
๐ Scaling: Overview โ
Goal: Minimize: Cost (i.e. amount of time spent training) Maximize: Performance Note๐ Note See ๐ค Performance and Scalability for more details ๐ข Training on a Single Device See also: Scientific AI at Scale: Distributed Training ๐ค Methods and tools for efficient training on a single GPU flowchart...
pbs-tui Figure 1: A terminal dashboard for monitoring PBS Pro schedulers ๐ Overview A terminal user interface built with Textual for monitoring PBS Pro schedulers at the Argonne Leadership Computing Facility. The dashboard surfaces job, queue, and node activity in a single view and refreshes itself automatically so...
Page 1 of 3