powRSS
Home Recent Glimpses Shuffle Blogs Random

Sam Foreman

Sam Foreman

Computational Scientist @ Argonne National Laboratory. AI Group @ Leadership Computing Facility (ALCF)

Top Words

  • training
  • foundation
  • models
  • supercomputers
  • aurora
  • auroragpt
  • model
  • scientific
  • scale
  • building

All posts (24)

In ancient times1, back in ~ 2022โ€“2023, virtually all (production) PyTorch code was designed to run on NVIDIA GPUs. In April 2023, AMD announced day-zero support for PyTorch 2.0 within the ROCm 6.0 ecosystem, leveraging new features like TorchDynamo for performance gantt title AMD and Intel PyTorch Enablement...
Iโ€™d like to try and post more this year. Ideally these would be less-polished, more-frequent updates on what Iโ€™m thinking about / working on. Ongoing Projects AuroraGPT: Large Language Models for Scientific Applications on leadership-class supercomputers1. Additional details can be found in some of my recent talks:...
๐Ÿง‘๐Ÿปโ€๐Ÿ’ป About Me ๐Ÿก samforeman.me UIUC (2015): Engineering Physics + Applied Mathematics University of Iowa (2015โ€“2019): PhD. Physics1 ANL (2019โ€“2022): Postdoctoral Researcher ANL (2022โ€“Present): Assistant Computational Scientist Member of the AI/ML Group at ALCF Current Research: AuroraGPT: Foundation Models for...
๐Ÿ“‰ Simple Experiment to Compare Validation Loss Cool Down Comparison Note๐Ÿ“‘ W&B Report See W&B Report: Cooling Down Checkpoints for more details. โ˜ƒ๏ธ Cooling Down 256 Nodes of Aurora: Cooled down over last 10%: W&B Run: volcanic-blaze-4312 Explicit command: LR_DECAY_STYLE=constant \ OPT=ipex.fusedlamb \...
๐Ÿง‘๐Ÿปโ€๐Ÿ’ป About Me ๐Ÿก samforeman.me UIUC (2015): Engineering Physics + Applied Mathematics University of Iowa (2015โ€“2019): PhD. Physics1 ANL (2019โ€“2022): Postdoctoral Researcher ANL (2022โ€“Present): Assistant Computational Scientist Member of the AI/ML Group at ALCF Current Research: AuroraGPT: Foundation Models for...
๐ŸŒ Distributed Training ๐Ÿš€ Scaling: Overview โœ… Goal: Minimize: Cost (i.e. amount of time spent training) Maximize: Performance Note๐Ÿ“‘ Note See ๐Ÿค— Performance and Scalability for more details In this talk, we will explore the intricacies of training foundation models on supercomputers. We will discuss the architecture of...
๐ŸŒŽ AERIS Figure 1: arXiv:2509.13523 ACM Gordon Bell Prize for Climate Modeling Finalist @ SCโ€™25 We demonstrate a significant advancement in AI weather and climate modeling with AERIS by efficient scaling of window-based transformer models. We have performed global medium-range forecasts with performance competitive...
Motivation When training on multiple data sources or domains, it is often desirable to smoothly interpolate between two distributions rather than switching abruptly. This ensures stable optimization and avoids sudden shifts in gradient statistics. We can achieve this with an annealing schedule that gradually shifts...
๐Ÿ‘€ Scaling: Overview โœ… Goal: Minimize: Cost (i.e. amount of time spent training) Maximize: Performance Note๐Ÿ“‘ Note See ๐Ÿค— Performance and Scalability for more details ๐Ÿข Training on a Single Device See also: Scientific AI at Scale: Distributed Training ๐Ÿค— Methods and tools for efficient training on a single GPU flowchart...
pbs-tui Figure 1: A terminal dashboard for monitoring PBS Pro schedulers ๐Ÿ‘€ Overview A terminal user interface built with Textual for monitoring PBS Pro schedulers at the Argonne Leadership Computing Facility. The dashboard surfaces job, queue, and node activity in a single view and refreshes itself automatically so...

Page 1 of 3