LAPIS: Linear Algebra Performance for Intermediate Subprograms

Experts in materials science or plasma physics want to focus on designing better simulations and more accurate models, instead of wrangling software dependencies and learning the advanced features of Kokkos. The LAPIS compiler helps in two ways. First, it lets scientists rapidly design and train machine-learning models in Python and then use those models in their C++ codes without interfacing between the two languages. Second, it lets users write custom algorithms in high-level Python and then turn them into expert-level Kokkos automatically. Unlike other linear algebra compilers, LAPIS processes an entire program at once and supports any sparse or dense tensor format. This means it can be used to generate any algorithm used in a modern simulation code. For example, LAPIS can generate a complete, efficient implementation of the preconditioned conjugate gradient (PCG) method in under one second. Writing the same code by hand could take a day. LAPIS supports both sparse and dense linear algebra, and the Kokkos code it generates can run on all supercomputer architectures.
This advancement was rooted in an Laboratory Directed Research & Development (LDRD) project.
Grayscale Lithography via Annealed Resin Engineering (GLARE)

Grayscale Lithography via Annealed Resin Engineering (GLARE) is a method for producing the world’s smallest, continuously varying grayscale images on transparent substrates for use as photolithography masks in standard semiconductor processing. GLARE is the first new practical true grayscale contact lithography technique developed in more than 30 years. It offers significant improvements over existing technologies by using non-specialty equipment, which broadens the potential market size and makes the technology more accessible to a wider range of users. The ability to produce masks quickly and affordably allows for rapid design iterations, avoiding the limitations of the traditional “first time right” approach that can stifle creativity and innovation. GLARE makes high-precision microfabrication accessible to a wider range of innovators, including those who may have previously been excluded due to high costs. Schematic representation of the GLARE process, illustrating the transformation of 2.5D polymer patterns into carbon masks that control light transmission.
Fenix: Rising from Failure

Fenix is a software fault-tolerance framework designed to enhance the performance and efficiency of large-scale computing. It utilizes existing resources and reduces the reliance on hardware reliability. This addresses the rising energy demands and environmental impacts associated with the modern computational power requirements of scientific computing and machine learning.
Fenix offers remarkable speed and scalability to applications running on the latest supercomputing hardware, enabling efficiency in complex code bases with minimal modification. When modelled against the latest failure rate data from Meta’s production AI clusters, Fenix’s in-memory checkpoints and PLR design is 58% faster than current GR solutions on 100K-GPU systems and 10 times faster on 1M-GPU systems. Fenix is a game-changing technology that will enable scientific computing applications to continue pushing the boundaries of science on hardware that is increasingly driven by AI compute demands to eschew reliability for greater performance and efficiency.
Pele Suite of Exascale Reacting Flow

The Pele suite of codes simulates turbulent reacting flows, which are critical to transportation, electricity generation, manufacturing, and national security, with unprecedented accuracy. Pele combines the software engineering needed to take advantage of exascale computing with a modular open-source architecture that allows integration with artificial intelligence so users can easily modify the code and to advance scientific discovery and de-risk technologies. National laboratory researchers, universities, and industry partners are using this capability to accelerate the development of new technologies and explore new scientific frontiers.
Led by National Laboratory of the Rockies with co-developers: Sandia National Laboratories; Lawrence Berkeley National Laboratory (LBNL); Oak Ridge National Laboratory (ORNL); Argonne National Laboratory (ANL); Lawrence Livermore National Laboratory (LLNL)
Persistent DynAMICS

Persistent DynAMICS is an intelligent sensing architecture that safeguards high-value assets by balancing cost and continuity inherent in large-scale monitoring. Rather than continuous, exhaustive data collection, Persistent DynAMICS adds a layer of decision-making on top of existing sensing systems delivering real-time activity tracking, improving detection sensitivity, and reducing false alarms. By adjusting sensor activity in real time, the system reduces unnecessary data collection, improves responsiveness, and makes more effective use of existing sensing infrastructure without requiring hardware replacement. The platform is designed for flexible deployment and integration so components can be containerized and incorporated into existing computing environments, enabling adoption without significant modification to existing hardware, software, or infrastructure.
Led by Los Alamos National Laboratory with co-developers: Sandia National Laboratories; Lawrence Livermore National Laboratory; Nevada National Security Site; Oak Ridge National Laboratory; Pacific Northwest National Laboratory