Publications Search

Kokkos Kernels: FY20 update

Rajamanickam, Sivasankaran; Berger-Vergiat, Luc; Acer, Seher; Dang, Vinh Q.; Ellingwood, Nathan D.; Harvey, Evan C.; Kelley, Brian M.; Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Presentation YEAR 2021

DOI OSTI

Kokkos Kernels: FY20 update

Berger-Vergiat, Luc; Rajamanickam, Sivasankaran; Dang, Vinh Q.; Ellingwood, Nathan D.; Kelley, Brian M.; Harvey, Evan C.; Wilke, Jeremiah; Acer, Seher

Abstract not provided.

More Details

TYPE Conference Presentation YEAR 2021

DOI OSTI

Codesign for the Masses

Lewis, Cannada; Hammond, Simon; Wilke, Jeremiah

In this position paper we will address challenges and opportunities relating to the design and codesign of application specific circuits. Given our background as computational scientists, our perspective is from the viewpoint of a highly motivated application developer as opposed to career computer architects

More Details

TYPE Other Report YEAR 2021

DOI OSTI

Evaluating Trade-offs in Potential Exascale Interconnect Technologies

Hemmert, Karl S.; Bair, Ray; Bhatale, Abhinav; Groves, Taylor; Jain, Nikhil; Lewis, Cannada; Mubarak, Misbah; Pakin, Scott D.; Ross, Robert; Wilke, Jeremiah

This report details work to study trade-offs in topology and network bandwidth for potential interconnects in the exascale (2021-2022) timeframe. The work was done using multiple interconnect models across two parallel discrete event simulators. Results from each independent simulator are shown and discussed and the areas of agreement and disagreement are explored.

More Details

TYPE Other Report YEAR 2020

DOI OSTI

An Evaluation of Ethernet Performance for Scientific Workloads

Proceedings of INDIS 2020: Innovating the Network for Data-Intensive Science, Held in conjunction with SC 2020: The International Conference for High Performance Computing, Networking, Storage and Analysis

Kenny, Joseph; Wilke, Jeremiah; Ulmer, Craig; Baker, Gavin M.; Knight, Samuel; Friesen, Jerrold A.

Priority-based Flow Control (PFC), RDMA over Converged Ethernet (RoCE) and Enhanced Transmission Selection (ETS) are three enhancements to Ethernet networks which allow increased performance and may make Ethernet attractive for systems supporting a diverse scientific workload. We constructed a 96-node testbed cluster with a 100 Gb/s Ethernet network configured as a tapered fat tree. Tests representing important network operating conditions were completed and we provide an analysis of these performance results. RoCE running over a PFC-enabled network was found to significantly increase performance for both bandwidth-sensitive and latency-sensitive applications when compared to TCP. Additionally, a case study of interfering applications showed that ETS can prevent starvation of network traffic for latency-sensitive applications running on congested networks. We did not encounter any notable performance limitations for our Ethernet testbed, but we found that practical disadvantages still tip the balance towards traditional HPC networks unless a system design is driven by additional external requirements.

More Details

TYPE Conference Poster YEAR 2020

OSTI Scopus

Opportunities and limitations of Quality-of-Service in Message Passing applications on adaptively routed Dragonfly and Fat Tree networks

Proceedings - IEEE International Conference on Cluster Computing, ICCC

Wilke, Jeremiah; Kenny, Joseph

Avoiding communication bottlenecks remains a critical challenge in high-performance computing (HPC) as systems grow to exascale. Numerous design possibilities exist for avoiding network congestion including topology, adaptive routing, congestion control, and quality-of-service (QoS). While network design often focuses on topological features like diameter, bisection bandwidth, and routing, efficient QoS implementations will be critical for next-generation interconnects. HPC workloads are dominated by tightly-coupled mathematics, making delays in a single message manifest as delays across an entire parallel job. QoS can spread traffic onto different virtual lanes (VLs), lowering the impact of network hotspots by providing priorities or bandwidth guarantees that prevent starvation of critical traffic. Two leading topology candidates, Dragonfly and Fat Tree, are often discussed in terms of routing properties and cost, but the topology can have a major impact on QoS. While Dragonfly has attractive routing flexibility and cost relative to Fat Tree, the extra routing complexity requires several VLs to avoid deadlock. Here we discuss the special challenges of Dragonfly, proposing configurations that use different routing algorithms for different service levels (SLs) to limit VL requirements. We provide simulated results showing how each QoS strategy performs on different classes of application and different workload mixes. Despite Dragonfly's desirable characteristics for adaptive routing, Fat Tree is shown to be an attractive option when QoS is considered.

More Details

TYPE Conference Poster YEAR 2020

OSTI Scopus

Modern problems require modern solutions: How modern CMake supports modern C++ in Kokkos

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2020

OSTI

Opportunities and limitations of Quality-of-Service in Message Passing applications on adaptively routed Dragonfly and Fat Tree networks

Proceedings - IEEE International Conference on Cluster Computing, ICCC

Wilke, Jeremiah; Kenny, Joseph

Avoiding communication bottlenecks remains a critical challenge in high-performance computing (HPC) as systems grow to exascale. Numerous design possibilities exist for avoiding network congestion including topology, adaptive routing, congestion control, and quality-of-service (QoS). While network design often focuses on topological features like diameter, bisection bandwidth, and routing, efficient QoS implementations will be critical for next-generation interconnects. HPC workloads are dominated by tightly-coupled mathematics, making delays in a single message manifest as delays across an entire parallel job. QoS can spread traffic onto different virtual lanes (VLs), lowering the impact of network hotspots by providing priorities or bandwidth guarantees that prevent starvation of critical traffic. Two leading topology candidates, Dragonfly and Fat Tree, are often discussed in terms of routing properties and cost, but the topology can have a major impact on QoS. While Dragonfly has attractive routing flexibility and cost relative to Fat Tree, the extra routing complexity requires several VLs to avoid deadlock. Here we discuss the special challenges of Dragonfly, proposing configurations that use different routing algorithms for different service levels (SLs) to limit VL requirements. We provide simulated results showing how each QoS strategy performs on different classes of application and different workload mixes. Despite Dragonfly's desirable characteristics for adaptive routing, Fat Tree is shown to be an attractive option when QoS is considered.

More Details

TYPE Conference Poster YEAR 2020

OSTI Scopus

Putting compilers in the simulation co-design loop with surrogate performance models

Wilke, Jeremiah; Lewis, Cannada

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2020

OSTI

Modern problems require modern solutions: How modern CMake supports modern C++ in performance portability

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2020

OSTI

Supercomputer in a workstation: simulation as a development platform for network architectures

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2020

OSTI

Kokkos Kernels

Rajamanickam, Sivasankaran; Acer, Seher; Berger-Vergiat, Luc; Dang, Vinh Q.; Ellingwood, Nathan D.; Kelley, Brian M.; Kim, Kyungjoo; Trott, Christian R.; Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2020

OSTI

The Exascale Computing Project: Hardware Evaluation for Interconnects

Hemmert, Karl S.; Wilke, Jeremiah; Ross, Rob; Groves, Taylor; Karlin, Ian

Abstract not provided.

More Details

TYPE Presentation YEAR 2020

OSTI

Kokkos CMake: Build Systems and for Modern C++

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Presentation YEAR 2019

OSTI

Evaluation of novel interconnect technologies for ASC applications

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Presentation YEAR 2019

OSTI

Supercomputer in a workstation: simulation as a development platform for network architectures

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Presentation YEAR 2019

OSTI

Significant Vendor Impact of Sandia?s Portals Networking Technology

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Presentation YEAR 2019

OSTI

Batched Linear Algebra in Kokkos Kernels

Rajamanickam, Sivasankaran; Berger-Vergiat, Luc; Dang, Vinh Q.; Ellingwood, Nathan D.; Kim, Kyungjoo; Mclendon, William; Trott, Christian R.; Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2019

OSTI

Build Systems and Package Management for Modern C++

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Presentation YEAR 2019

OSTI

Verifying Simulator Readiness for Evaluating Potential Exascale Interconnect Technologies [PowerPoint]

Hemmert, Karl S.; Wilke, Jeremiah; Kenny, Joseph; Lewis, Cannada; Bhatele, Abhinav; Georgakoudis, Giorgis; Pakin, Scott; Mubarak, Misbah; Groves, Taylor

Goals of the milestone are to: verify key hardware contention models in controlled environment; validate simulator readiness for future milestones; and, provide baseline to define cross-validation workflow across teams for ''bracketing'' results.

More Details

TYPE Other Report YEAR 2019

DOI OSTI

The programming and performance challenges of memory-semantic interconnects for scientific computing

Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Conference Poster YEAR 2019

OSTI

Kokkos Kernels

Rajamanickam, Sivasankaran; Berger-Vergiat, Luc; Dang, Vinh Q.; Ellingwood, Nathan D.; Kim, Kyungjoo; Mclendon, William; Trott, Christian R.; Wilke, Jeremiah

Abstract not provided.

More Details

TYPE Presentation YEAR 2018

OSTI

ECP Milestone Memo for 2.3.1.04.14

Wilke, Jeremiah

The DARMA many-task framework provides asynchronous communication and load balancing functionality. This functionality is embedded in standard, modern C++ through the use of the template wrapper classes similar to futures. DARMA previously functioned as a single, large repository. This simplified building and installation, but hindered agile development as individual components could not be easily updated or reused in other projects. DARMA components can now be developed independently and reused in other ECP projects. Through Spack and modern CMake, a complete DARMA package can be easily configured and installed with automatic dependency management for each of the configuration options.

More Details

TYPE Other Report YEAR 2018

DOI OSTI