Research Labs
All Research Labs
3D Deep Learning
Applied Research
Autonomous Vehicles
Deep Imagination
Publications
AI Playground
New and Featured
AI Art Gallery
NGC Demos
Research Areas
AI & Machine Learning
3D Deep Learning
Computer Vision
Robotics
All Areas
Careers
Academic Collaborations
Government Collaborations
Graduate Fellowship
Internships
Research Openings
Research Scientists
Meet the Team
Licensing
Skip to main content
Artificial Intelligence Computing Leadership from NVIDIA
Login
Research Labs
All Research Labs
3D Deep Learning
Applied Research
Autonomous Vehicles
Deep Imagination
Publications
AI Playground
New and Featured
AI Art Gallery
NGC Demos
Research Areas
AI & Machine Learning
3D Deep Learning
Computer Vision
Robotics
All Areas
Careers
Academic Collaborations
Government Collaborations
Graduate Fellowship
Internships
Research Openings
Research Scientists
Meet the Team
Licensing
Search
Search
Enter the terms you wish to search for.
Publications
Our publications provide insight into some of our leading-edge research.
Filters
Search
Apply
Filters
Filters
Publication Year
2025
(7)
2024
(1)
2023
(5)
2022
(5)
2021
(9)
2020
(6)
2019
(15)
2018
(14)
2017
(15)
2016
(9)
2015
(4)
2014
(1)
2012
(4)
2011
(4)
2010
(3)
2009
(2)
2008
(1)
2005
(1)
Facet Publication Year
Research Areas
High Performance Computing
(23)
Computer Architecture
(8)
Programming Languages, Systems and Tools
(6)
Resilience and Safety
(5)
Algorithms and Numerical Methods
(3)
Artificial Intelligence and Machine Learning
(3)
Computer Graphics
(3)
Networking
(3)
Real-Time Rendering
(3)
Autonomous Vehicles
(1)
Climate Simulation
(1)
Events
No Results Available
23 results found
High Performance Computing
Clear all
2021
2018
High Performance Computing
2021
GPS: A Global Publish-Subscribe Model for Multi-GPU Memory Management
Harini Muthukrishnan
,
Daniel Lustig
,
David Nellans
, Thomas Wenisch
Best Paper nominee
IEEE Micro Top Picks in Computer Architecture (Honorable Mention)
EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal in GPUs
Seung Won Min, Vikram Sharma Mailthody, Zaid Qureshi, Jinjun Xiong, Eiman Ebrahimi, Wen-mei Hwu
Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture
Seung Won Min, Kun Wu, Sitao Huang, Mert Hidayetoglu, Jinjun Xiong, Eiman Ebrahimi, Deming Chen,
Wen-mei Hwu
Suraksha: A Quantitative AV Safety Evaluation Framework to Analyze Safety Implications of Perception Design Choices
Hengyu Zhao,
Siva Hari
, Timothy Tsai,
Michael B. Sullivan
,
Steve Keckler
, Jishen Zhao
Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained Transfers
Harini Muthukrishnan
,
David Nellans
,
Daniel Lustig
, Jeffrey Fessler, Thomas Wenisch
Demystifying GPU Reliability: Comparing and Combining Beam Experiments, Fault Simulation, and Profiling
Fernando Fernandes dos Santos,
Siva Hari
, Pedro Martins Basso, Luigi Carro, Paolo Rech
Learning Sparse Matrix Row Permutations for Efficient SpMM on GPU Architectures
Atefeh Mehrabi,
Donghyuk Lee
,
Niladrish Chatterjee
, Danial J. Sorin, Benjamin C. Lee,
Mike O'Connor
Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture
Seung Won Min, Kun Wu, Sitao Huang, Mert Hidayetoglu, Jinjun Xiong, Eiman Ebrahimi, Deming Chen,
Wen-mei Hwu
Scaling Implicit Parallelism via Dynamic Control Replication
Michael Bauer
, Wonchan Lee, Elliott Slaughter, Zhihao Jia, Mario Di Renzo, Manolis Papadakis, Galen Shipman, Patrick McCormick,
Michael Garland
, Alex Aiken
2018
Dynamic Tracing: Memoization of Task Graphs for Dynamic Task-based Runtimes
Wonchan Lee, Elliott Slaughter,
Michael Bauer
, Sean Treichler, Todd Warszawski,
Michael Garland
, Alex Aiken
Exascale Deep Learning for Climate Analytics
Thorsten Kurth, Sean Treichler, Joshua Romero, Mayur Mudigonda, Nathan Luehr, Everett Phillips, Ankur Mahesh, Michael Matheson, Jack Deslippe, Massimiliano Fatica, Prabhat, Michael Houston
Evaluating and Accelerating High-Fidelity Error Injection for HPC
Chun-Kai Chang, Sangkug Lym, Nicholas Kelly,
Michael B. Sullivan
, Mattan Erez
Exploiting Idle Resources in a High-Radix Switch for Supplemental Storage
Matthias Blumrich
,
Ted Jiang
,
Larry Dennison
Fast, High Precision Ray/Fiber Intersection using Tight, Disjoint Bounding Volumes
Nikolaus Binder
,
Alex Keller
Massively Parallel Stackless Ray Tracing of Catmull-Clark Subdivision Surfaces
Nikolaus Binder
,
Alex Keller
Exascale Deep Learning for Climate Analytics
Thorsten Kurth, Sean Treichler, Joshua Romero, Mayur Mudigonda, Nathan Luehr, Everett Phillips, Ankur Mahesh, Michael Matheson, Jack Deslippe, Massimiliano Fatica, Prabhat, Michael Houston
CRUM: Checkpoint-Restart Support for CUDA's Unified Memory
Rohan Garg, Apoorve Mohan,
Michael B. Sullivan
, Gene Cooperman
Phantom Ray-Hair Intersector
Alexander Reshetov
,
David Luebke
Hamartia: A Fast and Accurate Error Injection Framework
Chun-Kai Chang, Sangkug Lym, Nicholas Kelly,
Michael B. Sullivan
, Mattan Erez
Isometry: A Path-Based Distributed Data Transfer System
Zhihao Jia, Sean Treichler, Galen Shipman, Patrick McCormick, Alex Aiken
Structurally Sparsified Backward Propagation for Faster Long Short-Term Memory Training
Maohua Zhu,
Jason Clemons
, Jeff Pool, Minsoo Rhu,
Steve Keckler
, Yuan Xie
Scalable Collectives for Distributed Asynchronous Many-Task Runtimes
Matthew Whitlock, Hemanth Kolla, Sean Treichler, Philippe Pebay, Janine C. Bennett
BabelFlow: An Embedded Domain Specific Language for Parallel Analysis and Visualization
Steve Petruzza, Sean Treichler, Valerio Pascucci, Peer-Timo Bremer