Learning to Track Instances without Video Annotations

Tracking segmentation masks of multiple instances has been intensively studied, but still faces two fundamental challenges: 1) the requirement of large-scale, frame-wise annotation, and 2) the complexity of two-stage approaches. To resolve these challenges, we introduce a novel semi-supervised framework by learning instance tracking networks with only a labeled image dataset and unlabeled video sequences. With an instance contrastive objective, we learn an embedding to discriminate each instance from the others.

Weakly-Supervised Physically Unconstrained Gaze Estimation

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly available and can be much more easily annotated with frame-level activity labels. In this work, we tackle the previously unexplored problem of weakly-supervised gaze estimation from videos of human interactions.

Contrastive Syn-to-Real Generalization

Training on synthetic data can be beneficial for label or data-scarce scenarios. However, synthetically trained models often suffer from poor generalization in real domains due to domain gaps. In this work, we make a key observation that the diversity of the learned feature embeddings plays an important role in the generalization performance.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective.

Sihao Liu

Sihao Liu is a Research Scientist in the Architecture Research Group at NVIDIA Research. His work spans computer architecture, programmable accelerators, and AI-assisted chip design. He studies how to make specialized hardware easier to build and program through compiler–architecture co-design, automated design-space exploration, and agentic tools for electronic design automation (EDA).

ShareMMU: Supporting Secure Address Translation Sharing among Untrusted Accelerators

The growing demand for accelerated computing is driving widespread deployment of multi-accelerator systems in the cloud and at the edge. These systems are temporally shared across processes that may not mutually trust one another and often include large pools of third-party, untrusted accelerators operating within a shared virtual address space. Due to the limited capability of the Input-Output Memory Management Unit (IOMMU) in supporting address translation, SoC designers are integrating large Shared Translation Lookaside Buffers (TLBs) that serve many accelerators and processes.

FastGen-PDD: Parallel Decoding Distillation for Image and Video Generation

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion.

Composing Detector Error Models

  • Fault-tolerant quantum programs thread together reusable logical operations, yet correcting errors requires comparing measurements across operation boundaries. Correlations tie each operation's error analysis to the computation around it. Adaptive computation makes this a runtime challenge: measurement results determine which operation comes next, so the error analysis must keep pace with execution.