Zihui Xue

Orcid: 0000-0001-7394-5169

According to our database1, Zihui Xue authored at least 38 papers between 2020 and 2026.

Collaborative distances:

Timeline

Legend:

Book  In proceedings  Article  PhD thesis  Dataset  Other 

Links

On csauthors.net:

Bibliography

2026
Personal Visual Context Learning in Large Multimodal Models.
CoRR, May, 2026

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models.
CoRR, April, 2026

SPOC: Spatially-Progressing Object State Change Segmentation in Video.
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2026

2025
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.
, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ,
Int. J. Comput. Vis., December, 2025

Seeing without Pixels: Perception from Camera Trajectories.
CoRR, November, 2025

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning.
CoRR, October, 2025

Seeing the Arrow of Time in Large Multimodal Models.
CoRR, June, 2025

Vid2Coach: Transforming How-To Videos into Task Assistants.
Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, 2025

Seeing the Arrow of Time in Large Multimodal Models.
Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, 2025

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning.
Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, 2025

REG: Rectified Gradient Guidance for Conditional Diffusion Models.
Proceedings of the Forty-second International Conference on Machine Learning, 2025

Progress-Aware Video Frame Captioning.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025

Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025

2024
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness.
CoRR, 2024

HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness.
Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, 2024

Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos.
Proceedings of the Computer Vision - ECCV 2024, 2024

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos.
Proceedings of the Computer Vision - ECCV 2024, 2024

Learning Object State Changes in Videos: An Open-World Perspective.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.
, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

Detours for Navigating Instructional Videos.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

2023
SUGAR: Efficient Subgraph-Level Training via Resource-Aware Graph Partitioning.
IEEE Trans. Computers, November, 2023

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.
CoRR, 2023

Egocentric Video Task Translation @ Ego4D Challenge 2022.
CoRR, 2023

Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment.
Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, 2023

The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation.
Proceedings of the Eleventh International Conference on Learning Representations, 2023

Egocentric Video Task Translation.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

Dynamic Multimodal Fusion.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

2022
The Modality Focusing Hypothesis: On the Blink of Multimodal Knowledge Distillation.
CoRR, 2022

Training-Free Robust Multimodal Learning via Sample-Wise Jacobian Regularization.
CoRR, 2022

SUGAR: Efficient Subgraph-level Training via Resource-aware Graph Partitioning.
CoRR, 2022

Co-advise: Cross Inductive Bias Distillation.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022

2021
Sampling Graphlets of Multiplex Networks: A Restricted Random Walk Approach.
ACM Trans. Web, 2021

What Makes Multimodal Learning Better than Single (Provably).
CoRR, 2021

What Makes Multi-Modal Learning Better than Single (Provably).
Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, 2021

Multimodal Knowledge Expansion.
Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, 2021

On Feature Decorrelation in Self-Supervised Learning.
Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, 2021

Anytime Depth Estimation with Limited Sensing and Computation Capabilities on Mobile Devices.
Proceedings of the Conference on Robot Learning, 8-11 November 2021, London, UK., 2021

2020
Sampling Graphlets of Multi-layer Networks: A Restricted Random Walk Approach.
CoRR, 2020


  Loading...