Kun Yuan
Orcid: 0000-0002-6030-8862Affiliations:
- Technical University of Munich, Center for Machine Learning, Munich, Germany
- University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France
According to our database1,
Kun Yuan authored at least 33 papers
between 2020 and 2026.
Collaborative distances:
Collaborative distances:
Timeline
Legend:
Book In proceedings Article PhD thesis Dataset OtherLinks
Online presence:
On csauthors.net:
Bibliography
2026
GaVA-CLIP:Refining Multimodal Representations With Clinical Knowledge and Numerical Parameters for Gait Video Analysis in Neurodegenerative Diseases.
IEEE J. Biomed. Health Informatics, May, 2026
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy.
CoRR, March, 2026
CoRR, January, 2026
npj Digit. Medicine, 2026
Medical Image Anal., 2026
Proceedings of the Fortieth AAAI Conference on Artificial Intelligence, 2026
2025
LapFM: A Laparoscopic Segmentation Foundation Model via Hierarchical Concept Evolving Pre-training.
CoRR, December, 2025
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature.
CoRR, December, 2025
CoRR, November, 2025
CoRR, October, 2025
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding.
CoRR, September, 2025
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model.
CoRR, June, 2025
Int. J. Comput. Assist. Radiol. Surg., June, 2025
EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy.
CoRR, May, 2025
CoRR, May, 2025
Medical Image Anal., 2025
Learning multi-modal representations by watching hundreds of surgical video lectures.
Medical Image Anal., 2025
Recognizing Surgical Phases Anywhere: Few-Shot Test-Time Adaptation and Task-Graph Guided Refinement.
Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2025, 2025
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery.
Proceedings of the AI for Clinical Applications - First International Workshops, 2025
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting.
Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2025, 2025
Multi-modal Representations for Fine-Grained Multi-Label Critical View of Safety Recognition.
Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2025, 2025
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining.
Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025
Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence, 2025
2024
Int. J. Comput. Assist. Radiol. Surg., July, 2024
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation.
Proceedings of the Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, 2024
HecVL: Hierarchical Video-Language Pretraining for Zero-Shot Surgical Phase Recognition.
Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2024, 2024
Enhancing Gait Video Analysis in Neurodegenerative Diseases by Knowledge Augmentation in Vision Language Model.
Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2024, 2024
2023
CholecTriplet2022: Show me a tool and tell me the triplet - An endoscopic vision challenge for surgical action triplet detection.
Medical Image Anal., October, 2023
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.
CoRR, 2023
2020
Proceedings of the 17th IEEE International Symposium on Biomedical Imaging, 2020
Towards Content-Independent Multi-Reference Super-Resolution: Adaptive Pattern Matching and Feature Aggregation.
Proceedings of the Computer Vision - ECCV 2020, 2020