Qiaomu Miao
Qiaomu Miao

About Me

Hi! I am Qiaomu Miao. I received my Ph.D. in Computer Science from Stony Brook University, advised by Prof. Dimitris Samaras and Prof. Minh Hoai Nguyen. My research expertise spans computer vision, vision-language models, multi-view geometry, generative AI, and human behavior analysis. Before that, I obtained my Bachelor’s and Master’s degrees in Computer Science at Tianjin University. During my Master’s study, I investigated the cognitive neural mechanisms of human vision.

Contact: qiamiao AT cs.stonybrook.edu

Publications
(2026). OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following. Conference on Neural Information Processing Systems (NeurIPS), 2026 (Spotlight).
(2026). Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer. European Conference on Computer Vision (ECCV), 2026.
(2026). Behavior-Based Skill Assessment for Open Surgery from Multi-View and Egocentric Videos. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026.
(2025). Multi-view Gaze Target Estimation. International Conference on Computer Vision (ICCV), 2025.
(2024). Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following. European Conference on Computer Vision (ECCV), 2024.
(2023). Patch-level gaze distribution prediction for gaze following. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023.
(2022). Study of detecting behavioral signatures within DeepFake videos. Arxiv Preprint, 2022.
Projects
Experience

Research Scientist Jul 2026 - Present

Meta

  • Working on content representation learning with multimodal input for large-scale autoregressive recommendation systems on Facebook Feed/Reels.


Student Researcher Sep 2025 - Apr 2026

Google

  • Developed eye segmentation and eye-tracking pipelines with large foundation models.


Technology Investigation Intern May 2022 - Aug 2022

Apple

  • Designed and trained activity recognition models on video datasets for augmented reality (AR) applications
  • Improved activity recognition performance by incorporating scene understanding concepts


Research Intern Jun 2021 - Aug 2021

Bytedance

  • Generated motion-transferred DeepFake videos with state-of-the-art lipsyncing and face reenactment models to investigate the behavioral signatures (speaking style and utterance, etc.) in person identification
  • Implemented a self-supervised DeepFake Detection model using behavioral and appearance-related features