Awesome Speech Audio
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Chaoyou Fu β most-cited papers & profile Β· Speech Audio
β authors
Β·
overview
Chaoyou Fu
11
papers Β·
0
citations Β·
15
h-index
Suzhou University of Science and Technology Β· Nanjing University
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
2026
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
2025
VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
2025
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
2025
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
2025
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
2025
Aligning Multimodal LLM with Human Preference: A Survey
2025
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
2025
BaseReward: A Strong Baseline for Multimodal Reward Model
2025
VITA-E: Natural Embodied Interaction With Concurrent Seeing, Hearing, Speaking, And Acting
2025
Topics
Evaluation
Vision-Language
Reinforcement Learning
Visual QA & Reasoning
Benchmarks
Manipulation
Perception
Control
Vision-Language Models
Efficiency