Awesome AI Agents
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yu-Chiang Frank Wang — most-cited papers & profile · AI Agents
← authors
·
overview
Yu-Chiang Frank Wang
11
papers ·
48
citations ·
44
h-index
National Taiwan University · Nvidia (United States)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
2025 · 46 citations
Receler: Reliable Concept Erasing of Text-to-Image Diffusion Models via Lightweight Erasers
2023 · 2 citations
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
2025
Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations
2025
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
2025
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
2025
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
2025
Zoom-zero: Reinforced Coarse-to-fine Video Understanding Via Temporal Zoom-in
2025
Wavelet Channel Attention Module with a Fusion Network for Single Image Deraining
2020
EoRA: Training-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation
2024
Top co-authors
Ankita Pasad
· 1
Arushi Goel
· 1
Boris Ginsburg
· 1
Chao-Han Huck Yang
· 1
Chenhui Chu
· 1
Hanrong Ye
· 1
Jinchuan Tian
· 1
Kunal Dhawan
· 1
Rafael Valle
· 1
Ryo Hachiuma
· 1
Shinji Watanabe
· 1
Shizhe Diao
· 1
Topics
Training Techniques
Video-Language
Vision-Language
Audio Understanding
Multimodal Audio
Speech Recognition
Manipulation
Human-Robot Interaction
Control
cs.IR