Awesome Generative Models
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Mike Zheng Shou — most-cited papers & profile · Generative Models
← authors
·
overview
Mike Zheng Shou
18
papers ·
36
citations ·
28
h-index
National University of Singapore
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
2025 · 17 citations
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
2025 · 4 citations
Factorized Learning For Temporally Grounded Video-language Models
2025 · 1 citations
Showui-\(π\): Flow-based Generative Models As GUI Dexterous Hands
2025
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
2026
World Action Models: The Next Frontier in Embodied AI
2026
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
2026
P-Flow: Prompting Visual Effects Generation
2026
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
2025
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
2025
MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping
2025
Ego-centric Predictive Model Conditioned on Hand Trajectories
2025
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
2025
WorldGUI: Dynamic Testing for Comprehensive Desktop GUI Automation
2025
Ego-centric Predictive Model Conditioned On Hand Trajectories
2025
Topics
Manipulation
Human-Robot Interaction
Video-Language
Vision-Language Models
Control
Benchmarks
Perception
Visual QA & Reasoning
Embodied & Agents
Planning