Awesome Computer Vision
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Wenxuan Song β most-cited papers & profile Β· Computer Vision
β authors
Β·
overview
Wenxuan Song
22
papers Β·
14
citations
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
2024 Β· 10 citations
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
2025 Β· 2 citations
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
2025 Β· 1 citations
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
2023 Β· 1 citations
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
2026
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
2026
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
2026
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
2026
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
2026
FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment
2026
Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline
2026
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
2025
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
2025
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
2025
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
2025
Topics
Manipulation
Control
Perception
Multi-Robot
Human-Robot Interaction
Sim-to-Real
Locomotion
Vision-Language Models
Video-Language
Embodied & Agents