Awesome Computer Vision
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Reading Packs
News
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Reading Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Hanbo Zhang — most-cited papers & profile · Computer Vision
← authors
·
overview
Hanbo Zhang
17
papers ·
244
citations ·
19
h-index
Shanghai Medical College of Fudan University · National University of Singapore · Shanghai Municipal Center For Disease Control Prevention · Southwest Jiaotong University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Fully Convolutional Grasp Detection Network with Oriented Anchor Box
2018 · 227 citations
Hindsight Trust Region Policy Optimization
2019 · 8 citations
REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point Clouds
2020 · 4 citations
Density-based Curriculum for Multi-goal Reinforcement Learning with Sparse Rewards
2021 · 2 citations
A Real-time Robotic Grasp Approach with Oriented Anchor Box
2018 · 1 citations
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
2023 · 1 citations
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
2023 · 1 citations
MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
2025
Chain-of-action: Trajectory Autoregressive Modeling For Robotic Manipulation
2025
Robot Operation Of Home Appliances By Reading User Manuals
2025
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
2025
Robot Operation of Home Appliances by Reading User Manuals
2025
What Matters in Building Vision-Language-Action Models for Generalist Robots
2024
ROI-based Robotic Grasp Detection for Object Overlapping Scenes
2018
Towards Unified Interactive Visual Grounding in The Wild
2024
Topics
cs.RO
Vision-Language Models
Video-Language
Manipulation
Perception
cs.AI
cs.CV
Embodied & Agents
Benchmarks
Visual QA & Reasoning