Awesome Computer Vision
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jiawei Zhou — most-cited papers & profile · Computer Vision
← authors
·
overview
Jiawei Zhou
19
papers ·
6
citations ·
12
h-index
Shanghai Jiao Tong University · China National Institute of Standardization
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Flow-SLM: Joint Learning of Linguistic and Acoustic Information for Spoken Language Modeling
2025 · 6 citations
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) i̊ghtarrow G$): Vision-Geometry Backbones over Language and Video Models
2026
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
2026
GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning
2026
PRISM: A Dual View of LLM Reasoning through Semantic Flow and Latent Computation
2026
Self-Improvement of Large Language Models: A Technical Overview and Future Outlook
2026
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
2026
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
2025
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
2025
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
2025
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
2025
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
2025
Safemvdrive: Multi-view Safety-critical Driving Video Synthesis In The Real World Domain
2025
Dynamic Token Reweighting for Robust Vision-Language Models
2025
Chunk-Distilled Language Modeling
2025
Topics
cs.CL
cs.AI
Vision-Language Models
Video-Language
Efficiency
Benchmarks
Evaluation
Vision-Language
Model Architecture
Audio Generation