Awesome AI for Code
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Xiawu Zheng — most-cited papers & profile · AI for Code
← authors
·
overview
Xiawu Zheng
20
papers ·
176
citations ·
18
h-index
Xiamen University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Motion-aware Latent Diffusion Models for Video Frame Interpolation
2024 · 8 citations
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
2025 · 3 citations
A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
2026
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
2026
Event-Anchored Frame Selection for Effective Long-Video Understanding
2026
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
2025
Training-Free Multimodal Large Language Model Orchestration
2025
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
2025
Mcp-zero: Active Tool Discovery For Autonomous LLM Agents
2025
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
2025
Motion-aware Latent Diffusion Models for Video Frame Interpolation
2024
A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
2023
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
2024
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
2024
Topics
Evaluation
Vision-Language Models
Benchmarks
Vision-Language
Video-Language
Model Architecture
Speech Translation
Text-to-Speech
Speech Recognition
Speech Enhancement