Awesome Computer Vision
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Reading Packs
News
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Reading Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Song Li — most-cited papers & profile · Computer Vision
← authors
·
overview
Song Li
11
papers ·
0
citations ·
13
h-index
Beijing Jiaotong University · Zhejiang University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
2025
Large Language Models Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluation
2025
Large Language Models Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluation
2025
LongCat-Flash-Omni Technical Report
2025
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
2025
Spatial-aware Speaker Diarization for Multi-channel Multi-party Meeting
2022
M3FGM:a node masking and multi-granularity message passing-based federated graph model for spatial-temporal data prediction
2022
M3FGM:a node masking and multi-granularity message passing-based federated graph model for spatial-temporal data prediction
2022
Boosting long-term forecasting performance for continuous-time dynamic graph networks via data augmentation
2023
Enhancing Multilingual Speech Recognition through Language Prompt Tuning and Frame-Level Language Adapter
2023
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
2024
Topics
Speech Recognition
Audio Understanding
Text-to-Speech
cs.LG
cs.AI
Speech Translation
Audio Generation
cs.CR
cs.CL
cs.SE