Awesome Multimodal
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Chenyang Lyu β most-cited papers & profile Β· Multimodal
β authors
Β·
overview
Chenyang Lyu
12
papers Β·
45
citations Β·
0
h-index
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
2023 Β· 27 citations
Retrieval-augmented Multi-modal Chain-of-thoughts Reasoning For Large Language Models
2023 Β· 14 citations
Marco-Voice Technical Report
2025 Β· 4 citations
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
2026
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
2026
Marco-Voice: A Unified Framework for Expressive Speech Synthesis with Voice Cloning
2026
Contextual and Seasonal LSTMs for Time Series Anomaly Detection
2026
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
2026
Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
2025
CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
2025
New Trends for Modern Machine Translation with Large Reasoning Models
2025
Is a Video worth $n\times n$ Images? A Highly Efficient Approach to Transformer-based Video Question Answering
2023
Topics
Model Architecture
Training Techniques
cs.SD
Vision-Language
In-Context Learning
RAG
Efficiency
cs.CL
eess.AS
Reinforcement Learning