Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Kanishka Rao — most-cited papers & profile · Multimodal
← authors
·
overview
Kanishka Rao
24
papers ·
866
citations ·
30
h-index
Chaitanya Bharathi Institute of Technology · ICFAI Foundation for Higher Education · Vignana Bharathi Institute of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
2023 · 270 citations
Personalized Speech recognition on mobile devices
2016 · 164 citations
Federated Learning for Emoji Prediction in a Mobile Keyboard
2019 · 164 citations
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
2023 · 101 citations
Streaming End-to-end Speech Recognition For Mobile Devices
2018 · 23 citations
Multilingual Speech Recognition With A Single End-To-End Model
2017 · 17 citations
Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions
2023 · 16 citations
Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions
2023 · 16 citations
Off-Policy Evaluation via Off-Policy Classification
2019 · 15 citations
RL-CycleGAN: Reinforcement Learning Aware Simulation-To-Real
2020 · 14 citations
RL-CycleGAN: Reinforcement Learning Aware Simulation-To-Real
2020 · 14 citations
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
2024 · 14 citations
Multi-Dialect Speech Recognition With A Single Sequence-To-Sequence Model
2017 · 10 citations
Exploring Architectures, Data and Units For Streaming End-to-End Speech Recognition with RNN-Transducer
2018 · 10 citations
Open-vocabulary Queryable Scene Representations for Real World Planning
2022 · 8 citations
Topics
Speech Recognition
Manipulation
Speech Translation
Perception
Multi-Robot
Control
Value-Based
Human-Robot Interaction
Text-to-Speech
Offline RL