STACK: Learning Composable Skills by Discovering Spatial and Temporal Structure with Foundation Models
Discovering reusable robot skills from a handful of demonstrations, then composing them to handle new scenes, constraints, and longer tasks.
I'm the co-founder and CEO of Verne, a robotics startup based in San Francisco.
We're building general-purpose robots for businesses, starting with manufacturers in the United States. We're backed by investors including Y Combinator, Standard Capital, and Crucible Capital.
My research spans robot learning, foundation models, and computer vision. I've worked with Jiajun Wu at Stanford and Shuran Song at Columbia, where I received my B.S. in computer science. At Apple, I worked on computer vision for health and fitness, gaze and gesture, and neural rendering for Apple Vision Pro.
Discovering reusable robot skills from a handful of demonstrations, then composing them to handle new scenes, constraints, and longer tasks.
Combining imitation learning and planning to turn language-annotated demonstrations into composable robot behaviors.
Learning how to interact with unfamiliar objects to discover their parts, reconstruct their geometry, and infer how they move.
Training convolutional and recurrent neural networks directly in memristor hardware for energy-efficient computing.

Aditya Sarathy, James P. Ochs, Yongyang Nie

Aditya Sarathy, Rene Aguirre Ramos, Umamahesh Srinivas, Yongyang Nie
Reviewer, Conference on Robot Learning (CoRL) · 2025, 2026