- September 04, 2026
- By Annie Krakower
- Illustration courtesy of SonoWorld
A PHOTO OF A BUSTLING CITYSCAPE or a majestic waterfall can be striking enough on its own. But picture this: What if you could virtually jump into those settings and take in all their sights—and sounds.
University of Maryland researchers are working to make that possible through SonoWorld, a new step in virtual reality technology that generates a 3D scene, complete with spatial audio, from a single image.
“When we are exploring our world, I can hear you, I can hear the music, I can hear the birds chirping,” says computer science Assistant Professor Ruohan Gao, who is working on the project along with Distinguished University Professor Ming C. Lin and doctoral students Derong Jin and Xiyi Chen. “Without this very important aspect, the world that we generate would be incomplete.”
Using artificial intelligence vision and language models, SonoWorld creates a 360-degree panorama from an original image and predicts meaningful objects, like a microwave or sink in a kitchen, in a process called semantic grounding. AI also infers what sounds would make sense from different sources—beeping buttons, dripping water—allowing the researchers to encode spatial audio. Users can then realistically experience auditory change in each ear as they turn their head.
Its applications include more efficient content creation for entertainment like gaming or film. The next step for the team is to incorporate dynamic motion into the scenes.
“That is the kind of fully immersive experience that we hope to deliver,” Lin says. “We call it ‘being there.’”
Issue
Fall 2026Types
Explorations