Ramani Duraiswami

Ramani Duraiswami

Professor, Department of Computer Science and the Institute for Advanced Computer Studies (UMIACS)
University of Maryland, College Park
Director, Physical Intelligence and Reality Lab (PIRL)
Affiliate appointments in Electrical and Computer Engineering, Applied Mathematics and Scientific Computation, the Artificial Intelligence Interdisciplinary Institute, and the Center for Automation Research

I work on machines that hear, and on the numerical methods that make physical models fast and differentiable. Three threads run through it: audio language models, scientific computing, and spatial perception — separate literatures that keep turning out to need each other.

Research

01Large Audio Language Models

Language models that listen — speech, sound and music as first-class input rather than transcribed text. Recent work includes the Audio Flamingo series of fully open audio-language models, now extended to video in Audio-Visual Flamingo and to music in Music Flamingo; GAMA and CompA on audio understanding and compositional reasoning; TEMPO on temporal grounding; and the MMAU benchmarks, which exist because the field could not otherwise say what these models actually understand.

Audio-language models, audio reasoning and captioning, benchmarking and evaluation, speech enhancement, retrieval-augmented audio.

02Scientific Computing and Differentiable Modeling

Fast multipole (demo) and boundary element methods for the Laplace and Helmholtz equations, and the translation theory underneath them. Recent work makes these solvers differentiable: hand-derived tangents and adjoints, so a wave solver can sit inside an optimization loop or a learned model rather than beside one. Much of this is joint work over twenty-five years with Nail A. Gumerov, who died in 2022; the codes are being opened and documented.

The same instinct runs into learning: GAIA studies geometry-adaptive operator learning for forward and inverse problems, where a surrogate has to respect a geometry it was not trained on. Older questions have a way of returning — an exact Navier–Stokes solution Gumerov and I found in 1998 turns out to speak to a 2026 blow-up construction.

Six frames of a gas bubble collapsing in gravity, from a sphere at t = 0.01 to a small flattened cavity at t = 0.925, computed with the boundary element method.
A gas bubble collapsing under gravity, solved on the Laplace boundary element library: sphere to jet touchdown at t = 0.925, volume down to one per cent of its initial value, energy conserved to two per cent with no smoothing. The physics of 1992, rebuilt in a day on libraries that make the numerics somebody else’s problem.

Fast multipole methods, boundary integral equations, adjoints and shape derivatives, GPU and heterogeneous computing, operator learning.

03Spatial Perception, with a focus on Audio

How listeners localize sound, and how to reproduce or measure that computationally: head-related transfer functions and their personalization, spherical microphone arrays, room acoustics, and audio-visual scene analysis. The physics and the perception are the same problem seen from two ends, which is why the solver work and this work share a laboratory.

Increasingly this meets the first branch directly: SPUR gives audio language models spatial understanding, and a differentiable multi-sphere scattering model lets spatial cues be studied analytically and optimized rather than only measured.

Map of left-ear head-related transfer function magnitude over azimuth and elevation at 4.1 kHz, showing a broad gain on the left side and a deep shadow on the right.
What a head does to sound at 4.1 kHz: the left-ear transfer function over 1,730 directions, computed on a scanned head with our own boundary element solver. Nearly 10 dB of gain on the near side, close to 30 dB of shadow on the far side, and the ragged structure in between is what tells a listener where a sound came from.

HRTFs, binaural audio, microphone arrays, source localization, room impulse responses.

Selected work

All publications · Google Scholar · Semantic Scholar · ACL Anthology · ORCID

Talks

All talks

Software

Spatial audio work from this laboratory underlies technology commercialized through VisiSonics and shipped in millions of devices.

Group

Around twenty graduate students work in the lab, across all three areas above, and eighteen have completed doctorates with me since 2000. Members and alumni. Enquiries from prospective students are welcome by email.

Background

Contact

ramanid@umd.edu · CV · LinkedIn · X
(301) 405-6710
Room 4244, Brendan Iribe Center
University of Maryland, College Park, MD 20742