Machine and Human Intelligence Group · Ma Lab

One cause or two?

How does your brain know that a sight and a sound belong together? A short film about causal inference in perception, and about the shapes of its two ingredients: the blur of our senses, and our expectations.

Transcript

The transcript needs JavaScript. The video's captions carry the same text.

Based on the study by Shuze Liu, Trevor Holland (joint first authors), Wei Ji Ma and Luigi Acerbi (joint senior authors), PLOS Computational Biology (2026).

Try it yourself

Weigh the odds

A flash appears straight ahead, and you hear a beep. Did they come from the same thing? Drag the sound, and watch the balance weigh one cause against two.

A balance weighs "one cause" against "two causes". Below it, a blue hill shows where sight places the flash, a green hill where hearing places the beep, and an amber curve shows your brain's guess of where the beep came from.

 

The balance and the curves are computed live by the standard Bayesian causal-inference observer (Körding et al., 2007), the same one that drives the film. The amber curve is your brain's guess of where the sound came from: the merged guess and the sound-alone guess, mixed in proportion to the odds. Try making hearing blurrier: a wide gap becomes easier to explain with one cause.

What we found

Two ingredients, drawn from the data

Weighing one cause against two takes two ingredients: how blurry each sense is in every direction, and where the brain expects things to be, its prior. Models of causal inference often assume equal blur everywhere and a plain bell-curve prior, and so does the demo above. Instead, we ran an experiment and inferred both shapes from people's answers, using flexible models.

  • 15 paid volunteers
  • 44,600 answers
  • brief flashes and short beeps
  • Where was it? Same place, or not?

Sharpest straight ahead, then the blur levels off

Each hill shows how blurry sight is in one direction. Dashed: the usual assumption, the same everywhere. Solid: the shape we found. That blur grows away from the centre was known; the new part is its shape: a steep rise just off-centre, then a plateau, within the tested range (±20° for sight, ±15° for sound).

Expectations: a sharp peak on a broad hill

The prior: where people expect things to be. Dashed: a plain bell curve. Solid: what we found, a sharp central peak on broad tails ("probably straight ahead, but maybe anywhere"), described by a mixture of a Gaussian and a Laplace distribution.

The shapes are illustrative. The fitted ones are in the paper (Fig 4).

Distilled into simple formulas (σ(s) = σ0 + k1(1 − e−k2|s|) for the blur, and a mixture of a Gaussian and a Laplace distribution for the prior), these shapes explain the volunteers' answers better than the standard assumptions. We commit to the shapes, not to these exact formulas: the data pin down the shapes but not a unique formula. All of this rests on one experiment. We expect similar shapes to hold in other studies, but that is still to be tested.

Also in the study

  • Hedging wins. In the tasks with both a flash and a beep, people's answers are best described by strategies that give weight to both "one cause" and "two causes", either averaging them (model averaging) or choosing each as often as it is likely (probability matching), rather than always going with the more likely one (model selection).
  • Two senses at once are blurrier. With a flash and a beep together, sight was about 1.15 times and hearing about 1.6 times blurrier than alone, perhaps because attention is split between them.
  • Sounds get stretched. The flashes spanned 40° and the beeps 30°. People placed the beeps as if their range were stretched to match, by a factor close to 4/3. This may be specific to the experiment's design.

A long road to this study

18 years from inception to publication

For Weiji (and by extension, for all authors), this has been the project with the longest road from inception to publication (18 years). He started this project as one of his very first projects after becoming an Assistant Professor at Baylor College of Medicine in 2008. It was a continuation of his earlier work on causal inference in multisensory perception with Ladan Shams, Ulrik Beierholm and Konrad Körding. Trevor, at the time an undergraduate at Rice University, was Weiji's very first lab member. He spent countless hours building a wooden structure to mount the speakers on. He covered the structure with a black cloth, on which the visual stimuli were then projected. They had to cover the walls with foam to reduce echoes. The data set was very rich (richer than any behavioural data from a multisensory experiment that Weiji knew of) but also extremely challenging to model. Manolis Froudarakis, as a rotating PhD student, made a valiant attempt. Then, the project went into the fridge until Luigi joined the lab in 2014.

Luigi picked up the analysis and, in 2015, presented early results in a talk at the Cosyne conference (Acerbi, Holland & Ma), arguing that the standard Bayesian models of causal inference could not explain people's answers. That talk is also where the marmot made its first appearance. But the full analysis was too demanding for the tools of the time, and the project went back into the fridge. In 2022, Shuze joined the team and rescued it. Using methods we had built in the meantime, such as BADS, Shuze wrote the analysis code, fitted all the models and wrote the first draft of the paper. And we found an answer to the puzzling results from 2015: the standard assumptions about blur and expectations do not hold for these data. With the shapes inferred from the data, Bayesian models describe the answers remarkably well.

More of our work

Causal inference, from motion to lines

Which way am I heading?

When you move, your eyes and the balance organs in your inner ear both tell you which way you are heading. Do you check whether they agree before combining them? We built a framework to compare causal-inference strategies in a fully Bayesian way. Asked directly whether the two cues matched, every participant took the conflict between them into account, though not always in a Bayesian way. Heading judgements alone could not rule out that some people simply fused the cues; combining both tasks could.

Acerbi, Dokka, Angelaki & Ma (2018). Bayesian comparison of explicit and implicit causal inference strategies in multisensory heading perception. PLOS Computational Biology 14(7): e1006110.

One line, or two?

Causal inference is not only about combining senses. Are two line segments, seen on either side of an occluder, parts of the same line? We varied how uncertain people's vision was by showing the segments at different distances from where they were looking. People took their own uncertainty into account, through a heuristic rule that departs slightly but systematically from the Bayes-optimal one.

Zhou, Acerbi & Ma (2020). The role of sensory uncertainty in simple contour integration. PLOS Computational Biology 16(11): e1006308.

More from our labs: the Machine and Human Intelligence Group in Helsinki and the Ma Lab at NYU.

For modellers

Fit models like this one

This part is for researchers who build computational models of perception and behaviour. The observer in the demo has a closed form. The study's observer, with direction-dependent blur and a peaked prior, does not: its likelihood needs numerical integration and is slow to evaluate, as is typical of models in computational neuroscience. Our lab builds tools for fitting and comparing such models. Models like these are realistic and hard to fit, so we also use them as benchmarks in several of our methods papers.

Start here

The study in miniature

A notebook that simulates a small version of this experiment, fits a constant-blur and a direction-dependent-blur observer with PyBADS, and compares them with PyVBMC. It runs on Google Colab in several minutes.

Fit your model

PyBADS · BADS

Bayesian Adaptive Direct Search: fast, robust optimization for maximum-likelihood and MAP fits when the objective is rough, noisy or slow to evaluate, with up to about 20 parameters. In the study, we used BADS for every fit with up to 20 parameters, and CMA-ES for the 40-parameter semiparametric fits.

pip install pybads

Get the posterior and the evidence

PyVBMC · VBMC

Variational Bayesian Monte Carlo: an approximate posterior over your parameters and an estimate of the model evidence, for model comparison, from a small budget of likelihood evaluations, even noisy ones. Best with up to about 10 parameters.

pip install pyvbmc

When you can only simulate

PyIBS · IBS

Inverse binomial sampling: unbiased, efficient estimates of the log-likelihood of models you can simulate but not write down, for data with discrete responses. Pair it with BADS or VBMC.

pip install pyibs

All our tools

Which tool do I need?

Our model-fitting page matches your problem to a tool, with the docs and papers for each, and more to explore.

Questions about our tools? Ask in our Discussions forum, and follow along for new releases. The data and analysis code of the study are at LSZ2001/Audiovisual-causal-inference.

Cite and reuse

Cite the study

Liu S, Holland T, Ma WJ, Acerbi L (2026). Distilling noise characteristics and prior expectations in multisensory causal inference. PLOS Computational Biology 22(5): e1014251. doi:10.1371/journal.pcbi.1014251

@article{liu2026distilling,
  title   = {Distilling noise characteristics and prior expectations
             in multisensory causal inference},
  author  = {Liu, Shuze and Holland, Trevor and Ma, Wei Ji and Acerbi, Luigi},
  journal = {PLOS Computational Biology},
  year    = {2026},
  volume  = {22},
  number  = {5},
  pages   = {e1014251},
  doi     = {10.1371/journal.pcbi.1014251}
}

Reuse the film

The film is licensed CC BY 4.0: show it in class, put it in a talk, cut it up, with credit. Download it with burned-in captions, or get the version without captions, plus a caption file, from the release page.

Suggested credit: “One cause or two?”, Machine and Human Intelligence Group, University of Helsinki, CC BY 4.0.

The narration was generated with ElevenLabs and is also subject to the ElevenLabs terms.

How the film was made

The film was made with Claude Code, Anthropic's AI coding assistant, running Claude Opus 5.5. Luigi Acerbi directed it and checked every claim against the paper. Everything in it is generated in code: the pictures are SVG animated with Remotion, the narration is text-to-speech, and the sound effects and score are synthesized with NumPy. The causal-inference scene is computed by the same Bayesian observer you can play with above. How it was made, and the source under the MIT licence, are on GitHub.

Acknowledgements

The experiment was run at Baylor College of Medicine. Shuze Liu is at Harvard University (Kempner Institute), Wei Ji Ma at New York University and Luigi Acerbi at the University of Helsinki; Trevor Holland was at Baylor College of Medicine. For this study, Luigi Acerbi was partly supported by the Research Council of Finland (grants 356498 and 358980) and acknowledges the research environment provided by ELLIS Institute Finland. We thank Emmanouil Froudarakis for a precursor to the analysis, and the staff of NYU's High Performance Computing services.

Follow along

More in the pipeline

We have plenty of new methods and tools on the way. Follow Luigi to hear about new releases and papers from the lab.