Your phone records three viewpoints per shot. The pipeline throws two away.

Figure: Li et al., Cornell / Berkeley / Adobe / Georgia Tech (CC BY 4.0) · Research
The three rear cameras on an iPhone see the same scene from viewpoints about five degrees apart, simultaneously, every time the shutter fires. An Apple Vision Pro's stereo pair sits roughly fifteen degrees apart. A Lytro Illum records a thirteen-by-thirteen grid of them at once. Shamus Li and colleagues point out that essentially every reconstruction pipeline takes one of those streams and discards the rest, then spends considerable effort inventing the parallax it just threw out.
Why it matters: The received wisdom is that consumer camera baselines are too small to matter — five degrees of separation between phone lenses is nothing next to walking around an object. So the field standardised on a moving monocular camera, and when the camera cannot move enough, on learned priors that hallucinate the missing views.
That reasoning holds only when the camera can move. It fails completely for a single exposure, and it fails for anything in motion: a stationary monocular camera watching a moving hand has no angular diversity at all, and no amount of prior fixes a measurement that was never taken.
The paper separates two regimes that usually get conflated. Sensor-limited multi-view is one sensor trading spatial resolution for angular resolution — a light field camera, where more viewpoints means fewer pixels each. Exposure-limited multi-view is several sensors on one device capturing the same instant, where the extra views are free and the only question is whether the baseline is wide enough to be worth anything.
By the numbers:
- On real single-exposure static scenes with a Lytro Illum: monocular 3DGS scores 26.24 dB and monocular SparseGS 27.75 dB. Feeding the same reconstruction all 81 sub-views takes plain 3DGS to 30.50 dB — past the specialist sparse-view method by 2.75 dB.
- With an Apple Vision Pro's stereo pair, single exposure: 17.57 dB monocular against 20.89 dB using both cameras, and 21.57 dB with FSGS on top.
- Depth is where the gain is starkest. From one exposure, mean absolute relative depth error falls from 0.0939 monocular to 0.0526 with stereo, and RMSE from 0.4588 to 0.3038 — roughly half the error, from hardware that was already in the device.
- On synthetic scenes at a ten-degree baseline, one exposure: 16.25 dB monocular, 24.78 dB with a light field camera, 26.39 dB with the multiplexed prototype the authors built.
- The dataset is the contribution that will outlast the analysis: static and dynamic real captures of each scene through an iPhone 15 Pro, an Apple Vision Pro and a Lytro Illum, which is three quite different angular sampling patterns of the same content.
Yes, but: The result does not generalise to the case most people are actually in, and the authors say so rather than letting the reader find out. On casual handheld video — someone walking around a scene with a phone — the extra cameras buy almost nothing: 21.75 dB monocular against 22.12 dB using all three iPhone lenses, and 27.89 against 28.44 for stereo.
Worse for the thesis, adding a diffusion prior to the monocular stream beats the multi-view capture outright. Monocular with Difix3D+ scores 23.67 dB where the iPhone's three cameras with the same prior score 23.25. The paper's own figure caption explains it plainly: camera motion already supplies angular diversity, so either added views or the prior can repair the ghosting, and you do not need both.
So the honest summary is narrower than the title suggests. Multi-camera capture matters when monocular capture is most constrained — one shot, or a fixed camera on a moving subject — and stops mattering once you are walking around. The authors state exactly this in their closing paragraph, which is more than many papers manage about their own headline claim.
There is also a measurement caveat worth reading the tables carefully for: single-view and three-view methods are scored against different camera sets, so the columns within Table 1 are not all directly comparable to each other.
The big picture: The interesting comparison is not multi-view against monocular; it is multi-view against the enormous research effort spent compensating for monocular. Sparse-view methods, regularisers and diffusion priors all exist because angular coverage is scarce. Here is a paper observing that on a large fraction of consumer hardware it is not scarce, merely discarded at the file format.
For anyone capturing dynamic scenes the practical instruction is unambiguous, because it is the one regime where the alternative does not exist. A fixed monocular camera cannot recover motion and geometry simultaneously — the hand smears — and the second lens is not an improvement but a precondition.




