SplatsThe evolution of media, in brief
RSS

Give it scattered keypoints and it matches a full splat reconstruction

A fairground drop-tower ride rendered twice: on the left from a degraded reconstruction, where the tower and foliage dissolve into white streaks and smears, and on the right after refinement, sharp and photographic against a clear sky

Figure: Vuong et al., Carnegie Mellon University · Research

Wen Jiang

Wen Jiang

Aug 28, 2026, 10:40 PM ET-Research

Every 3D representation produces its own flavour of artefact when you render it from somewhere the cameras never went, and the field has answered with a specialist repair method for each — one for splats, another for radiance fields, custom architectures and retraining apiece. Khiem Vuong, Deva Ramanan and Srinivasa Narasimhan at Carnegie Mellon propose that none of that is necessary, because a pretrained video model already knows what a real scene looks like.

Why it matters: Their observation is that a badly rendered fly-through is still a coherent video. The camera motion survives, the coarse layout survives, and what is broken is texture and local structure — which makes cleanup a video-to-video translation problem rather than a 3D one. So they take a pretrained video generative model, finetune it lightly, and feed it the broken sequence.

One addition does most of the work: a binary mask marking which pixels are already clean. Without it the model rewrites frames it should have left alone, and quality drops by 1.3 dB. With it, the output stays anchored to the training views the trajectory passes through and only invents where invention is needed.

The result is a single model handling four representations that would each normally get a bespoke pipeline.

By the numbers:

  • On DL3DV at three input views, the model scores 15.74 dB fed a mesh rendering, 15.52 fed sparse SfM points, 15.18 fed a 3DGS rendering and 14.22 fed NeRF. The scattered points beat the full Gaussian reconstruction.
  • That is the finding worth sitting with. Sparse SfM points are keypoints and nothing else — the only dense imagery reaching the model comes from the handful of clean training views the camera path crosses. It is enough.
  • Against the specialist post-hoc methods on 3DGS renderings, it leads at every view count: 15.18 / 17.65 / 19.76 dB across three, six and nine views, against 14.62 / 17.35 / 19.19 for the strongest prior work.
  • Twenty paired training videos already produce effective cleanup at 16.70 dB, with diminishing returns past a hundred. The pretrained model brings the priors; the paired data mostly teaches it the conditioning.
  • Cutting denoising from 50 steps to 5 makes it ten times faster — a 61-frame clip at 480×832 in 31 seconds on one H100 — and the PSNR goes up, from 17.65 to 18.02.

The reward trick: The most transferable idea here is how they train for 3D consistency without a 3D loss. They run structure-from-motion on the model's own output and use the recovered camera pose accuracy as a reward signal for direct preference optimisation.

It works, and the way it works is instructive. After that finetuning, PSNR moves from 17.51 to 17.65 — barely — while pose accuracy jumps from 61.12 to 68.32. The image metrics can hardly see the improvement that the method exists to produce.

That is now the fourth study in a fortnight where photometric scores and the thing actually being optimised diverge. Here the authors simply went and measured the other thing.

Yes, but: The absolute numbers are low — 15 to 20 dB — because this is the hard sparse-view regime, and readers used to seeing 30 dB on dense captures should not compare across the two.

It is also a generative model, so it invents content, and the paper meets that squarely rather than hedging. Their position is that hallucination is the point when filling unobserved regions and a failure only when it contradicts what was observed. As a preliminary way to tell the difference they run inference five times with different seeds and use the per-pixel variance as an uncertainty map — sensible, and explicitly labelled preliminary.

The practical constraints are the video model's: it works on clips at 480×832, not arbitrary single views at arbitrary resolution, and the fast configuration still costs half a minute of H100 time per clip. Real-time cleanup is named as future work behind distillation, not something demonstrated.

The big picture: If the representation feeding the cleanup barely matters, the argument for expensive per-scene optimisation weakens at the sparse end. Scattered keypoints plus a few real photographs plus a good video prior is a cheaper pipeline than fitting a Gaussian field and then repairing it, and on this benchmark it scores the same.

The authors make the modularity explicit: because the video prior is used almost unmodified, a better video model can be dropped in without redesigning anything. That is a claim the next generation of video models will test on their behalf.

Go deeper:

  • FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors on arXiv
  • Project page
⟵ Back to the brief

More stories

Six holographic reconstructions of laboratory equipment photographed against black through a HoloLens: a Bunsen burner and a rack of capped test tubes above, and below them a shredded, torn reconstruction of a mortar and pestle beside two further mortar-and-pestle models whose pestles are visibly deformed

PSNR said Gaussian splatting won. Seventeen people said it didn't.

2 hours ago

Key art for the plugin: a Gaussian-splat capture of a derelict stone barn with a corrugated roof, sitting on a white tile and surrounded by scattered blue and purple splat points, with the wireframe box of its tile bounds drawn around it

To put splats on the globe, he replaced the renderer

6 hours ago

The base of a Gaussian-splat capture of a pasta box shown twice. Above, the shadow beside it breaks into a hard blocky wedge, circled in red by the developer. Below, after the fix, the same shadow falls away as a smooth gradient

Babylon.js gave splats a hard ceiling

9 hours ago

The same view of a white bicycle leaning against a black bench on grass, rendered twice side by side — once from the uncompressed reconstruction and once from the compressed one — with no visible difference between them

738 MB to 3.2 MB, without touching the training loop

9 hours ago

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy