SplatsThe evolution of media, in brief
RSS
Daniel Habib

Daniel Habib

Aug 16, 2026, 7:20 AM ET-Research

Pull an object out of a splat you didn't capture

A query photo of a red handbag among clutter, beside a baseline extraction where the bag comes out surrounded by smeared background, beside Seed2GS's clean isolated bag on black

Figure: Ding et al., Chinese Academy of Sciences · Research

Zongjian Ding and colleagues at the Chinese Academy of Sciences, HKUST, Zhejiang University and the Beijing Institute of Technology report the highest published LERF-MASK accuracy for object extraction from a pre-built splat scene — with the scene frozen and no access to the cameras that built it.

Why it matters: Most 3D editing workflows receive a finished splat, not a capture session. The source images and reconstruction cameras are somebody else's, from months ago, and were never shipped with the asset.

Nearly every existing extraction method assumes otherwise — it wants the original cameras back, or it wants to train a per-scene representation for tens of minutes before you can ask it a question.

The insight: The authors separate two things that earlier methods tangle together: identity — which object do you mean — and coverage — where does it extend in 3D.

Identity is fixed once, from a single reliable reference mask chosen among open-vocabulary candidates. Coverage is then built up by lifting that seed and orbiting virtual viewpoints around it, with tracking carrying the seed forward instead of re-detecting the object in every view.

By the numbers:

  • 92.1% mean IoU on LERF-MASK, 3.7 points above the strongest scene-trained baseline and 7.6 above the closest camera-free one.
  • 9.3 seconds of measured compute-only latency, against tens of minutes for scene-trained methods.
  • 95.7% mIoU on 3D-OVS.
  • With one fixed reference view per scene the full pipeline still holds 91.1% mIoU.
  • Swapping the predicted seed for a ground-truth mask gains only 0.72 points — the seed selection is close to the ceiling.

The tell: In the comparison figure the baselines don't fail by missing the object — they smear. Pull out a red bag and you get the bag plus a haze of background gaussians dragged along with it.

That difference matters more than the mIoU gap for anyone actually editing: a clean extraction is an asset, a smeared one is a cleanup job.

What's next: The scene stays frozen throughout — masks supervise a single temporary foreground value per gaussian rather than modifying the representation, which is what makes the query cheap and repeatable.

Nine seconds is the number to watch. It is the difference between an offline batch process and something that can sit behind a click in an editor.

Go deeper:

  • Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes on arXiv
⟵ Back to the brief

More stories

Two smeared input photographs of a black car, beside the method's reconstructed view and the ground truth, with the grille crop enlarged underneath each

Two blurry photos are now enough for a 3D scene

2 hours ago

Two aerial views of the same reconstructed conifer forest: at ten simulated minutes a small orange and black burn scar, and at fifty minutes a black scar covering most of the frame

They set fire to a gaussian splat of a real forest

5 hours ago

Three columns: the raw compressive measurement as an unreadable scatter of speckle, the method's reconstruction of a plate of hot dogs and a vending machine, and the ground truth beside it

One exposure in, a whole 3D scene out

Yesterday

The same orchard branch twice: on the left the radiance field rendered photorealistically, on the right the semantic field, with apples picked out in red and foliage in green

Radiance fields stop being pictures and start being places

Yesterday

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy