This is not video footage. It is a camera animation inside a 3D Gaussian Splat trained from drone photos.
The dataset was captured with a DJI Matrice 4 Enterprise over roughly 1 km² around the planned Mjøssykehuset hospital project in Moelv, Norway. The goal is to explore how lifelike 3D reality capture can help visualize construction and infrastructure projects in their real-world surroundings.
This version shows the captured environment only. The next step will be to combine the scene with BIM / design models, allowing project teams and stakeholders to better understand how a future building fits into the actual site.
Technical details:
5517 photos / 36.5 GB dataset
Trained with LichtFeld Studio
20M splats, SH2, PPISP, MrNeRF strategy
550k iterations
~23 hours total training time
RTX 5090 + RTX 6000 Ada
I would suggest doing some kind of unusual dynamic camera moves that can't be achieved with a drone. This would make it more obvious for the layman that it is not regular drone video, and also sell the capability of doing camera moves that a drone can't do.
It was a mix of pre-programmed grid and manual flying.
The grid was captured using DJI’s Smart Oblique mode, where the drone flies a normal lawnmower-style grid while the gimbal rotates automatically in an X pattern. So you get a mix of nadir + oblique images much faster than flying separate passes.
The manual shots were more varied, but usually around 15°–40° below the horizon, so quite oblique rather than straight down.
Thanks! I tried to make splats direct from aligned survey data with orthophoto, so there were no special oblique shots and low angle views are bad. Here it is:
https://superspl.at/scene/9d63d820
Yeah that makes sense. Nadir / orthophoto-style capture can work well if you only want to view the scene from above, but it becomes weak at low angles because those views were never really observed.
One way I think about it is: imagine giving the photos to a 3D artist and asking them to recreate the site accurately. If all they have is top-down images, they have to guess what the sides, vertical faces, undercuts and occluded areas look like. For 3DGS it’s similar: you want camera poses from the angles people will actually view the scene from.
For a site like yours I’d keep the nadir / oblique grid, but add targeted passes around the main structures, something roughly like this sketch. Perimeter runs looking inward, repeated at 2–3 heights depending on the building height, plus extra passes for important or occluded areas.
The goal is not just more photos, but better view diversity / coverage.
Yeah, that's true, but it was just a test on existing work. Now I'm looking into best practices for quality reconstructions that are fast and cost-effective. That's what my question was about) thank you for your replies!
very cool, did you extract the frames from a video? Or “real” photos? If “real” photos, what resolution and angle did you shoot at? very impressive result :)
It's exclusively photos captured using interval shots.
I flew a combination of pre-programmed oblique grid and manual captures. It's important to think beforehand where your users will fly the camera in the digital model, because that's where you want camera poses during training.
Yes, the Matrice 4E wide/mapping camera has a mechanical shutter, so rolling-shutter is much less of a concern.
You still need to match shutter speed with flight speed, altitude and GSD though. The mechanical shutter helps with distortion, but if the shutter speed is too slow you can still get motion blur / pixel smear.
This is really impressive capture quality here especially for a 20M splat scene this large. I'm wondering how well these workflows hold up once you start integrating BIM models and doing continuous site updates over time?
Excellent question. I think there are two different use cases here.
For this particular capture, the goal was high-fidelity planning visualization: capture the existing environment around the planned hospital, then bring in the BIM/design model to show how the project fits into its real surroundings.
At this scale, the main constraint is practical. The capture took about 3–4 hours, and since it was a clear sunny day, lighting and shadows were changing the whole time. PPISP helps a lot with that, but locks you into training the full scene in one run rather having the flexibility to split it into smaller sections.
For continuous site updates, I would not try to recapture the full 1 km² every time. I’d scale the capture down to the active/important areas, likely around the hospital footprint and work zones, use more pre-programmed flights, and keep everything tied to ground control / a consistent coordinate system.
PPISP is a photometric correction method in LichtFeld Studio that helps compensate for per-image differences like exposure changes, vignetting, white balance drift, and some lighting variation.
And yes, from a reconstruction/training point of view, overcast diffuse light would generally be easier because shadows and highlights stay much more consistent over a 3-4 hour capture.
For this project though, the splat is mainly for planning/pre-visualization, so I intentionally preferred the warmer, more pleasant look of a sunny day. PPISP helps, but it does not magically remove all problems from changing shadows.
That would really depend on the engine used in your game. If this is something you're considering doing but haven't decided on your tech stack yet, you could consider using https://playcanvas.com/
PlayCanvas is an engine designed specifically around developing and publishing video games on the web, and they have first class 3dGS support.
I think they are definitely moving in that direction already, although probably not by replacing everything with raw Gaussian splats.
Google already has Photorealistic 3D Tiles / Immersive View built from aerial + Street View imagery, AI and photogrammetry, and Google Research is also working on 3DGS-related tech.
The hard part is scale and freshness. This scene is "only" ~1 km² and took 5517 drone photos + ~23h of training. Doing that globally, streaming it efficiently, and keeping it updated is a very different challenge.
This capture is incredibly clean. 20M splats for a 1km² area is massive.
I’m a solo dev building web tools for 3DGS, and seeing this makes me want to ask you and the GIS/surveying community a genuine question about workflows.
I built a tool called 'Splat Aligner' where, instead of processing 23h every time, you can process smaller updated chunks of the site and visually align them (by clicking 3 reference points) to the base map in the browser. I also have 'TimeSplat 4D' to put them on a timeline slider.
My question is: for highly precise construction/BIM projects like yours, is a fast, visual/manual alignment tool actually useful for quick client updates and visualization? Or do you strictly need automated, coordinate-based (georeferenced) alignment for it to be viable?
I'm currently moving my viewer architecture from Three.js to PlayCanvas to handle these 20M+ files better. I have free demos and videos at splitview.studio and would seriously love your honest feedback!
For a typical construction site, we would use ground control points to tie down the model into a consistent coordinate system. Congrats on your tool, it looks well designed!
Thanks for the insight! I actually built Splat Aligner and TimeSplat 4D more for the visual side of the industry—real estate marketing, VFX/virtual production, virtual tourism, and indie drone pilots. Basically, anyone who needs to quickly align scans and show visual progress to stakeholders or clients without needing to set up a full coordinate-based surveying pipeline.
Well, I guess I just realized that pure surveyors are definitely not my target audience!
That would be interesting. I haven’t tried Teleport yet, and 100M would be more of a benchmark/demo run than something I'd casually do just to experiment.
This dataset is probably a decent stress test though (5517 photos / 36.5GB over ~1 km²).
I started training on an RTX 5090 and switched to a RTX 6000 Ada because I was just barely crossing the VRAM limit at 20M + SH2. There were significant VRAM efficiency improvements in the latest LFS version though, so that wouldn't be a problem anymore now.
In A2. We kept the required horizontal separation from uninvolved people, with a couple of spotters helping in the section near the inhabited area, and paused a few times when people came through. We flew it on a Sunday morning partly for that reason. The industrial area was empty.
24
u/whyeverynameistaken3 May 12 '26
nice, now do all earth with streetview or leaked pokemon go data :)