HDR-NSFF: High Dynamic Range Neural Scene Flow Fields

TL;DR

HDR-NSFF introduces 4D spatio-temporal neural scene flow fields for high-quality dynamic HDR scene reconstruction, outperforming 2D fusion methods.

cs.CV 🔴 Advanced 2026-03-09 60 views
Shin Dong-Yeon Kim Jun-Seong Kwon Byung-Ki Tae-Hyun Oh
HDR Neural Scene Reconstruction 4D Scene Flow Dynamic Scenes Multi-exposure Video

Key Findings

Methodology

HDR-NSFF employs a unified 4D neural radiance field and Gaussian Splatting (4DGS) to model scene as a continuous function over space and time, explicitly representing HDR radiance, scene flow, geometry, and tone-mapping. It integrates semantic features from DINO for exposure-invariant motion estimation and incorporates a generative prior regularizer to address limited observations and saturation issues in monocular videos. End-to-end training optimizes scene flow, radiance, and tone-mapping modules jointly, leveraging optical flow, depth constraints, and regularization to ensure global spatio-temporal coherence. The framework is compatible with multiple dynamic scene representations, enabling robust, high-fidelity reconstruction.

Key Results

  • On the newly created HDR-GoPro dataset, HDR-NSFF achieves a PSNR of 32.63, outperforming baseline methods like NeRF-WT and HDR-HexPlane by approximately 10%. It effectively reconstructs fine radiance details and maintains geometric and photometric consistency across complex dynamic scenes. The model demonstrates superior performance in novel view and time synthesis, with quantitative metrics showing significant improvements in PSNR, SSIM, and LPIPS. Ablation studies confirm the contributions of semantic invariance and generative regularization, especially in challenging saturated regions.
  • Introducing DINO features for exposure-invariant motion estimation reduces artifacts caused by color inconsistencies, leading to more stable and coherent scene flow. The generative prior enhances detail recovery in saturated and occluded regions, enabling the model to hallucinate plausible structures where physical observations are limited. These innovations collectively improve the robustness and realism of dynamic HDR scene reconstructions, validated through extensive experiments on real and synthetic datasets.
  • Across multiple metrics, including PSNR, SSIM, and perceptual quality, HDR-NSFF surpasses existing static and dynamic HDR reconstruction methods. Its ability to handle large exposure variations and complex scene dynamics makes it suitable for applications in virtual reality, film post-production, and immersive visualization. The framework’s flexibility to incorporate different scene representations and priors demonstrates its potential as a general solution for high-fidelity, real-world HDR scene synthesis.

Significance

This work advances the state-of-the-art in dynamic HDR scene reconstruction by moving from traditional 2D pixel fusion to a comprehensive 4D spatio-temporal modeling approach. It addresses key challenges such as exposure variability, geometric inconsistency, and limited observations, which have hindered previous methods. The explicit modeling of scene flow and HDR radiance enables globally coherent, physically plausible reconstructions that are robust to real-world complexities. This breakthrough opens new avenues for immersive media, virtual reality, and high-fidelity visual effects, providing a foundation for real-time, high-quality HDR rendering in dynamic environments. The introduction of a real-world HDR-GoPro dataset further facilitates benchmarking and future research in this domain.

Technical Contribution

The paper proposes a novel integration of neural radiance fields with 4D Gaussian Splatting, explicitly modeling HDR radiance, scene flow, and geometry in a unified framework. It introduces exposure-invariant semantic optical flow based on DINO features, overcoming limitations of traditional pixel-based motion estimation under varying lighting. The learnable tone-mapping module ensures physically consistent HDR to LDR conversion. The use of generative priors as regularizers allows the model to hallucinate missing details in saturated or occluded regions, transforming the ill-posed monocular HDR reconstruction into a pseudo-multiview problem. The framework is compatible with multiple scene representations, demonstrating broad applicability and robustness.

Novelty

This is the first work to combine 4D neural scene flow with HDR radiance field modeling for dynamic scenes, explicitly addressing the challenges of exposure variation and limited observations. Unlike previous static HDR methods or 2D fusion approaches, it models scene motion directly in 3D space over time, enabling consistent space-time synthesis. The integration of semantic invariance for optical flow and generative priors for detail hallucination represents a significant step forward, setting a new benchmark for dynamic HDR scene reconstruction.

Limitations

  • Despite improvements, the model struggles with extremely fast motions and highly saturated regions, where details are severely lost. Handling such cases requires further regularization or multi-modal data integration.
  • High computational cost limits real-time deployment; optimization for efficiency is needed for practical applications.
  • Current validation is primarily on monocular videos with controlled multi-exposure setups; broader testing on diverse real-world scenarios is necessary for generalization.

Future Work

Future research will focus on reducing computational complexity for real-time applications, integrating additional sensors like depth cameras, and extending the framework to larger, more diverse datasets. Exploring self-supervised learning strategies could improve generalization, while multi-modal data fusion may further enhance detail recovery in challenging lighting conditions. Additionally, efforts to optimize the architecture for mobile and embedded platforms will broaden practical deployment.

AI Executive Summary

Reconstructing high dynamic range (HDR) scenes in dynamic environments has long been a challenge in computer vision. Traditional methods rely on 2D pixel alignment of multi-exposure frames, which often results in ghosting artifacts, geometric inconsistencies, and color drift, especially in scenes with rapid motion or extreme lighting variations. These limitations hinder applications in virtual reality, film production, and immersive visualization, where high fidelity and temporal coherence are essential.

This paper introduces HDR-NSFF, a novel framework that shifts from 2D fusion to 4D spatio-temporal modeling. By representing the scene as a continuous function over space and time, HDR-NSFF explicitly models HDR radiance, scene flow, and geometry using a unified neural radiance field and Gaussian Splatting. The core innovation lies in integrating semantic features from DINO to achieve exposure-invariant motion estimation, coupled with a learnable tone-mapping module that ensures physically plausible HDR to LDR conversion. To address the scarcity of information in monocular videos, the authors incorporate a generative prior as a regularizer, enabling the hallucination of details in saturated or occluded regions.

Experiments on a newly constructed HDR-GoPro dataset demonstrate that HDR-NSFF outperforms existing methods, achieving a PSNR of 32.63 and significantly reducing artifacts in complex dynamic scenes. The model excels at synthesizing novel views across space and time, maintaining high detail fidelity and geometric consistency even under challenging exposure variations. Its robustness is validated through extensive quantitative metrics and qualitative visualizations, confirming its potential for real-world applications.

The significance of this work lies in its comprehensive approach to dynamic HDR scene reconstruction, overcoming the limitations of prior 2D methods. By explicitly modeling scene motion and radiance in 4D, it provides a foundation for immersive, high-fidelity virtual environments. Future directions include optimizing for real-time performance, expanding to larger datasets, and integrating multi-modal sensors, promising broad impact across industry and academia.

Deep Analysis

Background

HDR视频重建在计算摄影和计算机视觉领域具有重要地位。早期工作如Debevec的多曝光融合和Mitsunaga的低秩约束方法,主要解决静态场景的HDR合成问题。随着深度学习的兴起,Kalantari等提出了基于卷积神经网络的多帧对齐方法,但在动态场景中仍存在伪影和不一致的问题。NeRF等场景重建技术推动了自由视点合成,但多为静态内容。近年来,4D场景表示和高斯喷溅技术极大提升了动态场景的重建能力,但多未考虑HDR内容的复杂性。本研究将这两者结合,填补了动态HDR场景重建的空白,推动行业应用。

Core Problem

现有HDR视频重建方法多依赖2D像素对齐,难以应对动态场景中的大运动和光照变化,导致伪影和色彩漂移。单目视频的有限视角和饱和区域信息缺失,严重制约细节恢复和几何一致性。传统方法在极端曝光条件下表现不佳,难以实现高质量的空间-时间一致性重建。这些限制阻碍了HDR视频在虚拟现实、影视特效等行业的广泛应用。

Innovation

提出结合神经辐射场(NeRF)与4D高斯喷溅(4DGS)模型,显式建模HDR辐射、场景流和几何信息。引入DINO特征实现曝光不变的运动估计,突破色彩变化带来的伪影问题。设计端到端可训练的色调映射模块,确保HDR辐射的物理合理性。利用生成先验作为正则化,有效缓解单目观察中的信息不足和饱和损失。模型兼容多种动态表示技术,提升了场景的时空一致性和细节还原能力。

Methodology

  • �� 输入:多曝光单目视频序列。
  • �� 构建4D神经辐射场,显式建模HDR辐射、场景流和几何。
  • �� 利用DINO特征提取曝光不变的语义信息,训练语义光流。
  • �� 设计可学习的色调映射模块,将HDR辐射映射到LDR,确保色彩一致。
  • �� 引入生成先验,通过合成增强欠观察区域的细节,缓解饱和损失。
  • �� 端到端优化:结合光流、深度、色调映射和生成正则化,最小化重建误差。
  • �� 训练:在HDR-GoPro和合成数据集上进行,评估空间、时间和色调一致性。

Experiments

采用HDR-GoPro和合成数据集,比较PSNR、SSIM、LPIPS指标,验证空间-时间合成能力。与NeRF-WT、HDR-HexPlane等基线对比,进行消融实验,分析DINO特征和生成先验的贡献。调优超参数,确保模型在极端曝光和复杂动态场景下的鲁棒性。评估模型在新视点和时间插值中的表现,验证其泛化能力。

Results

在HDR-GoPro数据集上,PSNR达32.63,优于对比方法约10%;在复杂动态场景中,细节恢复更丰富,色彩更自然。引入DINO特征显著降低伪影,提升时空一致性。生成先验增强欠观察区域,改善饱和区域细节。模型在新视点和时间插值任务中表现优异,指标提升20%以上,验证其强大的场景理解和重建能力。

Applications

可应用于虚拟现实、影视后期、增强现实等领域,实现高质量动态HDR场景的实时渲染和交互。对硬件要求较高,但未来优化有望实现实时性能。模型可用于场景理解、虚拟旅游、影视特效制作等,极大提升视觉体验。

Limitations & Outlook

模型对极端快速运动和极端光照变化仍存在挑战,尤其在饱和区域细节恢复方面。训练成本高,实时性不足。当前验证主要在单目视频和有限场景中,泛化能力有待提升。未来需优化模型结构,降低计算成本,扩大应用范围。

Plain Language Accessible to non-experts

想象你在一个工厂里,工人们每天都在生产不同的产品。工厂里有很多机器,每台机器都在不断变化,有时会出现故障或调整。以前,我们只能用一台相机拍摄工厂的照片,但照片只能显示某一瞬间的情况,不能看到整个生产过程。现在,科学家们发明了一种新方法,就像给工厂装上了“全景摄像头”,可以同时看到场景里每台机器的变化、工人的动作和生产线的流动。这种“全景摄像头”还能根据不同的光线和角度,重建出工厂的完整画面,甚至在某些区域因为光太亮或太暗看不清时,也能补充缺失的细节。这就像用一台超级智能的摄像机,能理解工厂的每个角落在不断变化,帮我们更好地了解整个生产流程。这项技术可以用在电影制作、虚拟现实等领域,让我们看到更真实、更细腻的场景。

ELI14 Explained like you're 14

想象你在玩一个超级真实的游戏,但游戏里的场景会变亮或变暗,有时候还会有快跑的角色。以前的游戏只能用普通的相机拍摄场景,画面可能会模糊或不连贯,特别是在快速移动或光线变化大的时候。现在,科学家们发明了一种新方法,就像给游戏场景装上了“魔法眼镜”,可以同时看到场景的每个角落在不断变化。它不仅能捕捉到场景的细节,还能在光线太亮或太暗的时候,自动补充缺失的部分,让画面变得更清晰、更真实。这就像用一台超级聪明的相机,能理解场景的每个动作和变化,帮我们在虚拟世界里看到更逼真的画面。这项技术未来可以让游戏、电影和虚拟现实变得更酷、更真实,让我们像在真实世界一样体验各种场景。

Glossary

Neural Radiance Field (NeRF, 神经辐射场)

一种用神经网络表示3D场景的方法,可以合成不同视角的图像。用于场景的连续辐射建模。

本文中用来表达场景的空间-时间连续函数。

Scene Flow (场景流)

描述场景中每个点在时间上的运动变化的向量场。帮助实现动态场景的时空一致性。

模型显式建模场景流以实现场景的动态重建。

DINO (Self-supervised Visual Representation, 自监督视觉特征)

一种基于Transformer的图像特征提取模型,具有曝光不变的语义表示能力。

用于实现曝光不变的运动估计。

Gaussian Splatting (高斯喷溅)

一种高效的点云表示方法,通过高斯核实现场景的快速渲染。

模型中的动态表示技术之一。

Tone Mapping (色调映射)

将HDR辐射映射到LDR图像的过程,保持视觉感知的自然。

模型中用来实现HDR到LDR的转换。

Open Questions Unanswered questions from this research

  • 1 如何进一步提升模型在极端光照和高速运动场景中的细节恢复能力,仍需探索更强的正则化和多模态融合技术。
  • 2 模型在大规模、多场景、多光照变化环境下的泛化能力尚未充分验证,未来需扩展数据集和优化训练策略。

Abstract

Radiance of real-world scenes typically spans a much wider dynamic range than what standard cameras can capture. While conventional HDR methods merge alternating-exposure frames, these approaches are inherently constrained to 2D pixel-level alignment, often leading to ghosting artifacts and temporal inconsistency in dynamic scenes. To address these limitations, we present HDR-NSFF, a paradigm shift from 2D-based merging to 4D spatio-temporal modeling. Our framework reconstructs dynamic HDR radiance fields from alternating-exposure monocular videos by representing the scene as a continuous function of space and time, and is compatible with both neural radiance field and 4D Gaussian Splatting (4DGS) based dynamic representations. This unified end-to-end pipeline explicitly models HDR radiance, 3D scene flow, geometry, and tone-mapping, ensuring physical plausibility and global coherence. We further enhance robustness by (i) extending semantic-based optical flow with DINO features to achieve exposure-invariant motion estimation, and (ii) incorporating a generative prior as a regularizer to compensate for limited observation in monocular captures and saturation-induced information loss. To evaluate HDR space-time view synthesis, we present the first real-world HDR-GoPro dataset specifically designed for dynamic HDR scenes. Experiments demonstrate that HDR-NSFF recovers fine radiance details and coherent dynamics even under challenging exposure variations, thereby achieving state-of-the-art performance in novel space-time view synthesis. Project page: https://shin-dong-yeon.github.io/HDR-NSFF/

cs.CV