NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
Problem
Framing
Novel-view synthesis lacked a continuous scene representation that could recover complex geometry, non-Lambertian appearance, and high-resolution views from sparse posed images. NeRF closes this gap by fitting a 5D radiance field with differentiable volume rendering, positional encoding, and hierarchical ray sampling, then surpasses prior methods on PSNR, SSIM, and LPIPS.
Currently Used Methods
Direct antecedents
- Local Light Field Fusion — sampled local light fields for forward-facing view synthesis.
- Limitation in context: weak extrapolation to complex geometry and lower fidelity than NeRF.
- Scene Representation Networks — implicit scene functions for novel-view rendering.
- Limitation in context: lower image quality on both synthetic and real scenes.
- Neural Volumes — learned volumetric scene representation with rendering.
- Limitation in context: discretized or sampled representations lose fidelity and efficiency.
- DeepVoxels — persistent 3D voxel features for object-centric view synthesis.
- Limitation in context: voxel storage scales poorly and struggles with complex appearance.
- @mildenhallNeRF2020 — continuous 5D radiance field with differentiable rendering.
- Limitation in context: rendering remains slow, around 30 seconds per frame.
Proposed Method
Architecture
NeRF maps a 5D query to density and view-dependent color . The core model is an 8-layer ReLU MLP with 256 channels; density depends on position, while color also conditions on viewing direction. Two networks are used: a coarse model for proposal sampling and a fine model for final rendering.

Loss / Objective
Training minimizes coarse and fine RGB reconstruction over a batch of rays.
Sampling Rule / Algorithm
Rendered color is the volume-rendered integral along each camera ray, approximated with quadrature.
Hierarchical sampling draws fine samples from a piecewise-constant PDF induced by coarse weights.
Training Procedure
- Optimizer: Adam
- Learning rate:
- Learning-rate decay: exponential, down to over optimization
- Batch size: 4096 rays
- Coarse samples per ray:
- Fine samples per ray:
- Position encoding frequencies:
- Direction encoding frequencies:
- Real-scene density noise: Gaussian, unit variance before ReLU
Evaluation
Datasets
- Diffuse Synthetic 360: 4 simple Lambertian objects
- Realistic Synthetic 360: 8 path-traced objects with complex materials
- Real Forward-Facing: 8 handheld cellphone scenes
Metrics
- PSNR
- SSIM
- LPIPS
Headline results
- Diffuse Synthetic 360: PSNR 35.13, SSIM 0.983, LPIPS 0.027
- Realistic Synthetic 360: PSNR 31.01, SSIM 0.947, LPIPS 0.081
- Real Forward-Facing: PSNR 26.50, SSIM 0.811, LPIPS 0.250
- Drums scene PSNR: 25.01 vs 22.58 for Neural Volumes
- Ficus scene PSNR: 30.13 vs 24.79 for Neural Volumes
Table 1: Per-scene PSNR on the realistic synthetic dataset
| Method | Chair | Drums | Ficus | Hotdog | Lego | Materials | Mic | Ship |
|---|---|---|---|---|---|---|---|---|
| SRN [42] | 26.96 | 17.18 | 20.73 | 26.81 | 20.85 | 18.09 | 26.85 | 20.60 |
| NV [24] | 28.33 | 22.58 | 24.79 | 30.71 | 26.08 | 24.22 | 27.78 | 23.93 |
| LLFF [28] | 28.72 | 21.13 | 21.79 | 31.41 | 24.54 | 20.72 | 27.48 | 23.22 |
| Ours | 33.00 | 25.01 | 30.13 | 36.18 | 32.54 | 29.62 | 32.91 | 28.65 |
Ablations
- Removing positional encoding drops realistic-synthetic PSNR from 31.01 to 28.77.
- Removing view dependence hurts specular scenes and lowers PSNR to 27.66.
- Removing hierarchical sampling lowers PSNR to 30.06.
- Using only 25 input images reduces PSNR to 27.78.
Method Strengths and Weaknesses
Strengths
- Continuous scene function avoids voxel-memory growth.
- Strong gains on complex synthetic scenes: PSNR 31.01, LPIPS 0.081.
- View-dependent color captures specular effects missing in ablations.
- Hierarchical sampling improves quality without dense uniform ray queries.
Weaknesses
- Rendering is slow: about 30 seconds per frame on a V100.
- Requires accurate camera poses and scene bounds.
- Assumes static scenes; no dynamics or transient effects.
- Per-scene optimization prevents instant deployment to new scenes.
Suggestions from the authors
- Accelerate rendering to avoid tens-of-seconds per frame inference.
- Extend the representation beyond static scenes.
- Reduce dependence on known poses and calibrated bounds.
- Improve handling of unbounded forward-facing real scenes.
Links
Prior Papers
No prior vault papers identified yet.
Further Papers
- @kerblGaussianSplatting2023 — replaces slow NeRF volume rendering with real-time anisotropic Gaussian scene primitives.
- @GenerativeInverseDesignof2023 — builds on neural fields for continuous 3D scene or object optimization.
- @mildenhallNeRF2020 — foundational neural radiance field paper for continuous view synthesis.