Phase 08: Generative AI

3D Generation

3D is the modality where 2D-to-3D leverage is strongest. The 2023 breakthrough was 3D Gaussian Splatting. The 2024-2026 generative push layers multi-view diffusion + 3D reconstruction on top to produce objects and scenes from a single prompt or photo. 3D content is painful: Representation. Meshes, point clouds, voxel grids, signed distance fields (SDFs), neural radiance fields (NeRFs), 3D Gaussians. Each has trade-offs. Data scarcity. ImageNet has 14M images. The largest clean 3D dataset (Objaverse-XL, 2023) has 10M objects, most low quality. Memory. A 512³ voxel grid is 128M voxels; a useful scene NeRF needs 1M samples/ray. Generation is harder than reconstruction. Supervision. For a 2D image you have the pixels. For 3D you usually have a handful of 2D views and have to lift to 3D. The 2026 stack separates the two problems. First, generate 2D multi-view images with a diffusion model. Second, fit a 3D representation (usually Gaussian splatting) to those images. 3D generation: multi-view diffusion + 3D reconstruction Represent a scene as a cloud of 1M 3D Gaussians. Each has 59 parameters: position (3), covariance (6, or quaternion 4 + scale 3), opacity (1), spherical-harmonics color (48 at degree 3, 3 at degree 0). Rendering = projection + alpha-compositing. Fast (100 fps at 1080p on a 4090). Differentiable. Fit by gradient descent against ground-truth photos. A scene fits…

3D Generation: 3D is the modality where 2D-to-3D leverage is strongest. The 2023 breakthrough was 3D Gaussian Splatting. The 2024-2026 generative push layers…

This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.

Browse the complete course catalog or open this lesson on GitHub.