Richard Geldreich's Blog: Fitting Neural Textures and PBR Material Maps with ES (No Backprop)
Richard Geldreich's Blog
Co-owner of Binomial LLC, working on GPU texture interchange. Open source developer, Open Geospatial Consortium member, graphics programmer, former video game developer. Worked previously at SpaceX (Starlink), Valve, Ensemble Studios (Microsoft), DICE Canada.
Friday, September 4, 2026
Fitting Neural Textures and PBR Material Maps with ES (No Backprop)
The repo, released on September 4, 2026 is here:https://github.com/richgel999/neural_texture_es2ES=Evolution Strategies.Here is a mirror of the README.md file on GitHub, with the Prior Art disclosure.Key phrases: neural texture compression, neural material compression, PBR material compression, derivative-free neural texture optimization, latent texture optimization, block-compressed neural texture, quantization-aware neural texture training, compressed-domain texture optimization.toy neural texture compression trained with Evolution StrategiesA small, self-contained C++ experiment: an RGB image (or up to four same-size RGB textures of one material) is encoded as a shared low-resolution latent texture plus a tiny MLP decoder, and both are trained entirely with Evolution Strategies — no backprop, no autodiff, no training framework. (An optional late-training polish, --mlp-fd, switches the decoder to numerical finite differences; still no backprop.) Dependencies are stb_image, stb_image_write, and OpenMP.Write-up: Fitting a neural texture decoder with ESI(u,v) ≈ MLP( bilinear(Z, u, v), phi(u,v) ) Z is the latent texture, phi a small positional encoding. At decode time each pixel bilinearly samples Z at its UV, appends phi, and runs the MLP. An optional second, coarser latent level (--latent2) is sampled at the same UV and its channels are concatenated onto the first level's.Results512×512 crop of kodim23, 3000 iterations, latent quantized to 8 bits after training, MLP weights counted as fp16:LatentPSNRbpp (raw)bpp (entropy coded)64×64×426.9 dB0.560.4764×64×828.2 dB1.070.87128×128×430.3 dB2.061.65128×128×832.2 dB4.073.27These runs used the original positional encoding uv,fourier:1 (--nfreq 1), which is why the evaluation command below passes it. The current default is uv only, which scored 0.26–0.32 dB higher where both were run (see METHOD.md); the results quoted later in this file use that default unless stated otherwise.The 128×128×8 run uses a 14 → 24 → 24 → 3 MLP (1035 weights, leaky ReLU, sigmoid output) and trains in about 150 s on a 32-thread CPU. Quantizing the latent to 8 bits costs 0.04 dB.Output of that run (out_128c8/): target crop, reconstruction after 3000 iterations, and the eight latent channels side by side.TargetReconstruction (32.2 dB)out_128c8/model.bin is the trained model; evaluate it with ntc kodim23.png --load out_128c8/model.bin --latent 128 128 8 --nfreq 1 --iters 0.A 4-layer materialThe PavingStones070 material (normal, roughness, albedo, AO; see Test images) trained jointly from one shared 128×128×4 + 64×64×4 latent and one 10 → 36 → 36 → 12 MLP (2172 weights), 3000 iterations, learning rate annealed over the second half, per-weight finite differences for the decoder over the last quarter. 8-bit latent: 2.64 bpp total, 0.66 bpp per texture. Left is the target, right the reconstruction from the quantized latent (out_m1234/). The training command wasntc m1.png m2.png m3.png m4.png --latent 128 128 4 --latent2 64 64 4 --mlp 36,36 --mlp-pairs 64 --iters 3000 --lr-anneal 0.5 0.05 --mlp-fd 0.75 --out out_m1234 Normal map, 23.2 dBRoughness, 31.5 dBAlbedo, 23.2 dBAmbient occlusion, 29.6 dBout_m1234/model.bin is the trained material; evaluate it with ntc m1.png m2.png m3.png m4.png --load out_m1234/model.bin --latent 128 128 4 --latent2 64 64 4 --mlp 36,36 --iters 0.How the ES training worksMLP: antithetic ES, 32 perturbation pairs per step, each pair evaluated on the same random 4096-pixel minibatch. The estimated gradient is fed to Adam. Optional late phases: full-image minibatch (--mlp-full), per-weight finite differences (--mlp-fd), or frozen decoder (--mlp-freeze).Latent: all latent values are perturbed at once and the full image is decoded twice per pair, 4 pairs per step. Each pixel's loss change is credited only to the (up to) 4 texels its bilinear tap reads, on each latent level.Ordinary ES already updates every parameter from one antithetic pair, but its variance grows with parameter count: the single scalar loss difference is the sum of thousands of separate local effects, and each texel's share is buried under everyone else's. Here each texel's loss difference is measured only over the pixels in its own bilinear footprint, so noise from the other ~16k texels is discarded instead of averaged. That credit assignment, not the per-evaluation cost, is what makes ES practical on a latent with 131k values using only 4 pairs per step.The latent stays fp32 during training; quantization is applied afterwards and reported at a configurable bit depth.BuildingCMake generating a Visual Studio solution (MSVC), or any C++17 compiler with OpenMP on Linux/WSL. Tested with MSVC 2022 and MSVC 2026 on Windows, and gcc 13 under WSL2. Note that std::normal_distribution differs between standard libraries, so the same seed gives slightly different results on each platform.cmake -S . -B build -G "Visual Studio 17 2022" -A x64 cmake --build build --config Release For Visual Studio 2026 use -G "Visual Studio 18 2026", which needs CMake 4.2 or newer (the CMake bundled with VS 2026 works).Runningntc image.png --out out --latent 128 128 8 --iters 3000 ntc albedo.png normal.png rough.png --weights 1,1,0.5 --out out_mat --latent 128 128 8 --latent2 64 64 4 Several positional images form a material: they must have the same size after cropping, share the latent and the MLP (3 outputs per texture), and can be weighted in the loss with --weights (relative; a weight of 0 drops that texture from training). Per-texture PSNRs are printed alongside the overall one, output files gain a _tK suffix, and the bitrate is reported both per material pixel and per texture.Run with no arguments, ntc trains on the checked-in kodim23.png using the default 64×64×4 latent. Images are located relative to the executable, so this works from the build directory as well as the repo root. A named image that cannot be found is an error; the synthetic fallback only applies when no image is named.Progress is printed to stdout (MSE, PSNR, quantized PSNR and bitrate, latent stats, throughput). Reconstructions, a latent visualization per level, and model.bin are written to the output directory periodically; side_by_side.png (target | 8-bit-latent decode) is written at the end.Useful options (ntc --help lists them all):FlagMeaning--latent W H Clatent texture size (default 64 64 4)--latent2 W H Coptional second (typically coarser) latent level, e.g. 32 32 4 (off)--weights w0,w1,...per-texture loss weights for a material, one per positional image (relative; 0 drops a texture)--lat2-sigma FES sigma for the second level (default: same as --lat-sigma)--lr-anneal START FINALdecay both learning rates linearly from 1× at START·iters to FINAL× at the end--mlp W1,W2,...hidden layer widths (default 24,24)--act leaky|relu|tanh|sinehidden activation--pos SPECpositional features: uv, fourier:N, dct:N, local, lfourier:N, lquad, ldct:N, ldct2:N, ldct4:N, none--nfreq Nshorthand for --pos uv,fourier:N (the encoding the headline table used)--qbits Nlatent bit depth for the reported quantized PSNR / bitrate--load model.bin --iters 0evaluate a saved model--mlp-pairs, --mlp-batch, --mlp-sigma, --mlp-lr, --mlp-every, --lat-pairs, --lat-sigma, --lat-lrES hyperparameters--mlp-fd START, --mlp-fd-h Hfrom START·iters, train the MLP by central finite differences per weight--mlp-full START, --mlp-full-pairs Nfrom START·iters, evaluate the MLP ES step on the full image--mlp-freeze STARTfrom START·iters, stop updating the MLP (latent-only phase)--lat-alttwo latent levels: perturb one level per pair, rotating, to remove cross-level crosstalkImages larger than 512×512 are center-cropped by default (--crop); all textures of a material are cropped identically and must match afterwards.Test imageskodim23.png, kodim01.png, kodim02.png: the Kodak test set (768×512, center-cropped to 512×512 by default). kodim23 is the parrots.m1.png … m4.png: a 4-layer cobblestone material at 512×512, in order normal map, roughness, albedo, ambient occlusion. Derived from PavingStones070 on ambientCG, released under Creative Commons CC0 1.0. Train it as a material with ntc m1.png m2.png m3.png m4.png ....Prior art disclosureThe blog post above and the single-texture results (the original repository, richgel999/neural_texture_es) were published on September 3, 2026. This repository, including the two-level latent, materials, and the items marked "added September 4", was published on September 4, 2026. The following are disclosed here as public prior art.Neural texture representations using learned latent grids with small neural decoders, and Evolution Strategies / simultaneous-perturbation methods for derivative-free optimization, are established ideas. The technically distinctive part explored here is their combination with the decoder's known spatial dependency structure: all latent values are perturbed simultaneously, antithetic full-image evaluations produce per-pixel loss differences, and each pixel's loss difference is attributed only to the latent texels actually read by that pixel's filtering footprint. This yields simultaneous, support-restricted ES estimates for every latent texel while discarding loss variation from pixels a given texel cannot affect. Estimates for neighboring texels still share pixels and the same perturbation draw, so they are correlated rather than independent.For a given latent value, the omitted per-pixel loss terms do not depend on that value's perturbation, so their products with it have zero expectation in the ordinary Gaussian ES estimator. Footprint attribution therefore removes them without bias, as a variance-reduction mechanism that follows directly from the decoder's dependency graph.Implemented in this repository:A low-resolution latent texture plus a small MLP decoder, with both the latent and the decoder optimized entirely by antithetic Evolution Strategies, without backpropagation or analytic derivatives (an optional finite-difference polish for the decoder is numerical, not autodiff).Support-restricted footprint attribution for latent ES: all latent values are perturbed simultaneously and the full image is decoded for +ε and −ε. Each pixel's loss difference is attributed only to the latent texels in that pixel's bilinear sampling footprint, so one antithetic decode pair produces simultaneous local ES estimates across the entire latent while excluding loss terms that cannot depend on each texel.Separate ES schedules matched to parameter support: minibatched, many-pair ES for the globally acting decoder weights, and full-image, few-pair footprint-attributed ES for the spatially local latent, interleaved every iteration, with the estimates fed through Adam.Late-training decoder phases (added September 4): from a chosen fraction of the run, the decoder step can switch to (a) antithetic ES evaluated on the full image, (b) per-weight central finite differences on a shared minibatch, a numerical gradient with no autodiff, with the decoder's Adam state reset at the switch because the ES phase's second-moment estimate otherwise throttles it, or (c) no decoder updates at all (latent-only phase). Measured on mario with the 128×128×4 + 64×64×4 configuration, a 36,36 decoder, 64 ES pairs, and the learning rate annealed over the second half: finite differences over the last quarter gained 0.27 dB at 3000 iterations and 0.42 dB at 6000 (31.38 → 31.80 dB), matching a 12000 iteration run in half the iterations; the Adam reset more than doubled the effect; the full-image and frozen phases changed nothing measurable.Learning-rate annealing under ES (added September 4): decaying both learning rates linearly over the second half of a run. Because ES gradient noise is re-injected every step, a fixed-rate Adam run settles at a jitter floor; annealing removed most of a visible texel-aligned artifact and gained 0.85 dB at 6000 iterations on mario with a single 128×128×8 latent and the ldct:2 positional input (31.36 → 32.21 dB) and 0.56 dB on the two-level configuration above (30.82 → 31.38 dB).Alternating-level perturbation (added September 4, --lat-alt): with two latent levels, each antithetic pair perturbs only one level, rotating across steps, with each level's gradient scaled by its own pair count. This removes cross-level crosstalk in the footprint attribution exactly; measured with a 64×64 second level it lost 0.17 dB, because halving each level's pair count cost more than the small crosstalk it removed.Post-training scalar quantization of the latent with per-channel scale, and reported bitrate at arbitrary latent bit depth.Pluggable positional encodings for the decoder, including cell-periodic cosine features of the bilinear cell offset (ldct:N), found to improve quality at fine latent resolution.Configurable decoder depth, width, and activation; saved models record the size of every latent level, MLP layout, activation, positional spec, and texture count (not the output mapping or the loss weights; see METHOD.md §7).Two-level latent pyramid (added September 4, --latent2): a second latent texture sampled at the same UV and concatenated onto the first, trained with the same footprint attribution applied once per level. Measured at 3000 iterations, 8-bit latent: kodim23 64×64×4 + 16×16×4 gives 27.47 dB at 0.59 bpp versus 27.16 dB at 0.55 bpp for 64×64×4 alone; mario 128×128×4 + 32×32×4 gives 29.08 dB at 2.18 bpp versus 28.82 dB at 2.05 bpp.Materials trained by ES (added September 4): up to four same-size RGB textures trained jointly from one shared latent and one MLP with three outputs per texture, with per-texture loss weights and per-texture reporting. Compressing a material's textures jointly from a shared latent is established practice in neural texture compression; what is disclosed here is doing it entirely derivative-free: the per-pixel loss sums over every texture's channels before footprint attribution, so one antithetic decode pair of the whole material yields the support-restricted ES estimate for every latent texel with respect to all textures at once, and the decoder is trained by ES or per-weight finite differences, with no backpropagation through any texture. The weights enter the per-pixel loss before attribution, so a zero weight removes that texture's influence on the latent gradient exactly. Results at 3000 annealed iterations with per-weight finite differences over the last quarter, 8-bit latent: the 4-layer PavingStones070 material (normal, roughness, albedo, AO) sharing 128×128×4 + 64×64×4 and a 36,36 decoder reaches 23.2 / 31.5 / 23.2 / 29.6 dB at 0.66 bpp per texture (2.64 bpp total); the first two layers alone reach 25.0 / 32.5 dB at 1.31 bpp per texture. Two unrelated photographs (kodim23 + mario) sharing 128×128×8 + 64×64×4 (annealed, no finite-difference phase) land at 30.4 and 29.6 dB at 2.3 bpp per texture, about what each gets alone at a similar per-texture bitrate, as expected when there is nothing to share.Described, not yet implemented:Quantization-aware training under ES: quantize (or block-compress) the latent inside the decode used for every ES evaluation. Because ES only observes loss values, any non-differentiable quantizer or codec can sit in the loop with no straight-through estimator or differentiable surrogate.Latents stored in standard GPU texture formats, inside the training loop. The latent texture is ultimately a GPU texture, so it can be held in any format the hardware samples natively: uncompressed fixed-point formats (A8R8G8B8, R8, RG8, 4-bit and 5:6:5 packings, RGBA16), or block-compressed formats (BC1–BC7, BC6H for signed or HDR latents, ASTC LDR and HDR at any block size). Because ES only observes loss values, the format's encode–decode round trip can sit inside every ES evaluation with no straight-through estimator or differentiable surrogate: the trainer sees the latent exactly as the GPU will. For block formats, loss attribution is per block rather than per texel, since one endpoint change moves every texel in the block; for per-texel formats the attribution is unchanged.Search directly in the encoded representation. Rather than training a float latent and encoding it, make the encoded texture itself the parameter vector and perturb its stored fields directly: quantized texel values for fixed-point formats, or block endpoints, partition and mode selectors, and per-texel indices for block-compressed formats, using ES with discrete perturbations or stochastic coordinate descent with accept/reject, exactly as a conventional texture encoder searches. There is then no encoder inside the loop at all, only the format's decoder, the bitrate is fixed by the format by construction, and the trainer and the texture compressor are the same program with the MLP inside its distortion metric.Non-overlapping perturbation phases: perturb only texels or blocks on one phase of a 2×2 grid per evaluation so that, under bilinear sampling, no two perturbed footprints overlap and neighbor crosstalk vanishes; cycle the phase to cover all parameters. (Level alternation above is the cross-level analogue and is implemented; the within-level version is not.)Latent initialization from the image itself as a warm start for ES training: box-downsample the target (all textures of a material) to each latent level's resolution, then project each texel's stacked channel vector (3T values, or a small local patch of them) onto the top C principal components of those vectors (PCA), so the initial latent is the best C- channel linear summary of the local image content instead of noise. The decoder then starts by learning the inverse projection, which is close to linear. Also decoder initialization from a previously trained model.Materials with non-RGB channel counts (single-channel roughness or AO, two-channel normals) and normal-map-aware losses.Training a material through a BRDF (rendering loss): put the shading model inside the ES evaluation, latent → decoded normal, albedo, roughness and AO → BRDF under one or more lights and views → rendered image → loss against the same rendering of the original maps. Today each map is fitted with its own pixel MSE and hand-set weights; a rendering loss instead weights every map by how much it changes the shaded result, which is what a game actually sees, and it requires no derivative of the BRDF, the tone mapping, or anything else in the pipeline. With a per-pixel shading model (no shadows or screen-space effects) the loss stays a sum over pixels, so footprint attribution applies unchanged; effects that read neighboring pixels enlarge the footprint and are handled the same way the coarse latent level is. Several lights or views per evaluation just sum more per-pixel terms.SPSA (Simultaneous Perturbation Stochastic Approximation) and Rademacher ES in place of Gaussian ES: Spall's SPSA perturbs every parameter by a random ±1 (Rademacher) step of size c, evaluates the two sides L± = L(θ ± cΔ), and estimates each component as ĝ_j = (L₊ − L₋) / (2c·Δ_j). Because Δ_j = ±1, 1/Δ_j = Δ_j, so this is (L₊ − L₋)·Δ_j / (2c): exactly the antithetic ES estimator used here with a ±1 direction instead of a Gaussian one, classically with a single pair per update. Two evaluations estimate the whole gradient regardless of parameter count, and the perturbation packs as one bit per parameter. It plugs into the footprint attribution unchanged, since the per-texel loss differences do not depend on the perturbation distribution, and the local credit assignment should make it far less noisy than global SPSA. The planned experiment is antithetic Rademacher perturbations plus the footprint attribution, for the latent and for the decoder, benchmarked against the Gaussian ES used now.Learned interpolation kernels expressed as a few global parameters rather than as decoder inputs.StatusThis is a deliberately simple research toy for learning and experimentation, not a codec. Nothing is tuned. Obvious next steps: quantization-aware training, block-compressed (BC/ASTC) latents inside the training loop, non-RGB channel counts for materials (single-channel roughness/AO, two-channel normals) with normal-map-aware losses, and alternative losses.
Posted by
Rich Geldreich
at
2:50 PM
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.
Newer Post
Older Post
Home
Subscribe to: Post Comments (Atom)
Blog Archive
▼
2026
(47)
▼
September
(15)
Notes on NNTC vs. neural texturing NNTC repo now on GitHub Neural texture decoders can suppress DCT ringing a... New GitHub repo release: "neural_block_textures" Neural block textures vs. GPU textures (like BC7) ... Neural block textures stored inside KTX2: a near-p... Neural block GPU texture compression: Prior Art di... Neural block texturing example Neural block texture compression trained with Evol... Neural block texture compression with CUDA Neural GPU block compression Fitting Neural Textures and PBR Material Maps with... Mirror of neural texture compression trained with ... ETC1S Chroma Artifact Filtering, Real-Time Encoder... Fitting a neural texture decoder with ES (no backp...
►
August
(2)
►
July
(1)
►
June
(6)
►
May
(2)
►
April
(6)
►
March
(6)
►
February
(5)
►
January
(4)
►
2023
(7)
►
May
(1)
►
April
(5)
►
March
(1)
►
2022
(7)
►
January
(7)
►
2021
(25)
►
November
(1)
►
February
(23)
►
January
(1)
►
2020
(12)
►
September
(1)
►
August
(1)
►
April
(5)
►
February
(3)
►
January
(2)
►
2019
(12)
►
October
(2)
►
September
(1)
►
April
(9)
►
2018
(43)
►
October
(2)
►
July
(3)
►
June
(6)
►
May
(9)
►
April
(17)
►
March
(2)
►
February
(4)
►
2017
(24)
►
November
(5)
►
September
(1)
►
August
(1)
►
June
(2)
►
April
(2)
►
March
(4)
►
February
(2)
►
January
(7)
►
2016
(77)
►
December
(3)
►
November
(2)
►
October
(19)
►
September
(33)
►
August
(5)
►
July
(3)
►
June
(1)
►
May
(4)
►
April
(3)
►
March
(1)
►
January
(3)
►
2015
(62)
►
December
(12)
►
November
(9)
►
October
(6)
►
June
(2)
►
May
(6)
►
April
(1)
►
March
(1)
►
February
(5)
►
January
(20)
►
2014
(34)
►
November
(1)
►
October
(2)
►
June
(6)
►
May
(3)
►
April
(2)
►
March
(11)
►
February
(5)
►
January
(4)
►
2013
(6)
►
November
(1)
►
October
(4)
►
August
(1)
►
2012
(4)
►
November
(1)
►
August
(1)
►
July
(2)
About Me
Rich Geldreich
Seattle, WA, United States Back in the day I worked for several years at Digital Illusions on things like the first shipping deferred shaded game ("Shrek" - 2001), software renderers, and game AI. Then, after working for Microsoft at Ensemble Studios for 5 years as engine lead on Halo Wars, I took a year off to create "crunch", an advanced DXTc texture compression library. I then worked 5 years at Valve, where I contributed to Portal 2, Dota 2, CS:GO, and the Linux versions of Valve's Source1 games. I was one of the original developers on the Steam Linux team, where I worked with a (somewhat enigmatic) multi-billionaire on proving that OpenGL could still hold its own vs. Direct3D. I also started the vogl (Valve's OpenGL debugger) project from scratch, which I worked on for over a year. In my spare time I work on various open source lossless and texture compression projects: crunch, LZHAM, miniz, jpeg-compressor, and picojpeg.
View my complete profile
Simple theme. Powered by Blogger. |
Richard Geldreich's work explores the fitting of neural textures and Physically Based Rendering (PBR) material maps by leveraging Evolution Strategies (ES) to perform derivative-free optimization without relying on backpropagation. The foundational experiment involves encoding an RGB image as a shared, low-resolution latent texture coupled with a small Multi-Layer Perceptron (MLP) decoder, which are then trained entirely using ES. This approach focuses on neural texture compression and optimizing material representations within a compressed domain, employing techniques such as latent texture optimization and quantization-aware training.
The method centers on embedding the optimization process within the decoder's spatial dependency structure. The novelty lies in the support-restricted footprint attribution for latent optimization: all latent values are perturbed simultaneously, and the full image is decoded for both positive and negative perturbations. This results in per-pixel loss differences that are attributed only to the latent texels encompassed by that pixel's bilinear sampling footprint. This mechanism effectively discards loss variation from texels that cannot affect a given pixel, providing simultaneous, support-restricted estimates for every latent texel while managing correlations between neighboring parameters. This attribution is justified because omitted per-pixel loss terms do not depend on the latent perturbation, making them irrelevant to the gradient estimate, thereby acting as a variance-reduction mechanism.
The training is implemented using antithetic ES, where 32 perturbation pairs are evaluated at each step, and the estimated gradient is fed into an Adam optimizer. The process can be refined through optional late-training phases, such as using per-weight central finite differences for the decoder, or freezing the decoder entirely to focus optimization solely on the latent space. Learning rate annealing is also incorporated by linearly decaying both learning rates over the latter half of the training iterations, which mitigates the jitter associated with ES gradient noise and results in measurable quality improvements.
The research further investigates how to incorporate advanced features into this framework. This includes using positional encodings for the decoder, such as Fourier features, which enhance the quality of the reconstruction. The framework also supports a two-level latent pyramid, where a coarser latent level is sampled and concatenated with the finer level, and the footprint attribution is applied separately to each level, effectively removing cross-level crosstalk in the optimization. The study also demonstrates the joint training of multiple material textures from a shared latent, where per-texture loss weights are used to control the influence of individual textures during the optimization.
Further theoretical explorations suggest avenues for expanding this framework. One proposed direction is integrating quantization-aware training directly within the ES loop, allowing the latent space to be quantized or block-compressed inside the decoder evaluation. Since ES only observes loss values, any non-differentiable quantization or codec can be incorporated without requiring a straight-through estimator. This leads to the possibility of training the texture compressor itself within the optimization loop, where the format's encode-decode round trip is treated as part of the distortion metric.
Additional insights explore using alternative perturbation methods, such as Simultaneous Perturbation Stochastic Approximation (SPSA) or Rademacher ES, which achieve gradient estimates by evaluating only two endpoints, effectively packing parameter perturbations into a single bit per parameter update. The work also considers initializing the latent from the image itself by projecting the local feature vectors onto principal components, providing a warm start that is closer to the true low-dimensional representation. Finally, the framework can be extended to handle materials with non-RGB channel counts (like single-channel roughness or normal maps) by incorporating rendering losses based on a BRDF evaluation, which allows the optimization to directly incorporate the visual consequences of shading without needing derivatives of the complex rendering pipeline. |