Backprop Alternative: Augmented Lagrangian Predictive Coding
Recorded: Sept. 14, 2026, 10 p.m.
| Original | Summarized |
Augmented Lagrangian Predictive Coding: training 1000-layer networks without backpropagation Augmented Lagrangian Predictive CodingTraining 1000-layer networks without backpropagation We introduce PC-ALM, a local alternative to backpropagation. PC-ALM trains residual MLPs up to 1000 layers, nearly matching backprop's performance despite using only layer-local dynamics. PC-ALM equips each layer with a feedback control dynamical system that distributes and propagates supervision credit throughout a network. Resources Paper Code Authors Jeffrey Seely Julian Gould Published Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly[1, 2]. How the brain solves the multilayer credit assignment problem without explicit use of backprop remains one of the fundamental unsolved problems in neuroscience (though not without progress[3, 4, 5]). We show that PC-ALM can successfully propagate supervision credit in 1000-layer neural networks, overcoming standard PC's signal decay problem[11] while remaining layer-local. We focus on deep, small-width networks, a regime in which PC tends to perform poorly. Paper: arxiv.org/abs/2605.31022 Predictive coding: each layer as a dynamical system Predictive coding minimizeθ,h12‖y−WLhL−1‖2subject tohi=σ(Wihi−1),i=1,…,L−1. Predictive coding inference for t=1,…,T learning Wi←Wi−ηθ∇WiFPCfor i=1,…,L Per mini-batch, a forward pass initializes the activations, followed by T inference steps and a single weight update. We set T proportional to network depth; the 1000-layer experiments below use T=2L. Each hi-update reduces the prediction errors between layers adjacent to i. This is because ∇hiFPC depends only on hi−1, hi, and hi+1. Inference requires only nearest-neighbor communication ("message passing") between layers. Explicitly, writing ri=hi−σ(Wihi−1) for the prediction error between layers i−1 and i, the inference update reads3: hi←hi−ηh(ri↑error below−Wi+1⊤(σ′⊙ri+1↑error above))i=1,…,L−1 PC inference in a 32-layer residual MLP (width 16, ReLU) at weight initialization. Credit = per-layer norm of the prediction error ri. Dashed reference: norm of the backprop adjoint (the loss gradient with respect to that layer's activations). Since minimizing the free energy with respect to each hi does not enforce the layer-wise constraints to hold exactly, PC results in a different learning trajectory compared to standard backpropagation. Augmented Lagrangian Predictive Coding To use the augmented Lagrangian for training, we make a simple modification to PC: Augmented Lagrangian predictive coding inference for t=1,…,T learning Wi←Wi−ηθ∇WiLfor i=1,…,L Here α is the dual step size. PC-ALM is primal descent, dual ascent on the augmented Lagrangian, versus PC's descent on the energy. Mechanistic interpretation PC vs PC-ALM in a two-layer, linear, scalar network, y^=w2w1x, with x,y clamped to data values. Left: training trajectories of BP, PC, and PC-ALM in weight space. In this simple model, both PC and PC-ALM work and converge to the solution manifold. Performance differences between PC and PC-ALM become apparent in deep-narrow networks. Control theory and credit assignment Results PC-ALM successfully trains 1000-layer MLPs on MNIST. We use the residual MLP setup from Innocenti et al. (2026)[18]. We train for five epochs. Our architecture is a simple MLP with residual skip connections, using weight parameterizations that stabilize backprop training at this depth. MNIST test accuracy vs depth (width N=32, ReLU activation). PC-ALM stays within ~2 percentage points of backprop across the whole range, including at 1000 layers. Image classification benchmarks Across benchmarks. PC-ALM consistently narrows the gap between standard PC and global backprop. As depth increases, PC’s accuracy falls off much faster than PC-ALM’s. PC-ALM propagation dynamics Ballistic vs diffusive credit propagation. Credit magnitude across layers during inference. PC's credit decays with depth; PC-ALM's spreads evenly across the network. Stable oscillatory transient responses Oscillatory inference dynamics. We consider a deep linear network (σ=identity) with depth L=8 and width N=1. With weights fixed, the coupled inference dynamics across all layers form a linear system, characterized by the eigenvalues shown on the left (Equation (17) of the paper). Right: a neuron’s activity and dual variable during inference. The parameter α is the dual step size. Setting α=0 recovers PC exactly, with all eigenvalues real; increasing α introduces complex eigenvalues and oscillatory dynamics. Increasing α too far eventually destabilizes the network (Equation (14) of the paper). Discussion 1The Neuro-AI and distributed optimization communities share a concern with locality, but there has been relatively little cross-talk between them. Our work on PC-ALM ties these threads together. Citation BIBTEX Copy @misc{seely2026pc-alm, FootnotesIn Rao & Ballard, the prediction arrives from the layer above, since feedback carries predictions of lower-level activity and feedforward carries the residuals (Rao & Ballard 1999). Our equations follow the supervised convention of Whittington & Bogacz (2017), who fix the input at the highest level of the hierarchy, so the two conventions are the same mathematics with the hierarchy direction relabeled.Compared to the more standard unconstrained formulation of deep learning optimization, in which only the parameter vector θ is optimized.At the output, define rL:=y−WLhL−1. Since the readout is linear, take σ′=1 for this edge.We can see this by rewriting the augmented Lagrangian in a more interpretable form. After completing the square and discarding terms whose derivative is 0, the augmented Lagrangian reads L=12‖y−WLhL−1‖2 +12∑i‖ri+λi‖2. That is, each primal step is a standard PC step but on a prediction error r shifted by the dual λ.PC-ALM can be interpreted in two ways. From an optimization perspective, PC-ALM is primal descent on L (w.r.t. h) and dual ascent on L (w.r.t. λ). From a control perspective, r and λ are the P and I terms of a feedback controller on each layer.ReferencesBackpropagation and the Brain Lillicrap, T.P., Santoro, A., Marris, L., Akerman, C.J. and Hinton, G., 2020. Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. DOI: 10.1038/s41583-020-0277-3Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit Assignment Ororbia, A., 2023. arXiv preprint arXiv:2312.09257. Dendritic cortical microcircuits approximate the backpropagation algorithm [PDF]Sacramento, J., Costa, R.P., Bengio, Y. and Senn, W., 2018. Advances in Neural Information Processing Systems, Vol 31. ‘Backpropagation and the brain’ realized in cortical error neuron microcircuits [link]Max, K., Jaras, I., Granier, A., Wilmes, K.A. and Petrovici, M.A., 2026. PLOS Computational Biology, Vol 22(4), pp. e1014164. DOI: 10.1371/journal.pcbi.1014164Backpropagation through space, time and the brain [link]Ellenberger, B., Haider, P., Benitez, F., Jordan, J., Max, K., Jaras, I., Kriener, L. and Petrovici, M.A., 2026. Nature Communications, Vol 17, pp. 66. DOI: 10.1038/s41467-025-66666-zDecoupled Neural Interfaces using Synthetic Gradients Jaderberg, M., Czarnecki, W.M., Osindero, S., Vinyals, O., Graves, A., Silver, D. and Kavukcuoglu, K., 2017. International Conference on Machine Learning (ICML), Vol 70, pp. 1627—1635. An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity Whittington, J.C.R. and Bogacz, R., 2017. Neural Computation, Vol 29(5), pp. 1229—1262. DOI: 10.1162/neco_a_00949Predictive Coding: A Theoretical and Experimental Review Millidge, B., Seth, A. and Buckley, C.L., 2021. arXiv preprint arXiv:2107.12979. A survey on neuro-mimetic deep learning via predictive coding [link]Salvatori, T., Mali, A., Buckley, C.L., Lukasiewicz, T., Rao, R.P., Friston, K. and Ororbia, A., 2026. Neural Networks, Vol 195, pp. 108161. DOI: https://doi.org/10.1016/j.neunet.2025.108161Learning on Arbitrary Graph Topologies via Predictive Coding Salvatori, T., Pinchetti, L., Millidge, B., Song, Y., Bao, T., Bogacz, R. and Lukasiewicz, T., 2022. Advances in Neural Information Processing Systems. {ePC}: Fast and Deep Predictive Coding in Digital Simulation [link]Goemaere, C., Oliviers, G., Bogacz, R. and Demeester, T., 2026. Proceedings of the 43rd International Conference on Machine Learning. Advancing Neuromorphic Computing With Loihi: A Survey of Results and Outlook Davies, M., Wild, A., Orchard, G., Sandamirskaya, Y., Guerra, G.A.F., Joshi, P., Plank, P. and Risbud, S.R., 2021. Proceedings of the IEEE, Vol 109(5), pp. 911—934. Handbuch der physiologischen Optik von Helmholtz, H., 1867. Leopold Voss.Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects Rao, R.P. and Ballard, D.H., 1999. Nature Neuroscience, Vol 2, pp. 79—87. DOI: 10.1038/4580A Theoretical Framework for Inference and Learning in Predictive Coding Networks Millidge, B., Song, Y., Salvatori, T., Lukasiewicz, T. and Bogacz, R., 2022. arXiv preprint arXiv:2207.12316. {μ}{PC}: Scaling Predictive Coding to 100+ Layer Networks Innocenti, F., Achour, E.M. and Buckley, C.L., 2025. arXiv preprint arXiv:2505.13124. Benchmarking Predictive Coding Networks — Made Simple Pinchetti, L., Qi, C., Lokshyn, O., Olivers, G., Emde, C., Tang, M., M’Charrak, A., Frieder, S., Menzat, B., Bogacz, R., Lukasiewicz, T. and Salvatori, T., 2025. arXiv preprint arXiv:2407.01163. On the Infinite Width and Depth Limits of Predictive Coding Networks Innocenti, F., Achour, E.M. and Bogacz, R., 2026. arXiv preprint arXiv:2602.07697. Multiplier and Gradient Methods Hestenes, M.R., 1969. Journal of Optimization Theory and Applications, Vol 4, pp. 303—320. A Method for Nonlinear Constraints in Minimization Problems Powell, M.J.D., 1969. Optimization, pp. 283—298. Multiplier Methods: A Survey Bertsekas, D.P., 1976. Automatica, Vol 12(2), pp. 133—145. Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers Boyd, S., Parikh, N., Chu, E., Peleato, B. and Eckstein, J., 2011. Foundations and Trends in Machine Learning, Vol 3(1), pp. 1—122. DOI: 10.1561/2200000016Training Neural Networks Without Gradients: A Scalable {ADMM} Approach Taylor, G., Burmeister, R., Xu, Z., Singh, B., Patel, A. and Goldstein, T., 2016. Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. On {ADMM} in Deep Learning: Convergence and Saturation-Avoidance Zeng, J., Lin, S., Yao, Y. and Zhou, D., 2021. Journal of Machine Learning Research, Vol 22. A Theoretical Framework for Back-Propagation LeCun, Y., 1988. Distributed Optimization of Deeply Nested Systems Carreira-Perpinan, M.A. and Wang, W., 2014. Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, PMLR 33. {ADMM} for Efficient Deep Learning with Global Convergence [link]Wang, J., Yu, F., Chen, X. and Zhao, L., 2019. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery \& Data Mining, pp. 111–119. Association for Computing Machinery. DOI: 10.1145/3292500.3330936Decoupling Backpropagation using Constrained Optimization Methods [link]Gotmare, A., Thomas, V., Brea, J. and Jaggi, M., 2018. ICML 2018 Workshop on Credit Assignment in Deep Learning and Deep Reinforcement Learning. Proximal Backpropagation [link]Frerix, T., Mollenhoff, T., Moeller, M. and Cremers, D., 2018. International Conference on Learning Representations. Lifted Neural Networks Askari, A., Negiar, G., Sambharya, R. and El Ghaoui, L., 2018. arXiv preprint arXiv:1805.01532. Lifted Proximal Operator Machines Li, J., Fang, C. and Lin, Z., 2019. Proceedings of the AAAI Conference on Artificial Intelligence, Vol 33, pp. 4181—4188. Fenchel Lifted Networks: A {L}agrange Relaxation of Neural Network Training Gu, F., Askari, A. and El Ghaoui, L., 2020. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, PMLR 108. Contrastive Learning for Lifted Networks Zach, C. and Estellers, V., 2019. British Machine Vision Conference (BMVC). Lifted {B}regman Training of Neural Networks Wang, X. and Benning, M., 2023. Journal of Machine Learning Research, Vol 24. A Unified Framework for Lifted Training and Inversion Approaches Wang, X., Valavanis, A., Mahmood, A., Mang, A., Benning, M. and Repetti, A., 2025. arXiv preprint arXiv:2510.09796. Neural Network Training as an Optimal Control Problem : — An Augmented Lagrangian Approach — [link]Evens, B., Latafat, P., Themelis, A., Suykens, J. and Patrinos, P., 2021. 2021 60th IEEE Conference on Decision and Control (CDC), pp. 5136–5143. IEEE. DOI: 10.1109/cdc45484.2021.9682842An Augmented Lagrangian Method for Training Recurrent Neural Networks [link]Wang, Y., Zhang, C. and Chen, X., 2025. SIAM Journal on Scientific Computing, Vol 47(1), pp. C22-C51. DOI: 10.1137/23M1627614Inferring Neural Activity Before Plasticity as a Foundation for Learning Beyond Backpropagation Song, Y., Millidge, B., Salvatori, T., Lukasiewicz, T., Xu, Z. and Bogacz, R., 2024. Nature Neuroscience. DOI: 10.1038/s41593-023-01514-1Predictive Coding Networks for Temporal Prediction Millidge, B., Tang, M., Osanlouy, M., Harper, N.S. and Bogacz, R., 2024. PLOS Computational Biology, Vol 20(4), pp. e1011183. DOI: 10.1371/journal.pcbi.1011183Learning Complex Temporal Dependencies via Local Synaptic Plasticity Ng-Kee-Kwong, J., Tang, M., Akam, T. and Bogacz, R., 2026. bioRxiv preprint. DOI: 10.64898/2026.07.09.737423Blockwise Self-Supervised Learning at Scale Siddiqui, S.A., Krueger, D., LeCun, Y. and Deny, S., 2024. Transactions on Machine Learning Research. Backpropagation and the BrainT.P. Lillicrap, A. Santoro, L. Marris, C.J. Akerman, G. Hinton.Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. 2020. DOI: 10.1038/s41583-020-0277-3Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit AssignmentA. Ororbia.arXiv preprint arXiv:2312.09257. 2023. Dendritic cortical microcircuits approximate the backpropagation algorithm [PDF]J. Sacramento, R.P. Costa, Y. Bengio, W. Senn.Advances in Neural Information Processing Systems, Vol 31. 2018. ‘Backpropagation and the brain’ realized in cortical error neuron microcircuits [link]K. Max, I. Jaras, A. Granier, K.A. Wilmes, M.A. Petrovici.PLOS Computational Biology, Vol 22(4), pp. e1014164. 2026. DOI: 10.1371/journal.pcbi.1014164Backpropagation through space, time and the brain [link]B. Ellenberger, P. Haider, F. Benitez, J. Jordan, K. Max, I. Jaras, L. Kriener, M.A. Petrovici.Nature Communications, Vol 17, pp. 66. 2026. DOI: 10.1038/s41467-025-66666-zDecoupled Neural Interfaces using Synthetic GradientsM. Jaderberg, W.M. Czarnecki, S. Osindero, O. Vinyals, A. Graves, D. Silver, K. Kavukcuoglu.International Conference on Machine Learning (ICML), Vol 70, pp. 1627—1635. 2017. Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit AssignmentA. Ororbia.arXiv preprint arXiv:2312.09257. 2023. Backpropagation and the BrainT.P. Lillicrap, A. Santoro, L. Marris, C.J. Akerman, G. Hinton.Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. 2020. DOI: 10.1038/s41583-020-0277-3An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic PlasticityJ.C.R. Whittington, R. Bogacz.Neural Computation, Vol 29(5), pp. 1229—1262. 2017. DOI: 10.1162/neco_a_00949Predictive Coding: A Theoretical and Experimental ReviewB. Millidge, A. Seth, C.L. Buckley.arXiv preprint arXiv:2107.12979. 2021. A survey on neuro-mimetic deep learning via predictive coding [link]T. Salvatori, A. Mali, C.L. Buckley, T. Lukasiewicz, R.P. Rao, K. Friston, A. Ororbia.Neural Networks, Vol 195, pp. 108161. 2026. DOI: https://doi.org/10.1016/j.neunet.2025.108161Learning on Arbitrary Graph Topologies via Predictive CodingT. Salvatori, L. Pinchetti, B. Millidge, Y. Song, T. Bao, R. Bogacz, T. Lukasiewicz.Advances in Neural Information Processing Systems. 2022. {ePC}: Fast and Deep Predictive Coding in Digital Simulation [link]C. Goemaere, G. Oliviers, R. Bogacz, T. Demeester.Proceedings of the 43rd International Conference on Machine Learning. 2026. Advancing Neuromorphic Computing With Loihi: A Survey of Results and OutlookM. Davies, A. Wild, G. Orchard, Y. Sandamirskaya, G.A.F. Guerra, P. Joshi, P. Plank, S.R. Risbud.Proceedings of the IEEE, Vol 109(5), pp. 911—934. 2021. Handbuch der physiologischen OptikH. von Helmholtz.Leopold Voss. 1867. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effectsR.P. Rao, D.H. Ballard.Nature Neuroscience, Vol 2, pp. 79—87. 1999. DOI: 10.1038/4580A Theoretical Framework for Inference and Learning in Predictive Coding NetworksB. Millidge, Y. Song, T. Salvatori, T. Lukasiewicz, R. Bogacz.arXiv preprint arXiv:2207.12316. 2022. A survey on neuro-mimetic deep learning via predictive coding [link]T. Salvatori, A. Mali, C.L. Buckley, T. Lukasiewicz, R.P. Rao, K. Friston, A. Ororbia.Neural Networks, Vol 195, pp. 108161. 2026. DOI: https://doi.org/10.1016/j.neunet.2025.108161Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effectsR.P. Rao, D.H. Ballard.Nature Neuroscience, Vol 2, pp. 79—87. 1999. DOI: 10.1038/4580An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic PlasticityJ.C.R. Whittington, R. Bogacz.Neural Computation, Vol 29(5), pp. 1229—1262. 2017. DOI: 10.1162/neco_a_00949{μ}{PC}: Scaling Predictive Coding to 100+ Layer NetworksF. Innocenti, E.M. Achour, C.L. Buckley.arXiv preprint arXiv:2505.13124. 2025. Benchmarking Predictive Coding Networks — Made SimpleL. Pinchetti, C. Qi, O. Lokshyn, G. Olivers, C. Emde, M. Tang, A. M’Charrak, S. Frieder, B. Menzat, R. Bogacz, T. Lukasiewicz, T. Salvatori.arXiv preprint arXiv:2407.01163. 2025. On the Infinite Width and Depth Limits of Predictive Coding NetworksF. Innocenti, E.M. Achour, R. Bogacz.arXiv preprint arXiv:2602.07697. 2026. {ePC}: Fast and Deep Predictive Coding in Digital Simulation [link]C. Goemaere, G. Oliviers, R. Bogacz, T. Demeester.Proceedings of the 43rd International Conference on Machine Learning. 2026. Multiplier and Gradient MethodsM.R. Hestenes.Journal of Optimization Theory and Applications, Vol 4, pp. 303—320. 1969. A Method for Nonlinear Constraints in Minimization ProblemsM.J.D. Powell.Optimization, pp. 283—298. 1969. Multiplier Methods: A SurveyD.P. Bertsekas.Automatica, Vol 12(2), pp. 133—145. 1976. Distributed Optimization and Statistical Learning via the Alternating Direction Method of MultipliersS. Boyd, N. Parikh, E. Chu, B. Peleato, J. Eckstein.Foundations and Trends in Machine Learning, Vol 3(1), pp. 1—122. 2011. DOI: 10.1561/2200000016Training Neural Networks Without Gradients: A Scalable {ADMM} ApproachG. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, T. Goldstein.Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. 2016. On {ADMM} in Deep Learning: Convergence and Saturation-AvoidanceJ. Zeng, S. Lin, Y. Yao, D. Zhou.Journal of Machine Learning Research, Vol 22. 2021. A Theoretical Framework for Back-PropagationY. LeCun. 1988. On the Infinite Width and Depth Limits of Predictive Coding NetworksF. Innocenti, E.M. Achour, R. Bogacz.arXiv preprint arXiv:2602.07697. 2026. Distributed Optimization of Deeply Nested SystemsM.A. Carreira-Perpinan, W. Wang.Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, PMLR 33. 2014. Training Neural Networks Without Gradients: A Scalable {ADMM} ApproachG. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, T. Goldstein.Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. 2016. {ADMM} for Efficient Deep Learning with Global Convergence [link]J. Wang, F. Yu, X. Chen, L. Zhao.Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery \& Data Mining, pp. 111–119. Association for Computing Machinery. 2019. DOI: 10.1145/3292500.3330936On {ADMM} in Deep Learning: Convergence and Saturation-AvoidanceJ. Zeng, S. Lin, Y. Yao, D. Zhou.Journal of Machine Learning Research, Vol 22. 2021. Decoupling Backpropagation using Constrained Optimization Methods [link]A. Gotmare, V. Thomas, J. Brea, M. Jaggi.ICML 2018 Workshop on Credit Assignment in Deep Learning and Deep Reinforcement Learning. 2018. Proximal Backpropagation [link]T. Frerix, T. Mollenhoff, M. Moeller, D. Cremers.International Conference on Learning Representations. 2018. Lifted Neural NetworksA. Askari, G. Negiar, R. Sambharya, L. El Ghaoui.arXiv preprint arXiv:1805.01532. 2018. Lifted Proximal Operator MachinesJ. Li, C. Fang, Z. Lin.Proceedings of the AAAI Conference on Artificial Intelligence, Vol 33, pp. 4181—4188. 2019. Fenchel Lifted Networks: A {L}agrange Relaxation of Neural Network TrainingF. Gu, A. Askari, L. El Ghaoui.Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, PMLR 108. 2020. Contrastive Learning for Lifted NetworksC. Zach, V. Estellers.British Machine Vision Conference (BMVC). 2019. Lifted {B}regman Training of Neural NetworksX. Wang, M. Benning.Journal of Machine Learning Research, Vol 24. 2023. A Unified Framework for Lifted Training and Inversion ApproachesX. Wang, A. Valavanis, A. Mahmood, A. Mang, M. Benning, A. Repetti.arXiv preprint arXiv:2510.09796. 2025. Neural Network Training as an Optimal Control Problem : — An Augmented Lagrangian Approach — [link]B. Evens, P. Latafat, A. Themelis, J. Suykens, P. Patrinos.2021 60th IEEE Conference on Decision and Control (CDC), pp. 5136–5143. IEEE. 2021. DOI: 10.1109/cdc45484.2021.9682842An Augmented Lagrangian Method for Training Recurrent Neural Networks [link]Y. Wang, C. Zhang, X. Chen.SIAM Journal on Scientific Computing, Vol 47(1), pp. C22-C51. 2025. DOI: 10.1137/23M1627614Inferring Neural Activity Before Plasticity as a Foundation for Learning Beyond BackpropagationY. Song, B. Millidge, T. Salvatori, T. Lukasiewicz, Z. Xu, R. Bogacz.Nature Neuroscience. 2024. DOI: 10.1038/s41593-023-01514-1Predictive Coding Networks for Temporal PredictionB. Millidge, M. Tang, M. Osanlouy, N.S. Harper, R. Bogacz.PLOS Computational Biology, Vol 20(4), pp. e1011183. 2024. DOI: 10.1371/journal.pcbi.1011183Learning Complex Temporal Dependencies via Local Synaptic PlasticityJ. Ng-Kee-Kwong, M. Tang, T. Akam, R. Bogacz.bioRxiv preprint. 2026. DOI: 10.64898/2026.07.09.737423Blockwise Self-Supervised Learning at ScaleS.A. Siddiqui, D. Krueger, Y. LeCun, S. Deny.Transactions on Machine Learning Research. 2024. |
Predictive coding provides a framework for understanding unconscious processing rooted in Helmholtz's theories, and it can be viewed as a layer-local alternative to backpropagation for training deep neural networks. Standard predictive coding (PC) models the network dynamics through diffusive coupling between layers, where each layer attempts to minimize prediction errors by passing error signals upward. Mathematically, this involves minimizing a free energy function, FPC, which combines the supervised loss with a penalty for constraint violations between adjacent layer activations. Training involves alternating inference steps, where the network updates its states based on local prediction errors, followed by weight updates. While PC can successfully train networks on simple tasks, it suffers from a signal decay problem, particularly in deep, narrow architectures, as supervision must propagate through sequential local compromises, leading to weak credit signals reaching earlier layers. Augmented Lagrangian Predictive Coding (PC-ALM) is introduced as a method to enhance the propagation of these supervision credit signals while maintaining the layer-local dynamics characteristic of predictive coding. PC-ALM achieves this by replacing the standard PC energy with an augmented Lagrangian formulation, incorporating Lagrange multipliers, denoted as $\lambda_i$, at each layer. This modification draws upon the mathematical structure found in distributed optimization methods, suggesting that the multipliers can encode the necessary credit assignment information. The process involves running a primal descent step on the augmented Lagrangian with respect to the hidden states ($h$) and a dual ascent step with respect to the multipliers. Mechanistically, the augmented Lagrangian approach introduces a dual variable ($\lambda_i$) per layer, which, when combined with the prediction errors, functions analogously to the proportional and integral terms of a feedback controller at each layer. This framework suggests that the dual variables accumulate and recover the exact backpropagation credit signals, particularly in linear networks. The method is interpreted through both an optimization perspective, where PC-ALM performs primal descent on the augmented Lagrangian, and a control perspective, where the prediction errors ($r$) and the dual variables ($\lambda$) act as the primal and integral terms of a feedback controller. The overall goal of PC-ALM is to facilitate gradient computation and credit distribution across a network using only layer-local dynamics. This approach effectively distributes supervision credit across the network, leading to improved performance compared to standard PC. Experiments demonstrate that PC-ALM successfully trains 1000-layer multilayer perceptrons on tasks like MNIST, achieving performance nearly matching that of standard backpropagation using only layer-local dynamics. Furthermore, PC-ALM has been shown to propagate credit more effectively than PC, exhibiting a ballistic credit propagation wavefront across the network, which is faster than the diffusive propagation observed in standard PC. This dynamic advantage allows PC-ALM to narrow the gap between predictive coding and global backpropagation, showing superior performance on image classification benchmarks and demonstrating that credit signals spread more evenly across the network depths compared to PC, which suffers from signal decay. The approach positions credit assignment as a problem solvable through constrained optimization methods, relating PC-ALM to concepts like lifted networks and augmented Lagrangian methods used in distributed training. This work explores the trade-off between different optimization regimes, noting that while PC offers a prospective configuration that might improve sample efficiency, PC-ALM improves credit propagation. The research suggests a tension between preserving the prospective configuration and optimizing for efficient credit distribution, hinting at a fundamental trade-off that can be tuned by parameters such as the number of inference steps and the dual step size. The motivation stems from the observation that the multipliers in a standard Lagrangian system encode backpropagation credit signals, suggesting a biologically plausible mechanism for how local neural dynamics can compute global credit assignment. |