首页 > AI前沿 > Backprop Alternative: Augmented Lagrangian Predictive Coding

Backprop Alternative: Augmented Lagrangian Predictive Coding

Hacker News 2026-09-15 02:03 2 阅读 查看原文
Citation Footnotes In Rao & Ballard, the prediction arrives from the layer above, since feedback carries predictions of lower-level activity and feedforward carries the residuals (Rao & Ballard 1999). Our equations follow the supervised convention of Whittington & Bogacz (2017), who fix the input at the highest level of the hierarchy, so the two conventions are the same mathematics with the hierarchy direction relabeled. Compared to the more standard unconstrained formulation of deep learning optimization, in which only the parameter vector θ is optimized. At the output, define rL:=y−WLhL−1. Since the readout is linear, take σ′=1 for this edge. We can see this by rewriting the augmented Lagrangian in a more interpretable form. After completing the square and discarding terms whose derivative is 0, the augmented Lagrangian reads L=12‖y−WLhL−1‖2 +12∑i‖ri+λi‖2. That is, each primal step is a standard PC step but on a prediction error r shifted by the dual λ. PC-ALM can be interpreted in two ways. From an optimization perspective, PC-ALM is primal descent on L (w.r.t. h) and dual ascent on L (w.r.t. λ). From a control perspective, r and λ are the P and I terms of a feedback controller on each layer. References Backpropagation and the Brain Lillicrap, T.P., Santoro, A., Marris, L., Akerman, C.J. and Hinton, G., 2020. Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. DOI: 10.1038/s41583-020-0277-3 Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit Assignment Ororbia, A., 2023. arXiv preprint arXiv:2312.09257. Dendritic cortical microcircuits approximate the backpropagation algorithm  [PDF] Sacramento, J., Costa, R.P., Bengio, Y. and Senn, W., 2018. Advances in Neural Information Processing Systems, Vol 31. ‘Backpropagation and the brain’ realized in cortical error neuron microcircuits  [link] Max, K., Jaras, I., Granier, A., Wilmes, K.A. and Petrovici, M.A., 2026. PLOS Computational Biology, Vol 22(4), pp. e1014164. DOI: 10.1371/journal.pcbi.1014164 Backpropagation through space, time and the brain  [link] Ellenberger, B., Haider, P., Benitez, F., Jordan, J., Max, K., Jaras, I., Kriener, L. and Petrovici, M.A., 2026. Nature Communications, Vol 17, pp. 66. DOI: 10.1038/s41467-025-66666-z Decoupled Neural Interfaces using Synthetic Gradients Jaderberg, M., Czarnecki, W.M., Osindero, S., Vinyals, O., Graves, A., Silver, D. and Kavukcuoglu, K., 2017. International Conference on Machine Learning (ICML), Vol 70, pp. 1627—1635. An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity Whittington, J.C.R. and Bogacz, R., 2017. Neural Computation, Vol 29(5), pp. 1229—1262. DOI: 10.1162/neco_a_00949 Predictive Coding: A Theoretical and Experimental Review Millidge, B., Seth, A. and Buckley, C.L., 2021. arXiv preprint arXiv:2107.12979. A survey on neuro-mimetic deep learning via predictive coding  [link] Salvatori, T., Mali, A., Buckley, C.L., Lukasiewicz, T., Rao, R.P., Friston, K. and Ororbia, A., 2026. Neural Networks, Vol 195, pp. 108161. DOI: https://doi.org/10.1016/j.neunet.2025.108161 Learning on Arbitrary Graph Topologies via Predictive Coding Salvatori, T., Pinchetti, L., Millidge, B., Song, Y., Bao, T., Bogacz, R. and Lukasiewicz, T., 2022. Advances in Neural Information Processing Systems. {ePC}: Fast and Deep Predictive Coding in Digital Simulation  [link] Goemaere, C., Oliviers, G., Bogacz, R. and Demeester, T., 2026. Proceedings of the 43rd International Conference on Machine Learning. Advancing Neuromorphic Computing With Loihi: A Survey of Results and Outlook Davies, M., Wild, A., Orchard, G., Sandamirskaya, Y., Guerra, G.A.F., Joshi, P., Plank, P. and Risbud, S.R., 2021. Proceedings of the IEEE, Vol 109(5), pp. 911—934. Handbuch der physiologischen Optik von Helmholtz, H., 1867. Leopold Voss. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects Rao, R.P. and Ballard, D.H., 1999. Nature Neuroscience, Vol 2, pp. 79—87. DOI: 10.1038/4580 A Theoretical Framework for Inference and Learning in Predictive Coding Networks Millidge, B., Song, Y., Salvatori, T., Lukasiewicz, T. and Bogacz, R., 2022. arXiv preprint arXiv:2207.12316. { μ}{PC}: Scaling Predictive Coding to 100+ Layer Networks Innocenti, F., Achour, E.M. and Buckley, C.L., 2025. arXiv preprint arXiv:2505.13124. Benchmarking Predictive Coding Networks — Made Simple Pinchetti, L., Qi, C., Lokshyn, O., Olivers, G., Emde, C., Tang, M., M’Charrak, A., Frieder, S., Menzat, B., Bogacz, R., Lukasiewicz, T. and Salvatori, T., 2025. arXiv preprint arXiv:2407.01163. On the Infinite Width and Depth Limits of Predictive Coding Networks Innocenti, F., Achour, E.M. and Bogacz, R., 2026. arXiv preprint arXiv:2602.07697. Multiplier and Gradient Methods Hestenes, M.R., 1969. Journal of Optimization Theory and Applications, Vol 4, pp. 303—320. A Method for Nonlinear Constraints in Minimization Problems Powell, M.J.D., 1969. Optimization, pp. 283—298. Multiplier Methods: A Survey Bertsekas, D.P., 1976. Automatica, Vol 12(2), pp. 133—145. Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers Boyd, S., Parikh, N., Chu, E., Peleato, B. and Eckstein, J., 2011. Foundations and Trends in Machine Learning, Vol 3(1), pp. 1—122. DOI: 10.1561/2200000016 Training Neural Networks Without Gradients: A Scalable {ADMM} Approach Taylor, G., Burmeister, R., Xu, Z., Singh, B., Patel, A. and Goldstein, T., 2016. Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. On {ADMM} in Deep Learning: Convergence and Saturation-Avoidance Zeng, J., Lin, S., Yao, Y. and Zhou, D., 2021. Journal of Machine Learning Research, Vol 22. A Theoretical Framework for Back-Propagation LeCun, Y., 1988. Distributed Optimization of Deeply Nested Systems Carreira-Perpinan, M.A. and Wang, W., 2014. Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, PMLR 33. {ADMM} for Efficient Deep Learning with Global Convergence  [link] Wang, J., Yu, F., Chen, X. and Zhao, L., 2019. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery \& Data Mining, pp. 111–119. Association for Computing Machinery. DOI: 10.1145/3292500.3330936 Decoupling Backpropagation using Constrained Optimization Methods  [link] Gotmare, A., Thomas, V., Brea, J. and Jaggi, M., 2018. ICML 2018 Workshop on Credit Assignment in Deep Learning and Deep Reinforcement Learning. Proximal Backpropagation  [link] Frerix, T., Mollenhoff, T., Moeller, M. and Cremers, D., 2018. International Conference on Learning Representations. Lifted Neural Networks Askari, A., Negiar, G., Sambharya, R. and El Ghaoui, L., 2018. arXiv preprint arXiv:1805.01532. Lifted Proximal Operator Machines Li, J., Fang, C. and Lin, Z., 2019. Proceedings of the AAAI Conference on Artificial Intelligence, Vol 33, pp. 4181—4188. Fenchel Lifted Networks: A {L}agrange Relaxation of Neural Network Training Gu, F., Askari, A. and El Ghaoui, L., 2020. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, PMLR 108. Contrastive Learning for Lifted Networks Zach, C. and Estellers, V., 2019. British Machine Vision Conference (BMVC). Lifted {B}regman Training of Neural Networks Wang, X. and Benning, M., 2023. Journal of Machine Learning Research, Vol 24. A Unified Framework for Lifted Training and Inversion Approaches Wang, X., Valavanis, A., Mahmood, A., Mang, A., Benning, M. and Repetti, A., 2025. arXiv preprint arXiv:2510.09796. Neural Network Training as an Optimal Control Problem : — An Augmented Lagrangian Approach —  [link] Evens, B., Latafat, P., Themelis, A., Suykens, J. and Patrinos, P., 2021. 2021 60th IEEE Conference on Decision and Control (CDC), pp. 5136–5143. IEEE. DOI: 10.1109/cdc45484.2021.9682842 An Augmented Lagrangian Method for Training Recurrent Neural Networks  [link] Wang, Y., Zhang, C. and Chen, X., 2025. SIAM Journal on Scientific Computing, Vol 47(1), pp. C22-C51. DOI: 10.1137/23M1627614 Inferring Neural Activity Before Plasticity as a Foundation for Learning Beyond Backpropagation Song, Y., Millidge, B., Salvatori, T., Lukasiewicz, T., Xu, Z. and Bogacz, R., 2024. Nature Neuroscience. DOI: 10.1038/s41593-023-01514-1 Predictive Coding Networks for Temporal Prediction Millidge, B., Tang, M., Osanlouy, M., Harper, N.S. and Bogacz, R., 2024. PLOS Computational Biology, Vol 20(4), pp. e1011183. DOI: 10.1371/journal.pcbi.1011183 Learning Complex Temporal Dependencies via Local Synaptic Plasticity Ng-Kee-Kwong, J., Tang, M., Akam, T. and Bogacz, R., 2026. bioRxiv preprint. DOI: 10.64898/2026.07.09.737423 Blockwise Self-Supervised Learning at Scale Siddiqui, S.A., Krueger, D., LeCun, Y. and Deny, S., 2024. Transactions on Machine Learning Research.