Levenberg-Marquardt: A hybrid approach for improving performance in neural networks

Authors

DOI:

https://doi.org/10.71112/fdzkm898

Keywords:

Levenberg–Marquardt, Neural Networks, Optimization, Convergence, Supervised Learning

Abstract

This work analyzes the application of the Levenberg–Marquardt (LM) algorithm as an optimization method for training multilayer neural networks. This approach combines the speed of the Gauss–Newton method with the stability of gradient descent, dynamically adjusting a damping parameter that regulates its behavior. A theoretical framework was developed to explain its mathematical formulation based on the Taylor expansion, the computation of the gradient and the Jacobian, and its adaptation to supervised learning. Experimental results show that LM significantly improves both the convergence speed and the accuracy of the model compared to first-order methods, achieving a noticeably lower mean squared error. This confirms its effectiveness in optimizing highly nonlinear and complex functions in neural networks.

Downloads

Download data is not yet available.

References

A. Ranganath, I. R. (2024). Quasi-Adam: Accelerating Adam Using Quasi-Newton Approximations. 2024 International Conference on Machine Learning and Applications (ICMLA). https://doi.org/10.1109/icmla61862.2024.00107 DOI: https://doi.org/10.1109/ICMLA61862.2024.00107

Adhirai Subramaniyam, R. M. (2019). Taylor and Gradient Descent-Based Actor Critic Neural Network for the Classification of Privacy Preserved Medical Data. Big data. https://doi.org/10.1089/big.2018.0166 DOI: https://doi.org/10.1089/big.2018.0166

Adrien B. Taylor, J. H. (2015). Exact Worst-Case Performance of First-Order Methods for Composite Convex Optimization. SIAM J. Optim. https://doi.org/10.1137/16m108104x DOI: https://doi.org/10.1137/16M108104X

Adrien B. Taylor, Y. D. (2021). An optimal gradient method for smooth strongly convex minimization. Mathematical Programming. https://doi.org/10.1007/s10107-022-01839-y DOI: https://doi.org/10.1007/s10107-022-01839-y

Alexandru Damian, E. N. (2022). Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability. ArXiv. https://doi.org/10.48550/arxiv.2209.15594

Antoine Lesage-Landry, J. A. (2020). Second-Order Online Nonconvex Optimization. IEEE Transactions on Automatic Control. https://doi.org/10.1109/tac.2020.3040372 DOI: https://doi.org/10.1109/TAC.2020.3040372

Aryan Mokhtari, Q. L. (2017). Network Newton Distributed Optimization Methods. IEEE Transactions on Signal Processing. https://doi.org/10.1109/tsp.2016.2617829 DOI: https://doi.org/10.1109/TSP.2016.2617829

C. Uriarte, M. B. (2024). Optimizing Variational Physics-Informed Neural Networks Using Least Squares. ArXiv. https://doi.org/10.1016/j.camwa.2025.02.022 DOI: https://doi.org/10.1016/j.camwa.2025.02.022

Chunfu Guo, Y. S. (2023). A combination of library search and Levenberg-Marquardt algorithm in optical scatterometry. Thin Solid Films. https://doi.org/10.1016/j.tsf.2023.139670 DOI: https://doi.org/10.1016/j.tsf.2023.139670

D. Rotondo, M. J. (2024). A Weighted Linearization Approach to Gradient Descent Optimization. 2024 European Control Conference (ECC). https://doi.org/10.23919/ecc64448.2024.10590954 DOI: https://doi.org/10.23919/ECC64448.2024.10590954

Damien Lee, K. N.-H.-W. (2023). A Continual Learning Algorithm Based on Orthogonal Gradient Descent Beyond Neural Tangent Kernel Regime. IEEE Access. https://doi.org/10.1109/access.2023.3303869 DOI: https://doi.org/10.1109/ACCESS.2023.3303869

Dongsheng Guo, Z.‐Y. N. (2017). Novel Discrete-Time Zhang Neural Network for Time-Varying Matrix Inversion. IEEE Transactions on Systems, Man, and Cybernetics: Systems. https://doi.org/10.1109/tsmc.2017.2656941 DOI: https://doi.org/10.1109/TSMC.2017.2656941

Gleich, D. (2017). TRUST REGION METHODS. https://doi.org/10.1007/978-0-387-40065-5_4 DOI: https://doi.org/10.1007/978-0-387-40065-5_4

Hawraz N. Jabbar, B. A. (2021). Two-versions of descent conjugate gradient methods for large-scale unconstrained optimization. Indonesian Journal of Electrical Engineering and Computer Science. https://doi.org/10.11591/ijeecs.v22.i3.pp1643-1649 DOI: https://doi.org/10.11591/ijeecs.v22.i3.pp1643-1649

Huanshui Zhang, H. W. (2024). Optimization methods rooted in optimal control. Sci. China Inf. Sci. https://doi.org/10.1007/s11432-024-4207-5 DOI: https://doi.org/10.1007/s11432-024-4207-5

Ivan Perez Avellaneda, L. A. (2023). Output Reachability of Chen-Fliess series: A Newton-Raphson Approach. 2023 57th Annual Conference on Information Sciences and Systems (CISS). https://doi.org/10.1109/ciss56502.2023.10089740 DOI: https://doi.org/10.1109/CISS56502.2023.10089740

Jaehoon Lee, L. X.-D. (2019). Wide neural networks of any depth evolve as linear models under gradient descent. Journal of Statistical Mechanics: Theory and Experiment. https://doi.org/10.1088/1742-5468/abc62b DOI: https://doi.org/10.1088/1742-5468/abc62b

Jamie M. Taylor, D. P. (2022). A Deep Fourier Residual Method for solving PDEs using Neural Networks. ArXiv. https://doi.org/10.48550/arxiv.2210.14129 DOI: https://doi.org/10.1016/j.cma.2022.115850

Jesse Chan, C. G. (2020). Efficient computation of Jacobian matrices for entropy stable summation-by-parts schemes. J. Comput. Phys. https://doi.org/10.1016/j.jcp.2021.110701 DOI: https://doi.org/10.1016/j.jcp.2021.110701

John Taylor, W. W. (2022). Optimizing the optimizer for data driven deep neural networks and physics informed neural networks. ArXiv. https://doi.org/10.48550/arxiv.2205.07430

Marcus Carlsson, V. N. (2025). The bilinear Hessian for large scale optimization. ArXiv. https://doi.org/10.48550/arxiv.2502.03070

Mou Wu, N. X. (2020). RNN-K: A Reinforced Newton Method for Consensus-Based Distributed Optimization and Control Over Multiagent Systems. IEEE Transactions on Cybernetics. https://doi.org/10.1109/tcyb.2020.3011819 DOI: https://doi.org/10.1109/TCYB.2020.3011819

P. Baldi, K. C. (2016). Parameterized neural networks for high-energy physics. The European Physical Journal C. https://doi.org/10.1140/epjc/s10052-016-4099-4 DOI: https://doi.org/10.1140/epjc/s10052-016-4099-4

R. B. Yunus, N. Z.-Y. (2024). An improved accelerated 3-term conjugate gradient algorithm with second-order Hessian approximation for nonlinear least-squares optimization. Journal of Mathematics and Computer Science. https://doi.org/10.22436/jmcs.036.03.02 DOI: https://doi.org/10.22436/jmcs.036.03.02

R. Dehghani, M. M. (2018). The modified quasi‐Newton methods for solving unconstrained optimization problems. International Journal of Numerical Modelling: Electronic Networks. https://doi.org/10.1002/jnm.2459 DOI: https://doi.org/10.1002/jnm.2459

Shuvomoy Das Gupta, R. F. (2023). Nonlinear conjugate gradient methods: worst-case convergence rates via computer-assisted analyses. Mathematical Programming. https://doi.org/10.1007/s10107-024-02127-7 DOI: https://doi.org/10.1007/s10107-024-02127-7

Taylor Simons, D.-J. L. (2019). A Review of Binarized Neural Networks. Electronics. https://doi.org/10.3390/electronics8060661 DOI: https://doi.org/10.3390/electronics8060661

Weilong Liao, Q. Z. (2022). An SQP Algorithm for Structural Topology Optimization Based on Majorization–Minimization Method. Applied Sciences. https://doi.org/10.3390/app12136304 DOI: https://doi.org/10.3390/app12136304

Yuan-Yuan Cheng, L. L.-M. (2025). Eigenvalue Blending for Projected Newton. Computer Graphics Forum. https://doi.org/10.1111/cgf.70027 DOI: https://doi.org/10.1111/cgf.70027

Yunjin Tong, S. X. (2020). Symplectic Neural Networks in Taylor Series Form for Hamiltonian Systems. ArXiv. https://doi.org/10.1016/j.jcp.2021.110325 DOI: https://doi.org/10.1016/j.jcp.2021.110325

Zheyuan Hu, Z. S. (2023). Hutchinson Trace Estimation for High-Dimensional and High-Order Physics-Informed Neural Networks. ArXiv. https://doi.org/10.1016/j.cma.2024.116883 DOI: https://doi.org/10.1016/j.cma.2024.116883

Published

2026-07-02

Issue

Section

Computational Sciences

How to Cite

Morales Rodríguez, E. A. (2026). Levenberg-Marquardt: A hybrid approach for improving performance in neural networks. Multidisciplinary Journal Epistemology of the Sciences, 3(3 Edición Especial), 88-112. https://doi.org/10.71112/fdzkm898