MATHEMATICAL ANALYSIS OF A TRACE AWARE REINFORCEMENT ARCHITECTURE FOR MOBILE EDGE COMPUTING
Main Article Content
Abstract
This paper develops an applied mathematical framework for reinforcement learning in Mobile Edge Computing (MEC). The offloading process is formulated as a Markov Decision Process (MDP), where the reward function balances delay and energy consumption. Eligibility traces are introduced to improve temporal credit assignment, formalized as exponential decay operators that propagate influence across state–action pairs. A dual‑head value decomposition, is analyzed to reduce estimator variance and sharpen action discrimination. Theoretical results establish convergence properties under bounded rewards and finite state–action spaces, with proofs based on contraction mappings of the Bellman operator. Compact experiments confirm that the proposed Trace‑Aware Reinforcement Architecture (TARA) achieves smoother convergence and improved delay–energy trade‑offs compared to baseline methods. Overall, this study demonstrates how applied mathematical tools—Poisson processes, stochastic analysis, and operator theory—can rigorously model reinforcement learning in dynamic MEC environments, providing both theoretical guarantees and practical insights for distributed computing systems.