We assess the ability of Reinforcement Learning (RL) agents to form reliable, consistent market views by constructing four families of agents tasked with generating views to dynamically manage a diversified equity portfolio via the Black-Litterman (BL) model. Using narrow and wide datasets, we explore the impact of information density on View Generation (VG). We evaluate how Artificial Neural Network (ANN) architecture and temporal awareness (look-back windows) affect VG, and compare performance across on-policy (PPO) and off-policy (SAC) algorithms. We find that RL agents are able to conduct VG reliably and consistently provided they possess an information-dense observation space, an optimal ANN architecture, and sufficient temporal awareness. Both algorithm classes train functional policies provided adequate reward functions are used in each case; specifically, on-policy PPO requires a highly specific reward function, called IRR, whereas off-policy SAC performs best with a generic one. We conclude that RL agents are able to dynamically manage portfolios through the BL model which are able to outperform benchmarks reliably. However, they seem to suffer from shifts in market regime as posited by the Adaptive Market Hypothesis (AMH). Furthermore, we establish that our framework is valid to conduct empirical tests of both the AMH and the Efficient Market Hypothesis (EMH).
Can Reinforcement Learning Agents Generate Reliable Market Views? Some evidence from dynamic portfolio management through a RL-BL hybrid approach.
ANDRETTA, EDOARDO
2025/2026
Abstract
We assess the ability of Reinforcement Learning (RL) agents to form reliable, consistent market views by constructing four families of agents tasked with generating views to dynamically manage a diversified equity portfolio via the Black-Litterman (BL) model. Using narrow and wide datasets, we explore the impact of information density on View Generation (VG). We evaluate how Artificial Neural Network (ANN) architecture and temporal awareness (look-back windows) affect VG, and compare performance across on-policy (PPO) and off-policy (SAC) algorithms. We find that RL agents are able to conduct VG reliably and consistently provided they possess an information-dense observation space, an optimal ANN architecture, and sufficient temporal awareness. Both algorithm classes train functional policies provided adequate reward functions are used in each case; specifically, on-policy PPO requires a highly specific reward function, called IRR, whereas off-policy SAC performs best with a generic one. We conclude that RL agents are able to dynamically manage portfolios through the BL model which are able to outperform benchmarks reliably. However, they seem to suffer from shifts in market regime as posited by the Adaptive Market Hypothesis (AMH). Furthermore, we establish that our framework is valid to conduct empirical tests of both the AMH and the Efficient Market Hypothesis (EMH).| File | Dimensione | Formato | |
|---|---|---|---|
|
ANDRETTA_EDOARDO_888603.pdf
accesso aperto
Dimensione
11.27 MB
Formato
Adobe PDF
|
11.27 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14247/29941