<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with OASIS Tables with MathML3 v1.4 20241031//EN" "https://jats.nlm.nih.gov/archiving/1.4/JATS-archive-oasis-article1-4-mathml3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" dtd-version="1.4" article-type="research-article" xml:lang="en"><front><journal-meta><journal-title-group><journal-title xml:lang="ru">Успехи кибернетики</journal-title></journal-title-group><issn publication-format="electronic">2712-9942</issn></journal-meta><article-meta><article-categories><subj-group><subject>Other</subject></subj-group></article-categories><title-group><article-title xml:lang="ru">Ансамблевый метод глубокого обучения с подкреплением для управления инвестиционным портфелем на российском фондовом рынке</article-title><trans-title-group xml:lang="en"><trans-title>Ensemble Deep Reinforcement Learning Approach for Portfolio Management in the Russian Equity Market</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Кобзев</surname><given-names>А. А.</given-names></name><name xml:lang="en"><surname>Kobzev</surname><given-names>A. A.</given-names></name></name-alternatives><xref ref-type="aff" rid="aff1"/><xref ref-type="aff" rid="aff2"/><email>artem.kobzev.2001@mail.ru</email><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0009-2369-1741</contrib-id></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Крахмалев</surname><given-names>О. Н.</given-names></name><name xml:lang="en"><surname>Krakhmalev</surname><given-names>O. N.</given-names></name></name-alternatives><xref ref-type="aff" rid="aff1"/><xref ref-type="aff" rid="aff2"/><email>onkrakhmalev@fa.ru</email><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-9388-4137</contrib-id></contrib><aff-alternatives id="aff1"><aff><institution xml:lang="en">Financial University under the Government of the Russian Federation</institution><city xml:lang="en">Moscow</city><country xml:lang="en">Russian Federation</country></aff></aff-alternatives><aff-alternatives id="aff2"><aff><institution xml:lang="ru">Финансовый университет при Правительстве Российской Федерации</institution><city xml:lang="ru">Москва</city><country xml:lang="ru">Российская Федерация</country></aff></aff-alternatives></contrib-group><pub-date pub-type="epub" iso-8601-date="2026-06-30"><day>30</day><month>06</month><year>2026</year></pub-date><volume>7</volume><issue>2</issue><fpage>119</fpage><lpage>125</lpage><history><date date-type="received" iso-8601-date="2026-04-06"><day>06</day><month>04</month><year>2026</year></date><date date-type="accepted" iso-8601-date="2026-05-14"><day>14</day><month>05</month><year>2026</year></date></history><self-uri xlink:href="https://ru.jcyb.ru/nisii_tech/article/view/509" xlink:title="https://ru.jcyb.ru/nisii_tech/article/view/509">https://ru.jcyb.ru/nisii_tech/article/view/509</self-uri><self-uri content-type="pdf" xlink:href="publication-f84b97a5-e3af-42d7-8779-3d424934f421.pdf" xlink:title="PDF"/><abstract xml:lang="ru"><p>в статье исследуется ансамблевый подход к управлению инвестиционным портфелем на российском фондовом рынке на основе глубокого обучения с подкреплением. Цель работы состоит в воспроизведении базовой ансамблевой архитектуры на данных российского рынка, в проверке ее переносимости и в выявлении тех модификаций, которые, действительно, улучшают качество торговли на длительный период. В качестве исходного решения рассматривается ансамбль из трех алгоритмов принятия решений, для которого последовательно анализируются расширение набора признаков, добавление макроэкономических переменных, штрафов за риск и механизма непрерывной адаптации одного из агентов в процессе торговли. Эксперименты проведены на данных 2015–2025 годов по ликвидным российским акциям, а итоговое сравнение выполнено на периоде 2023–2025 годов. Показано, что наибольший вклад в результат дает механизм непрерывного обучения, при котором агент дообучается на сделках любого активного участника ансамбля. Лучшая конфигурация обеспечивает суммарную доходность 61,7 процента и превосходит пассивные ориентиры по абсолютной доходности, однако не решает проблему защиты капитала на затяжном падающем рынке при высокой ключевой ставке.</p></abstract><abstract xml:lang="en" abstract-type="summary"><p>we studied an ensemble-based deep reinforcement learning approach to portfolio management in the Russian stock market. The study aimed to reproduce a baseline ensemble architecture using Russian equity market data, evaluate its transferability, and identify modifications that improve trading performance over a long investment horizon. We implemented an initial model consisting of three decision-making agents and then extended the analysis by incorporating a broader set of features, macroeconomic indicators, risk-adjusted reward functions, and a mechanism for continuous adaptation of one agent during live trading. We trained and evaluated the models on data from 2015 to 2025 for liquid Russian equities, and we conducted the final performance comparison on the out-of-sample period from 2023 to 2025. The results show that the main driver of performance improvement is continuous off-policy training, where one agent is updated using trading data generated by any active agent in the ensemble. The best-performing configuration achieves a cumulative return of 61.7 percent and outperforms passive benchmark strategies in absolute return. However, the results also reveal a structural limitation: a long-only ensemble without an explicit allocation mechanism to low-risk assets does not provide sufficient capital protection during prolonged bear market conditions combined with high interest rates.</p></abstract><kwd-group xml:lang="ru"><kwd>глубокое обучение с подкреплением</kwd><kwd>управление инвестиционным портфелем</kwd><kwd>ансамблевые торговые стратегии</kwd><kwd>непрерывная адаптация</kwd><kwd>управление риском</kwd><kwd>российский фондовый рынок</kwd></kwd-group><kwd-group xml:lang="en"><kwd>deep reinforcement learning</kwd><kwd>portfolio management</kwd><kwd>ensemble trading strategies</kwd><kwd>continuous adaptation</kwd><kwd>risk-aware optimization</kwd><kwd>Russian stock market</kwd></kwd-group></article-meta></front><back><ref-list><ref id="ref1"><mixed-citation publication-type="other" xml:lang="ru">Markowitz H. M. Portfolio Selection. The Journal of Finance. 1952;7(1):77–91. DOI: 10.1111/j.1540-6261.1952.tb01525.x.</mixed-citation></ref><ref id="ref2"><mixed-citation publication-type="other" xml:lang="ru">Black F., Litterman R. Global Portfolio Optimization. Financial Analysts Journal. 1992;48(5):28–43. DOI: 10.2469/faj.v48.n5.28.</mixed-citation></ref><ref id="ref3"><mixed-citation publication-type="other" xml:lang="ru">Yang Z., Zhu Y., Guo J., Liu X.-Y., Zhong S., Walid A. Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy. ICAIF ’20. 2020:1–8. DOI: 10.1145/3383455.3422540.</mixed-citation></ref><ref id="ref4"><mixed-citation publication-type="other" xml:lang="ru">Mnih V., Kavukcuoglu K., Silver D. et al. Human-Level Control through Deep Reinforcement Learning. Nature. 2015;518:529–533. DOI: 10.1038/nature14236.</mixed-citation></ref><ref id="ref5"><mixed-citation publication-type="other" xml:lang="ru">Lillicrap T. P., Hunt J. J., Pritzel A. et al. Continuous Control with Deep Reinforcement Learning. arXiv:1509.02971. 2016. DOI: 10.48550/arXiv.1509.02971.</mixed-citation></ref><ref id="ref6"><mixed-citation publication-type="other" xml:lang="ru">Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347. 2017. DOI: 10.48550/arXiv.1707.06347.</mixed-citation></ref><ref id="ref7"><mixed-citation publication-type="other" xml:lang="ru">Haarnoja T., Zhou A., Abbeel P., Levine S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. Proceedings of ICML. 2018:1861–1870.</mixed-citation></ref><ref id="ref8"><mixed-citation publication-type="other" xml:lang="ru">Liu S., Rui J., Gao J. et al. FinRL-Meta: Market Environments and Benchmarks for Data-Driven Financial Reinforcement Learning. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).</mixed-citation></ref><ref id="ref9"><mixed-citation publication-type="other" xml:lang="ru">Theate T., Ernst D. An Application of Deep Reinforcement Learning to Algorithmic Trading. Expert Systems with Applications. 2021;173:114632. DOI: 10.1016/j.eswa.2021.114632.</mixed-citation></ref><ref id="ref10"><mixed-citation publication-type="other" xml:lang="ru">Khetarpal K., Riemer M., Rish I., Precup D. Towards Continual Reinforcement Learning: A Review and Perspectives. Journal of Artificial Intelligence Research. 2022;75:1401–1476. DOI: 10.1613/jair.1.13673.</mixed-citation></ref><ref id="ref11"><mixed-citation publication-type="other" xml:lang="ru">Artzner P., Delbaen F., Eber J.-M., Heath D. Coherent Measures of Risk. Mathematical Finance. 1999;9(3):203–228. DOI: 10.1111/1467-9965.00068.</mixed-citation></ref><ref id="ref12"><mixed-citation publication-type="other" xml:lang="ru">Rockafellar R. T., Uryasev S. Conditional Value-at-Risk for General Loss Distributions. Journal of Banking &amp; Finance. 2002;26(7):1443–1471. DOI: 10.1016/S0378-4266(02)00271-6.</mixed-citation></ref></ref-list></back></article>
