Avoiding Overfitting in Variable-Order Markov Models: a Cross-Validation Approach

Secchini, Valeria; Garcia-Bernardo, Javier; Janský, Petr

物理学 > 物理与社会

arXiv:2501.14476v1 (physics)

[提交于 2025年1月24日 ]

标题：避免变量阶马尔可夫模型中的过拟合：一种交叉验证方法

标题： Avoiding Overfitting in Variable-Order Markov Models: a Cross-Validation Approach

Authors:Valeria Secchini, Javier Garcia-Bernardo, Petr Janský

摘要： Higher$\text{-}$order Markov chain models are widely used to represent agent transitions in dynamic systems, such as passengers in transport networks. They capture transitions in complex systems by considering not only the current state but also the path of previously visited states. For example, the likelihood of train passengers traveling from Paris (current state) to Rome could increase significantly if their journey originated in Italy (prior state). Although this approach provides a more faithful representation of the system than first$\text{-}$order models, we find that commonly used methods$-$relying on Kullback$\text{-}$Leibler divergence$-$frequently overfit the data, mistaking fluctuations for higher$\text{-}$order dependencies and undermining forecasts and resource allocation. Here, we introduce DIVOP (Detection of Informative Variable$\text{-}$Order Paths), an algorithm that employs cross$\text{-}$validation to robustly distinguish meaningful higher$\text{-}$order dependencies from noise. In both synthetic and real$\text{-}$world datasets, DIVOP outperforms two state$\text{-}$of$\text{-}$the$\text{-}$art algorithms by achieving higher precision, recall, and sparser representations of the underlying dynamics. When applied to global corporate ownership data, DIVOP reveals that tax havens appear in 82$\%$ of all significant higher$\text{-}$order dependencies, underscoring their outsized influence in corporate networks. By mitigating overfitting, DIVOP enables more reliable multi$\text{-}$step predictions and decision$\text{-}$making, paving the way toward deeper insights into the hidden structures that drive modern interconnected systems.

主题：	物理与社会 (physics.soc-ph) ; 社会与信息网络 (cs.SI); 一般经济学 (econ.GN)
引用方式：	arXiv:2501.14476 [physics.soc-ph]
	(或者 arXiv:2501.14476v1 [physics.soc-ph] 对于此版本)
	https://doi.org/10.48550/arXiv.2501.14476

提交历史

来自： Valeria Secchini [查看电子邮件]
[v1] 星期五， 2025 年 1 月 24 日 13:18:11 UTC (11,651 KB)

物理学 > 物理与社会

标题：避免变量阶马尔可夫模型中的过拟合：一种交叉验证方法

标题： Avoiding Overfitting in Variable-Order Markov Models: a Cross-Validation Approach

提交历史

获取论文：

参考文献与引用

收藏

文献和引用工具

与本文相关的代码，数据和媒体

演示

推荐器和搜索工具

arXivLabs：与社区合作伙伴的实验项目

物理学 > 物理与社会

标题： 避免变量阶马尔可夫模型中的过拟合：一种交叉验证方法 显示英文标题

标题： Avoiding Overfitting in Variable-Order Markov Models: a Cross-Validation Approach

提交历史

获取论文：

参考文献与引用

BibTeX 格式的引用

收藏

文献和引用工具

与本文相关的代码，数据和媒体

演示

推荐器和搜索工具

arXivLabs：与社区合作伙伴的实验项目

标题：避免变量阶马尔可夫模型中的过拟合：一种交叉验证方法