When Learning Reproduces the Quantile: A Theoretical Analysis of Logistic-Regression Threshold Selection and the Role of Bounded Truncation in Extreme Insurance Loss Modelling

Auteurs

  • Amina NADIM Faculté des sciences Juridiques, Economiques et Sociales Ain sebaa, Université Hassan II, Casablanca, Maroc
  • Tarek ZARI Faculté des sciences Juridiques, Economiques et Sociales Ain sebaa, Université Hassan II, Casablanca, Maroc

Mots-clés :

Extreme value theory; threshold selection; peaks over threshold; Hill estimator; logistic regression; generalized Pareto distribution; non-life insurance; Value-at-Risk; Expected Shortfall

Résumé

The choice of the threshold above which a loss is treated as extreme remains the central open problem in applying peaks-over-threshold methods to non-life insurance, governing a bias-variance trade-off with direct consequences for reserves, premium loading, and solvency capital. Since the foundational results of Pickands (1975), Balkema and de Haan (1974), and Hill (1975), threshold selection has relied mainly on graphical diagnostics and bias-correction methods, none of which resolve the choice of threshold itself; recent work has proposed automating this choice through supervised classifiers, most notably logistic regression, on the claim that this yields an objective, reproducible threshold. This paper makes two contributions that reposition that claim. First, working within the standard semi-parametric framework of extreme value theory, we establish the paper's central theoretical result (Proposition 1): when the sole covariate is the loss amount and labels are generated by an empirical quantile rule, the threshold returned by logistic-regression selection converges asymptotically, under stated regularity conditions, to the labelling quantile itself, so that the procedure recovers, rather than discovers, a threshold fixed implicitly at the labelling stage. Second, we extend Hill's (1975) estimator, via the bounded-domain correction of Nuyts (2010), to contractually right-truncated insurance losses, a structural feature of insurance data left unaddressed by prevailing threshold-selection methods. The analysis culminates in a unified decision framework situating heuristic, statistical, graphical, and learning-based methods within a common bias-variance logic. The contribution is conceptual and theoretical.

Classification JEL: C13; C24; G22

Paper type: Theoretical Research

Téléchargements

Publiée

2026-08-06

Comment citer

NADIM, A., & ZARI, T. (2026). When Learning Reproduces the Quantile: A Theoretical Analysis of Logistic-Regression Threshold Selection and the Role of Bounded Truncation in Extreme Insurance Loss Modelling. International Journal of Accounting, Finance, Auditing, Management and Economics, 7(9), 361–375. Consulté à l’adresse https://ijafame.org/index.php/ijafame/article/view/2557

Numéro

Rubrique

Articles