Over the past 20 years Artificial intelligence (AI) has become an important cornerstone of Responsible Gambling (RG) efforts, with automated systems now central to detecting players at risk of harm. Millions of people globally participate in gambling, and AI has become a key technology operators use to identify problematic behaviour through behavioural markers to trigger early interventions.
This technology is having an important impact on the industry and has increasingly become embedded in gambling regulations globally, yet the field faces a profound and worrying challenge: the effectiveness of these AI systems is fundamentally unmeasured.
To address the issue of transparency in these systems we are pleases to collaborate with an experienced research team to publish a pre-print of a new paper titled The Need for Benchmarks to Advance AI-Enabled Player Risk Detection in Gambling.
The Measurement Crisis: Why We Can’t Compare AI Systems
The core issue is one of transparency. We have arguably powerful, predictive systems deployed to protect vulnerable individuals, but we have no standardized, objective way to verify if they actually work—or to compare one system against another. This is the AI Paradox of Responsible Gambling: the technology designed to increase safety has become opaque precisely when transparency is most crucial.
Today, different operators use proprietary, siloed systems; vendors make performance claims that cannot be independently verified; and regulators lack an objective baseline to evaluate approved tools. This fragmentation means that when a new model is introduced, the industry cannot confidently determine if it represents genuine progress or simply marketing noise. For researchers, the absence of a shared benchmark makes it impossible to build cumulatively on prior work. The result is a knowledge vacuum, preventing real innovation and progress in harm prevention.
The next critical innovation required is not the advancement of current methods or the creation of new models, but the development of a framework to measure them.
The Critical Need for Benchmarking
This is where the concept of benchmarking comes in. The paper argues that a conceptual framework for benchmarking is the necessary next step to advance the field. In this context, benchmarking means the structured and repeatable assessment of AI models using three crucial elements:
- Standardized Datasets: A common pool of anonymized player data for testing.
- Clearly Defined Tasks: Precise definitions of what the model must predict (e.g., predicting a level of risk).
- Agreed-Upon Performance Metrics: Universal metrics to measure accuracy, sensitivity (catching true risk), and precision (avoiding false alarms).
The goal of this framework is to enable objective, comparable, and longitudinal evaluation of player risk detection systems. Without shared metrics, operators are left with a difficult trade-off between sensitivity and precision; an inaccurate model that cries “wolf” too often can lead to unnecessary intervention, which may undermine player trust in responsible gambling programs.