
This past week, I attended NeurIPS 2025 (Conference on Neural Information Processing Systems) in San Diego. For those unfamiliar, NeurIPS is the premier annual event on the artificial intelligence and machine learning calendar. It’s a gathering of both the academic and industrial research communities that presents the latest innovations and, more importantly, sets the research agenda for years to come.
This was my first NeurIPS and I was astounded by the sheer amount of content. Tutorials, workshops, keynotes, oral presentations, and posters… wow, the poster sessions! There were two 3-hour poster sessions across three days of the conference in a room that would rival the size of the G2E or ICE exhibit halls. See the pic below—there must have been at least 1,000 posters in each session!

While there is everything from robotics to computer vision, and math theory to deep learning on display, over this post I will highlight my top 3 presentations from the week and contextualize what these might mean for the gambling industry.
1. The Golden Era of Benchmarking (and Why Gambling Needs to Catch Up)
NeurIPS Session: Tutorial — The Science of Benchmarking: What’s Measured, What’s Missed, and What’s Next
Read more: https://benchmarking.science/
This tutorial was incredibly timely. One of the prevailing sentiments I heard throughout the week is that the AI/ML community is currently living through “the golden era” of benchmarking. As models become more capable, the science of how the community evaluates them is moving rapidly to keep up.
This stands in stark contrast to where we are in the gambling industry. While the wider AI community (as well as domains like healthcare and finance) is debating the nuances of evaluation, in gambling, the field hasn’t reached any era of benchmarking yet—we’re (arguably) at ground zero! Myself and co-authors have recently taken a step to address this gap regarding player risk detection, and you can read more about our position and vision in our recent preprint: arXiv:2511.21658. But going back to this session, two critical takeaways stood out to me about benchmarking in the AI/ML community that we in the gambling field should be aware of:
i) Moving from “Task-Driven” to “Capability-Driven” Benchmarks There is a shift away from testing isolated tasks toward “real-world” or “capability” benchmarks. Here I provide a few examples for the interested reader:
OpenAI’s GDPval (openai.com/index/gdpval/) evaluates how well agents perform economic tasks that actually contribute to GDP.
VendingBench (andonlabs.com/evals/vending-bench) tests an agent’s ability to maintain a vending machine business over time.
ARC AGI (https://arcprize.org/arc-agi) evaluates an agent’s ability to efficiently learn new skills and solve open-ended reasoning puzzles.
While our recent paper reflects on a specific task (player risk detection), this tutorial inspired me to think bigger. Perhaps the next should be a capability benchmark for an AI agent that manages the entire player protection pipeline. It’s not so hard to imagine a near future where a multi-agent framework handles everything from model development and implementation to interpreting outputs and even interacting with players directly…and we’ll need a way to evaluate that system.
ii) The Critical Importance of Construct Validity Another major issue raised during this tutorial—and brilliantly articulated in a related poster I saw by Andrew M. Bean (University of Oxford)—is the concept of Construct Validity (arXiv:2511.04703).
Simply put: Are we actually measuring what we think we are measuring with these benchmarks? Bean’s research provides quantitative evidence that this is majorly lacking in current AI benchmarks. So, as we attempt to adopt AI evaluation in the gambling sector—whether for player risk detection, fraud detection, or AI chatbots for treatment/counseling—we must ensure we have construct validity. If our metric says a model is “safe” or “accurate”, does it actually correspond to real-world player safety and accuracy?
2. The Artificial Hivemind: Will AI Kill Differentiation in Gambling?
NeurIPS Session: Oral Presentation — Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
Read more: https://arxiviq.substack.com/p/neurips-2025-artificial-hivemind
This presentation, which won a best paper award, detailed a study on how LLMs behave “in the wild”. The authors analyzed open-ended interactions. So not simple multiple-choice questions, they looked at how models respond to more complex tasks like planning, problem-solving, and brainstorming—the exact things you likely use your favorite chatbot for every day. The findings were quite shocking…these models suffer from severe “intra-model” and “inter-model” repetition.
In simple terms, not only does a single model tend to repeat itself, but different models from different companies are starting to sound exactly the same. They are converging on a “mean” standard of output.
What This Means for Gambling: GenAI is currently being deployed to assist with marketing copy, promotional materials, game assets, and perhaps even underlying game mechanics. But if everyone is leveraging the same foundation models and prompting them with similar instructions—for example, “write an engaging retention email for a VIP player” or “design a new slot theme based on mythology”—does this put brand differentiation at risk?
Just for fun, I entered the prompt “Write an engaging retention email for a VIP casino player” into both ChatGPT and Claude. Look at the opening sentence from each:
ChatGPT: The VIP Lounge hasn’t been the same without you. As one of our most valued players, you’ve always brought energy, excitement, and a little bit of magic every time you joined us — and we’d love to welcome you back.
Claude: We’ve noticed you haven’t visited us in a while, and frankly, the tables haven’t been the same without you. As one of our most valued players, your experience matters deeply to us.
This is a simple example, but you can see the exact same words are used. And it’s not all about specific words, the structure and tone of outputs can also converge to similarity. Brands may want to consider how best to leverage this technology…while it can help be more efficient is this at the expense of something greater?
3. Preventing “Digital Heroin”: What AI Safety Can Learn from Responsible Gambling
NeurIPS Session: Oral Presentation — Real-Time Hyper-Personalized Generative AI Should Be Regulated to Prevent the Rise of “Digital Heroin”
Read more: https://openreview.net/forum?id=1IpHkK5Q8F
For my final highlight, I’m flipping things. During this presentation, I realized there is an opportunity for the AI field to learn from us.
Soon AI will be capable of generating real-time, hyper-personalized content designed specifically to maximize user engagement, creating a new class of addictive media. While the problem statement during this presentation was spot on, the proposed solutions felt all too familiar to anyone in our sector. The presenter recommended interventions like time-outs and behavioral analytics based on “warning signs”.
These are the tools we are all familiar with. But as we know in the gambling industry, the mere existence of a RG tools doesn’t solve the problem. The missing piece in the presentation—and currently in the wider AI safety conversation—is the measurement problem, something we have had to reckon with for a long time.
How do you accurately identify when a user is “at-risk” based solely on behavioral signals?
This is the exact challenge the gambling industry has been struggling with for years. While we have begun to identify important behavioral markers like chasing losses, spike play, and frequent depositing, we also acknowledge the limitations of these markers in providing a truly reliable signal.
Unlike other addictive disorders, gambling is uniquely linked to digital environments and real-time data tracking. This presents a unique opportunity for collaboration. There is clear potential for transfer learning here: we can share our frameworks for identifying harm and also, perhaps more importantly, the lessons we have learned. If Generative AI is indeed heading toward a “Digital Heroin” moment, the gambling industry’s experience with harm minimization might just be the best playbook they have.
In closing…
I’m already looking forward to next year’s conference. If this week proved anything, it’s that NeurIPS is where the future is forged, and I’ll definitely be back in 2026 to see what’s next for AI and the gambling industry.