Brass pins on felt, one pair bound with red thread
Write-up

Finding colluders who share nothing.

Eric Keller  ·  October 2026  ·  about 9 minutes

Most collusion tooling in online poker links accounts by what they share: a device, an IP range, a payment method. That catches one person running several accounts. It does very little against two real people on two phones who have agreed to go easy on each other or to move chips between them. This is a write-up of how far you can get by looking only at how the cards were played.

The benchmark

Slash ran "Detect Suspicious Value Transfers in Poker" on Kaggle, closing in September 2026. 370+ teams got two million hands from 400 tables of 30 players, with every hole card visible. Each table had a development phase and an evaluation phase. The labels were thin on purpose: 372 colluding pairs, 1,488 pairs confirmed clean, and about 156,000 pairs nobody had looked at.

Three behaviours were described. Directed transfer is chip dumping. Soft play is two players avoiding each other's money. Coordinated isolation is two players squeezing a third out of pots. A fourth type existed only in the evaluation data and was never described.

The score had three parts:

Where we finished, and what happened after

At the deadline we were 5th on the final board with 0.92817. Late submissions on Kaggle still return a private score, which is useful and also dangerous: with enough of them you can fit the hidden test set and learn nothing. So we set rules first. Every change was chosen on held-out development data, its expected score was written down before submission, and we capped ourselves at twelve late submissions in total.

VersionWhat changedPrivate
DeadlineGradient-boosted pair models, a listing rule for evidence0.92817
v2Opponent-aware residuals, plus a pair-level sequence model0.93449
Winning entry0.93514
v3Player-specific baselines inside the sequence model0.93562

A few of the late submissions were spent pulling scores apart into their three components. That showed where the gap actually was. On pair ranking, v2 scored 0.9885 against the winner's 0.9858. On evidence (0.7218 vs 0.7296) and type (0.9818 vs 0.9919), the winner was ahead. These are late, unofficial scores, and the official standings don't change.

What worked

1. Asking how A plays while B is still in the hand

Start with a model of normal play: given the position, the price, the stacks, the action so far and the player's own cards, how likely is a fold, a call or a raise? Train it cross-fitted by table, so no player is ever scored by a model that saw them.

Then, for every two players A and B, take every decision A made while B was still live. Compare what A did with what the model expected, using a likelihood-ratio test and per-action z-scores, overall and when A was facing B's bet. A soft-play partner folds to one opponent far more than the spot justifies. A chip dumper calls and pays off in spots where nobody calls.

This idea came from another team's write-up (8th place). It was the largest single gain we found: one gradient-boosted model went from 0.958 to 0.969 average precision on our held-out view. The soft-play family, the hardest one, gained the most.

A ledger of handwritten tallies
Every decision gets an expected distribution. The signal is in the gap between expected and actual, per opponent.

2. A sequence model over the hands two players shared

One example is one pair; one token is one hand they were both dealt, in order. Each token carries:

That last part lets the model explain a hand away. If A's strange folds line up with his real partner being at the table, his innocent tablemates stop inheriting his signal. Bolting that on afterwards as a demotion rule made things worse every time we tried it. Inside the model it worked.

Three things mattered more than architecture:

3. Player-specific baselines, in the right place

We built a second policy model that also knows each player's own tendencies (how often they fold to a bet, raise when checked to, and so on), computed without the current hand. Its log loss on held-out actions is 17% lower than the population model's.

Fed to the gradient-boosted pair models, it did nothing in three different forms, and twice made things slightly worse. Those models already rebuild each player's baseline from the pair features. Fed to the sequence model, which has no such features, it was the step that moved us past the winning score.

A feature that is strong on its own tells you very little about what it adds to a model that already sees most of the same information. We had to learn this several times.

What didn't work

An open case folder with circled hands
The evidence part of the score rewards listing the right five hands in the right order. It is still where the most points are left.

How we kept it honest

Most of the hidden positives in this data are unlabelled pairs that look exactly like the labelled ones. A validation set that quietly removes them, which was our first one, rewards a model for ranking them low. That is exactly the wrong behaviour.

We kept two views. One drops the suspected hidden positives; the other counts them as positives. We shipped only what won on both. That rule killed several changes that looked good on the first view alone. One of them scored a perfect 1.000 probability of improvement on the lenient view and then barely moved the leaderboard.

What changes on real data

Some things get easier. An operator holds every hole card, so the "what did each player actually have" features we relied on are available in production, not just in a benchmark.

Some things get harder:

If you run games and want to see what this finds on your own hands, there's a short pilot on the main page.