Most collusion tooling in online poker links accounts by what they share: a device, an IP range, a payment method. That catches one person running several accounts. It does very little against two real people on two phones who have agreed to go easy on each other or to move chips between them. This is a write-up of how far you can get by looking only at how the cards were played.
The benchmark
Slash ran "Detect Suspicious Value Transfers in Poker" on Kaggle, closing in September 2026. 370+ teams got two million hands from 400 tables of 30 players, with every hole card visible. Each table had a development phase and an evaluation phase. The labels were thin on purpose: 372 colluding pairs, 1,488 pairs confirmed clean, and about 156,000 pairs nobody had looked at.
Three behaviours were described. Directed transfer is chip dumping. Soft play is two players avoiding each other's money. Coordinated isolation is two players squeezing a third out of pots. A fourth type existed only in the evaluation data and was never described.
The score had three parts:
- 70% for ranking 112,540 pairs (average precision);
- 20% for listing up to five hands that prove each real case, in the order the host listed them;
- 10% for naming the type.
Where we finished, and what happened after
At the deadline we were 5th on the final board with 0.92817. Late submissions on Kaggle still return a private score, which is useful and also dangerous: with enough of them you can fit the hidden test set and learn nothing. So we set rules first. Every change was chosen on held-out development data, its expected score was written down before submission, and we capped ourselves at twelve late submissions in total.
| Version | What changed | Private |
|---|---|---|
| Deadline | Gradient-boosted pair models, a listing rule for evidence | 0.92817 |
| v2 | Opponent-aware residuals, plus a pair-level sequence model | 0.93449 |
| Winning entry | 0.93514 | |
| v3 | Player-specific baselines inside the sequence model | 0.93562 |
A few of the late submissions were spent pulling scores apart into their three components. That showed where the gap actually was. On pair ranking, v2 scored 0.9885 against the winner's 0.9858. On evidence (0.7218 vs 0.7296) and type (0.9818 vs 0.9919), the winner was ahead. These are late, unofficial scores, and the official standings don't change.
What worked
1. Asking how A plays while B is still in the hand
Start with a model of normal play: given the position, the price, the stacks, the action so far and the player's own cards, how likely is a fold, a call or a raise? Train it cross-fitted by table, so no player is ever scored by a model that saw them.
Then, for every two players A and B, take every decision A made while B was still live. Compare what A did with what the model expected, using a likelihood-ratio test and per-action z-scores, overall and when A was facing B's bet. A soft-play partner folds to one opponent far more than the spot justifies. A chip dumper calls and pays off in spots where nobody calls.
This idea came from another team's write-up (8th place). It was the largest single gain we found: one gradient-boosted model went from 0.958 to 0.969 average precision on our held-out view. The soft-play family, the hardest one, gained the most.
2. A sequence model over the hands two players shared
One example is one pair; one token is one hand they were both dealt, in order. Each token carries:
- the hand's action and chip-flow features;
- the policy residuals for each direction;
- "graph context": how each player did against their strongest other partner in that same hand.
That last part lets the model explain a hand away. If A's strange folds line up with his real partner being at the table, his innocent tablemates stop inheriting his signal. Bolting that on afterwards as a demotion rule made things worse every time we tried it. Inside the model it worked.
Three things mattered more than architecture:
- Train on the population that gets scored. The host only scored pairs with 38 or more shared hands. A quarter of our first training set sat below that. Fixing it was worth more than any layer we added.
- Bag seeds. Single runs ranged from 0.919 to 0.950 on the same data. Three seeds averaged gave 0.955, and gains stopped at about six.
- Fuse in log-odds, at the model level. The sequence model's scores are very peaked, so mixing its ranks into the final ranking only added noise. Adding its log-odds to each gradient-boosted model before the final blend helped on every view we checked.
3. Player-specific baselines, in the right place
We built a second policy model that also knows each player's own tendencies (how often they fold to a bet, raise when checked to, and so on), computed without the current hand. Its log loss on held-out actions is 17% lower than the population model's.
Fed to the gradient-boosted pair models, it did nothing in three different forms, and twice made things slightly worse. Those models already rebuild each player's baseline from the pair features. Fed to the sequence model, which has no such features, it was the step that moved us past the winning score.
A feature that is strong on its own tells you very little about what it adds to a model that already sees most of the same information. We had to learn this several times.
What didn't work
- A tabular foundation model as an extra member. Two of the top three teams used one. For us it helped one validation view and hurt the stricter one at every weight.
- "Partner present" contrasts on their own. Strong univariately (0.996 AUC), worth nothing on top of the model.
- Hard negatives (a known colluder paired with someone he didn't collude with) in the sequence model. Ranking got worse: 0.894 to 0.862.
- An exact expected-AP decoder for the evidence lists. The host's listing rule is nested, which makes plain sorting already optimal. It gained 0.00004.
- Cheap detectors for the undescribed fourth type. The type is real, and labelling it correctly is worth about 0.002. Neither of our cheap signals picked out the right pairs.
How we kept it honest
Most of the hidden positives in this data are unlabelled pairs that look exactly like the labelled ones. A validation set that quietly removes them, which was our first one, rewards a model for ranking them low. That is exactly the wrong behaviour.
We kept two views. One drops the suspected hidden positives; the other counts them as positives. We shipped only what won on both. That rule killed several changes that looked good on the first view alone. One of them scored a perfect 1.000 probability of improvement on the lenient view and then barely moved the leaderboard.
What changes on real data
Some things get easier. An operator holds every hole card, so the "what did each player actually have" features we relied on are available in production, not just in a benchmark.
Some things get harder:
- Real cheaters adapt.
- Friends-only and club games have fewer "everyone else" to compare against, so groups of regulars can look like rings. Any real deployment needs a false-positive check on exactly those groups.
- New attacks need a look of their own, such as the hole-card leaks through compromised table tools that surfaced this month.
- Calibration needs real cases. The model has to be calibrated against cases a security team has already closed. Synthetic labels can't do that.
If you run games and want to see what this finds on your own hands, there's a short pilot on the main page.
