W poprzednim wpisie ograłeś sześciu przeciwników w kamień, papier, nożyce. Nauczyłeś się tam trzech rzeczy: czym jest exploit, że każdy exploit ma swój exploit i że istnieje strategia, której nie da się pokonać.
Ale kamień, papier, nożyce to nie poker. Czegoś tam brakowało.
Dwóch rzeczy. Zakrytych kart i zakładów. W RPS obaj wiemy wszystko i płacimy tyle samo. W pokerze nie widzisz mojej ręki, a stawką sterujemy betami. I dopiero z tych dwóch składników rodzi się coś, czego w RPS nie było wcale: blef.
Dziś dam ci całego pokera w wersji mini. Sześć kart, dwie rundy, trzy decyzje. I trzech przeciwników do ogrania.
Czym jest Leduc Hold'em? To uproszczony poker badawczy: talia 6 kart (J, Q, K w dwóch kolorach), jedna karta własna, jedna wspólna, dwie rundy zakładów. Zachowuje blef, value betting i pozycję pełnego pokera, ale jest na tyle mały, że da się go policzyć do końca. Standard w badaniach nad AI (RLCard, PettingZoo).
Poker skurczony do sześciu kart
Ta gra nazywa się Leduc Hold'em. Nie wymyślili jej pokerzyści, tylko badacze AI: potrzebowali pokera na tyle małego, żeby dało się go policzyć do końca, i na tyle prawdziwego, żeby zachował blef, value bet i pozycję. Do dziś jest standardowym poligonem w badaniach nad game AI (środowiska RLCard i PettingZoo).
Zasady:
- Talia: 6 kart. Walet, Dama, Król, każda ranga w dwóch kolorach.
- Obaj wrzucamy ante 1 chip. Dostajesz jedną zakrytą kartę. Przeciwnik też.
- Runda 1: bet albo raise kosztuje 2 chipy. Potem odsłania się jedna karta wspólna.
- Runda 2: bet albo raise kosztuje 4 chipy. Max dwa raise'y na rundę.
- Showdown: para z kartą wspólną bije wszystko. Bez pary wygrywa wyższa karta.
Tyle. Całe zasady. A jednak siedzi w nich zaskakująco dużo pokera: masz rękę, której przeciwnik nie widzi, masz turn zmieniający układ sił i masz decyzje o tym, ile pieniędzy wchodzi do puli. Czyli masz wszystko, czego potrzebuje blef.
Gra 1: przeciwnik, który nigdy nie blefuje
Pierwszy bot jest szczery do bólu. Betuje i raisuje tylko wtedy, kiedy ma Króla albo parę. Z resztą czeka i modli się o darmowy showdown.
Zagraj z nim kilkanaście rozdań, zanim czytasz dalej. Spróbuj wyczuć, co robić z Waletem.
Zauważyłeś? Jego zagrania mówią ci wszystko. Bet znaczy siłę, więc twoje Q i J mogą uciekać za darmo. A jego check znaczy słabość, więc... no właśnie.
Właśnie odkryłeś blef. Kiedy on czeka, betujesz KAŻDĄ kartą. Także Waletem, najgorszą kartą w talii. I to zarabia.
Blef zarabia na jego foldzie, nie na twojej karcie.
Walet w twojej ręce nie zrobił się lepszy. Zarabia, bo przeciwnik wyrzuca wszystko poniżej Króla. Płaci ci nie twoja karta, tylko jego błąd: składa za często. W RPS exploitem był papier na czyjś kamień. Tu exploitem jest bet na czyjś strach.
Gra 2: przeciwnik, który nigdy nie pasuje
Drugi bot to odwrotność pierwszego. Raisuje kiedy tylko może i nie spasował w życiu ani razu.
Pierwsza pokusa: skoro tamtego dało się blefować, to tego też. Sprawdź. Serio, spróbuj go zablefować Waletem.
Boli, co?
Blef zarabia tylko wtedy, gdy przeciwnik umie spasować. Ten nie umie, więc każdy twój blef to podpalanie chipów. Kontra na niego jest dokładnie odwrotna: czysty value bet. Z Królem i parami raisuj ile się da, on zapłaci każdym śmieciem. Z Waletem wyrzucaj i nie oglądaj się za siebie.
Postaw te dwa boty obok siebie. Ten sam Walet: przeciwko pierwszemu betujesz nim jak szalony, przeciwko drugiemu wyrzucasz go bez mrugnięcia. Ta sama karta, dwie przeciwne decyzje. Bo exploit nie celuje w kartę. Celuje w konkretny błąd konkretnego przeciwnika.
Trzy karty to cały zakres
W Leducu przeciwnik ma jedną z trzech rang. Zawsze. Możesz w głowie ogarnąć jego CAŁY możliwy zakres: Walet, Dama albo Król, i tyle.
I to jest dokładnie to myślenie, które w prawdziwym pokerze nazywa się graniem przeciwko zakresowi, a nie przeciwko "ręce, którą mu przypisałem". W No Limit Hold'em przeciwnik może mieć 1326 kombinacji kart startowych. Nikt nie widzi ich wszystkich naraz. Ale mechanizm jest ten sam co przy trzech: jego akcje odcinają kawałki zakresu, a ty grasz przeciwko temu, co zostało.
Leduc pozwala ci poczuć ten mechanizm w skali, w której mieści się on w pamięci operacyjnej. Bet drugiego gracza w rundzie 2, kiedy na stole leży Dama? Odetnij z jego zakresu to, czym by nie betował. To, co zostało, mówi ci, czy twój Król to jeszcze value, czy już tylko bluff-catcher.
Gra 3: równowaga, której nie wymyśliłem
Została ostatnia obietnica z tamtego wpisu: strategia, której nie da się pokonać. W RPS była banalna, równy mix 1/3. W Leducu nikt jej ręcznie nie wypisze: jest mieszana, zależy od karty, rundy i całej historii betów.
Więc jej nie wymyśliłem. Wytrenowałem ją.
Bot poniżej gra strategią z algorytmu CFR (Counterfactual Regret Minimization): program grał sam ze sobą setki tysięcy rozdań i po każdym korygował swoje częstotliwości tam, gdzie żałował decyzji. To ta sama rodzina metod, którą Uniwersytet Alberty rozwiązał heads-up limit hold'em (bot Cepheus, 2015). Solvery, których używasz do nauki, to kuzyni tego podejścia.
Uczciwość wymaga liczb, więc je zmierzyłem. Po treningu szukałem najlepszej możliwej kontry na tego bota: grając jako pierwszy, najlepsza znaleziona kontra traci ok. 0,08 chipa na rozdanie. Grając jako drugi, zarabia ok. 0,08. Suma wychodzi praktycznie zero, czyli bot siedzi w równowadze, a różnica między tymi liczbami to nie jego słabość.
To cena pozycji.
Pozycja to nie frazes z kursów. W tej grze zaczynający oddaje ok. 0,08 chipa na rozdanie nawet grając perfekcyjnie. Zmierzone, nie wymyślone.
Zagraj z nim. Blefuj, nie blefuj, graj tight, graj agro. Patrz na wykres i poczuj różnicę: z dwoma poprzednimi botami linia szła w górę, jak tylko znalazłeś ich błąd. Tu nie ma czego znaleźć.
Podpowiedź w grze pokazuje metrykę treningu tego bota.
Tak wygląda GTO w Leducu
Grałeś z równowagą w ciemno. Czas ją zobaczyć. Poniżej masz chart jak z pokerowego trenera: realne frekwencje bota, z którym właśnie grałeś. Złoto to bet albo raise, niebieski to check albo call, szary to fold. Długość paska to częstotliwość.
Zanim zaczniesz scrollować, cztery rzeczy, których szukaj:
- Walet otwiera betem 9% czasu. Najgorszą kartą w talii. To nie kaprys, to wbudowana kwota blefu: na tyle mała, że nie da się jej ukarać, na tyle realna, że nie możesz ignorować jego betów.
- Drabinka po twoim checku: 28%, 85%, 100%. Tak często betuje kolejno z J, Q i K. Zauważ: z każdą kartą betuje CZASEM. Nie ma karty, która gra tylko jedną akcję w każdej sytuacji.
- Para z Waletem na boardzie: check 100%. Najsilniejsza możliwa ręka w tym spocie, a równowaga czeka. Slow-play pod check-raise, policzony, nie wyczuty. Tego nie wymyślisz przy stole.
- OOP vs IP na samym dole. Ta sama karta, ten sam moment gry, a w pozycji betujesz częściej: Walet 9% vs 28%, Król 77% vs 100%. Pod parami pasków zmierzona cena samej kolejności działania: ok. 0,16 chipa na rozdanie.
To jest sedno grania mieszanego, o które wszyscy pytają przy solverach. Frekwencja przy akcji nie znaczy "solver nie umiał się zdecydować". Znaczy: ta karta MUSI czasem robić jedno, a czasem drugie, w konkretnej proporcji, bo inaczej twój przeciwnik dostaje darmowy exploit.
Mieszanie frekwencji to nie brak decyzji. To jest decyzja.
Jedno zastrzeżenie, bo uczciwość: to jest równowaga wytrenowana lokalnie na moje potrzeby i jedna z możliwych. Gry tego typu mają wiele równowag; inny solver może wypluć inne liczby przy tym samym zerowym exploitability. Kierunki i mechanizmy zostają te same.
Czym różni się prawdziwy poker?
Skalą, nie naturą. 52 karty zamiast 6. Cztery ulice zamiast dwóch. Dowolne sizingi zamiast sztywnych 2 i 4. Drzewo decyzji rośnie do rozmiarów, w których nawet solver liczy przybliżenia, a my uczymy się przybliżeń tych przybliżeń.
Ale siły w grze pozostają te same trzy, które właśnie przećwiczyłeś. Blef, który zarabia na cudzym foldzie. Value, które zarabia na cudzym callu. I równowaga, która pilnuje proporcji między nimi, żebyś to ty exploitował, a nie ciebie.
Jak ktoś ci następnym razem powie, że "GTO to granie bez wyobraźni", pokaż mu bota numer jeden i spytaj, czemu przegrywa z Waletem. A jak powie, że "blef to sztuka czytania duszy", pokaż mu bota numer dwa.
Cały poker w sześciu kartach. Wszystkie trzy gry w jednym miejscu: pawel.life/tools/leduc.
Podobało się? To jest część serii o GTO bez bullshitu: kamień, papier, nożyce → Kuhn (3 karty) → Leduc. Daj znać, czy robić trzeci, przyciskiem na dole albo pisząc do mnie.
In the previous post you beat six opponents at rock, paper, scissors. You learned three things there: what an exploit is, that every exploit has its own exploit, and that there exists a strategy that cannot be beaten.
But rock, paper, scissors is not poker. Something was missing.
Two things. Hidden cards and bets. In RPS we both know everything and pay the same. In poker you can't see my hand, and we control the stakes with bets. And only from those two ingredients comes something RPS never had: the bluff.
Today I'm giving you all of poker in a mini version. Six cards, two rounds, three decisions. And three opponents to beat.
What is Leduc Hold'em? A simplified research poker: a 6-card deck (J, Q, K in two suits), one private card, one community card, two betting rounds. It keeps the bluff, value betting and position of full poker while staying small enough to compute completely. A standard in AI research (RLCard, PettingZoo).
Poker shrunk to six cards
The game is called Leduc Hold'em. It wasn't invented by poker players but by AI researchers: they needed a poker small enough to compute completely, yet real enough to keep the bluff, the value bet and position. To this day it's a standard testbed in game AI research (the RLCard and PettingZoo environments).
The rules:
- Deck: 6 cards. Jack, Queen, King, each rank in two suits.
- We both ante 1 chip. You get one private card. So does the opponent.
- Round 1: a bet or raise costs 2 chips. Then one community card is revealed.
- Round 2: a bet or raise costs 4 chips. Max two raises per round.
- Showdown: a pair with the community card beats everything. No pair, higher card wins.
That's it. The whole rulebook. And yet there's a surprising amount of poker in there: you hold a hand your opponent can't see, there's a turn card that shifts the balance, and there are decisions about how much money goes in. Which means there's everything a bluff needs.
Game 1: the opponent who never bluffs
The first bot is honest to a fault. He bets and raises only with a King or a pair. With everything else he checks and prays for a free showdown.
Play a dozen hands before you read on. Try to work out what to do with a Jack.
Did you notice? His actions tell you everything. A bet means strength, so your Q and J can run away for free. And his check means weakness, so... exactly.
You just discovered the bluff. When he checks, you bet EVERY card. Including the Jack, the worst card in the deck. And it makes money.
A bluff earns on his fold, not on your card.
The Jack in your hand didn't get any better. It earns because the opponent folds everything below a King. What pays you is not your card but his mistake: he folds too often. In RPS the exploit was paper against someone's rock. Here the exploit is a bet against someone's fear.
Game 2: the opponent who never folds
The second bot is the first one's opposite. He raises whenever he can and has never folded in his life.
The first temptation: if the previous one could be bluffed, so can this one. Check for yourself. Seriously, try bluffing him with a Jack.
Hurts, doesn't it?
A bluff only makes money when the opponent is capable of folding. This one isn't, so every bluff of yours is burning chips. The counter is exactly the opposite: pure value betting. Raise your Kings and pairs as much as the rules allow, he'll pay you off with any trash. Fold your Jacks and don't look back.
Put the two bots side by side. The same Jack: against the first one you bet it relentlessly, against the second you fold it without blinking. Same card, two opposite decisions. Because an exploit doesn't target a card. It targets a specific mistake of a specific opponent.
Three cards are a whole range
In Leduc your opponent holds one of three ranks. Always. You can hold his ENTIRE possible range in your head: Jack, Queen or King, and that's it.
And that is exactly the kind of thinking real poker calls playing against a range, not against "the hand I decided he has". In No Limit Hold'em an opponent can have 1,326 starting combos. Nobody sees them all at once. But the mechanism is the same as with three: his actions cut pieces off his range, and you play against what's left.
Leduc lets you feel that mechanism at a scale that actually fits in working memory. The opponent bets in round 2 with a Queen on board? Cut from his range whatever he wouldn't bet. What remains tells you whether your King is still value or already just a bluff-catcher.
Game 3: an equilibrium I didn't make up
One promise from the previous post remains: the strategy that cannot be beaten. In RPS it was trivial, an even 1/3 mix. In Leduc nobody writes it out by hand: it's mixed, and it depends on your card, the round and the whole betting history.
So I didn't make it up. I trained it.
The bot below plays a strategy produced by CFR (Counterfactual Regret Minimization): the program played hundreds of thousands of hands against itself, and after each one it corrected its frequencies wherever it regretted a decision. It's the same family of methods the University of Alberta used to solve heads-up limit hold'em (the Cepheus bot, 2015). The solvers you study with are cousins of this approach.
Honesty requires numbers, so I measured them. After training I searched for the best possible counter to this bot: playing first, the best counter I found loses about 0.08 chips per hand. Playing second, it earns about 0.08. The sum is practically zero, which means the bot sits at equilibrium, and the gap between those numbers is not its weakness.
It's the price of position.
Position is not a coaching cliché. In this game the first player gives up about 0.08 chips per hand even playing perfectly. Measured, not invented.
Play him. Bluff, don't bluff, play tight, play aggro. Watch the chart and feel the difference: against the two previous bots the line climbed as soon as you found their mistake. Here there is nothing to find.
The in-game hint shows this bot's training metric.
This is what GTO looks like in Leduc
You've been playing against the equilibrium blind. Time to see it. Below is a chart straight out of a poker trainer: the real frequencies of the bot you just played. Gold is bet or raise, blue is check or call, gray is fold. Bar length is frequency.
Before you start scrolling, four things to look for:
- The Jack opens with a bet 9% of the time. The worst card in the deck. Not a whim. A built-in bluff quota: small enough that it can't be punished, real enough that you can't ignore his bets.
- The ladder after your check: 28%, 85%, 100%. That's how often he bets with J, Q and K respectively. Notice: every card bets SOMETIMES. No card plays just one action in every spot.
- A pair with the Jack on board: check 100%. The strongest possible hand in that spot, and equilibrium checks it. A slow-play set up for a check-raise. Computed, not felt. You would not invent this at the table.
- OOP vs IP at the very bottom. Same card, same moment of the game, yet in position you bet more: the Jack 9% vs 28%, the King 77% vs 100%. Below the bar pairs sits the measured price of acting order alone: about 0.16 chips per hand.
This is the heart of mixed play, the thing everyone asks about with solvers. A frequency next to an action doesn't mean "the solver couldn't make up its mind". It means: this card MUST sometimes do one thing and sometimes the other, in a specific proportion, because otherwise your opponent gets a free exploit.
Mixing frequencies is not indecision. It is the decision.
One caveat, for honesty's sake: this is an equilibrium trained locally for my purposes, and one of many. Games like this have multiple equilibria; a different solver can spit out different numbers with the same zero exploitability. The directions and mechanisms stay the same.
How is real poker different?
In scale, not in nature. 52 cards instead of 6. Four streets instead of two. Any bet size instead of a fixed 2 and 4. The decision tree grows to a size where even solvers compute approximations, and we learn approximations of those approximations.
But the forces in the game remain the same three you just practiced. The bluff, which earns on someone's fold. Value, which earns on someone's call. And equilibrium, which keeps the proportions between them so that you are the one exploiting, not the one being exploited.
Next time someone tells you "GTO is playing without imagination", show them bot number one and ask why it loses to a Jack. And when they say "bluffing is the art of reading souls", show them bot number two.
All of poker in six cards. All three games in one place: pawel.life/tools/leduc.
Enjoyed this? It's part of the no-bullshit GTO series: rock, paper, scissors → Kuhn (3 cards) → Leduc. Tell me whether to make a third one, with the button below or by writing to me.