GTO. Game Theory Optimal. These days it’s hard to hear the word poker without hearing the word GTO: "Play GTO.", "GTO has the answer", "Solvers play GTO", "GTO can’t be beaten" etc.
But ask someone to explain what it actually is, and somehow it gets harder to hear anything sensible. Maybe someone repeats a memorized rule and that’s it.
Today I’ll show you what it really is, using a simple schoolyard game: rock, paper, scissors. What does it have to do with poker? Quite a lot, actually.
Let’s go.
Where did all this GTO come from
Game theory is a branch of mathematics about, among other things, decisions whose outcome depends not only on you, but also on what others do. John Nash (btw, I recommend the movie "A Beautiful Mind") proved in 1950 that in such games there is an equilibrium point: a set of strategies where nobody gains by changing anything on their own. He got a Nobel Prize for it. GTO in poker is exactly that equilibrium, carried over into the world of cards.
A solver is a program that computes this equilibrium for a specific hand. It gets ranges, stacks and sizings, then plays against itself millions of times, improving both strategies in turns, until neither side can gain anything anymore. Solvers entered poker a little over a decade ago. But one important thing: a solver does not understand poker. It doesn’t know what a bluff is or what value is. It just iterates to equilibrium. Understanding WHY the output looks the way it does is our job, not its.
And one caveat: a solver computes an approximation of equilibrium under the assumptions you feed it. Change the ranges, the sizings or the distance to equilibrium, and the answer changes. A solver is not an oracle. It’s a calculator. It solves the equation you give it. Many people go there looking for the answer to "Did I play this well?". But that answer will always be limited by the inputs you provide. You have a lot of influence over the answer you get. And how are you supposed to stay objective here?
What "optimal" means
Game Theory Optimal is a strategy that mathematically cannot be beaten in the long run. Whatever your opponent comes up with, he won’t gain an edge over you. In the worst case you break even, and if he makes any mistake at all, you profit.
There is one very important assumption in this definition, often skipped. When looking for a GTO strategy we assume that the opponent knows our strategy. All of it. He sees it laid out in the open and can adjust however he wants. And his best possible result against our strategy is breaking even. That’s exactly how solvers work: they play one against the other until they reach the point where neither side can gain an edge.
The measure of success: EV
Before we start playing, let’s agree on how we measure success. A won round is +1, a tie 0, a lost round −1. EV, expected value, is the average of these outcomes weighted by what your opponent plays and how often. It tells you how much you earn if you repeat a play many, many times.
Example. The opponent plays nothing but rock. What do you do to squeeze the most out of his tendency? Answer that question for yourself.
Show answer
You play paper. You win 100% of the time, so:
EV(paper) = 1·(+1) + 0·0 + 0·(−1) = +1
That’s the theory. In practice you’ll see EV in every game below. You can also find all the variants in one place in the tool pawel.life/tools/rps-gto.
Rock, paper, scissors
You probably know the rules. Rock beats scissors, scissors beat paper, paper beats rock.
Let’s start with strategies that are NOT GTO. I’m telling you straight up: I throw rock 100% of the time. What do you do? You play paper. Every round. EV(paper) = +1.00, you win everything. My strategy is exploitable to the bone.
Click Hint in the game if you want a ready-made counter.
This time the opponent improves. He still loves rock but throws it less often: rock 80%, paper 10%, scissors 10%. What is your maximum exploit? Answer before you click.
Show answer
You change nothing, you keep playing paper only. Let’s count:
EV(paper) = 0.8·(+1) + 0.1·0 + 0.1·(−1) = +0.70 per round
The leak is smaller, but you’re still crushing him. Notice on the chart that the line jumps around more than in the previous game. That’s variance: sometimes the opponent hits scissors into your paper and you lose individual rounds, even though the play is profitable.
The opponent sees it and changes his strategy once more, this time: rock 50%, paper 30%, scissors 20%. What do you do?
Show answer
Your strategy stays the same. The max exploit is still 100% paper, but you earn a lot less.
EV(paper) = 0.5·(+1) + 0.3·0 + 0.2·(−1) = +0.30 per round
You see the pattern. As long as the opponent’s mix is unbalanced, you can always find a counter-strategy with positive EV against it. Notice that variance grows again as the edge gets smaller.
How much to extract from your opponent’s mistakes?
Let’s say I know your leak: you like rock and subconsciously throw it 40% of the time, the rest split evenly (30%). The theoretical max for me is playing paper only:
EV(paper) = 0.4·(+1) + 0.3·0 + 0.3·(−1) = +0.10 per round
This is max exploit. I’m squeezing every possible point of EV out of your mistake. But +0.10 is a small edge: on the chart you can watch variance hide it for the first dozen or so rounds. Small edges are realized over the long run, not within a single session.
And the second catch. I play paper only, so after a few rounds you see what’s going on. And you adjust.
Same thing at the table. You noticed someone folds too often to 3-bets, so you 3-bet him wider. But if you start doing it every hand, there’s a good chance he catches on. And you lose your future shots at that exploit. Finding an exploit is an art. Using it so it earns as much as possible is an art too.
The one strategy you cannot beat
Rock, paper and scissors in equal parts, each 1/3 of the time, chosen at random.
How much do you earn here?
Show answer
EV(paper) = 1/3·(+1) + 1/3·0 + 1/3·(−1) = 0.00
Zero. Scissors only? Same zero. You mix evenly like the opponent? Zero again. There is no adjustment that wins anything here.
This is what unexploitable means. No knowledge about the opponent’s mix gives you anything. And it works both ways: any mix other than 1/3, 1/3, 1/3 can be exploited and beaten in this game. Play for a while and watch the chart: whatever you come up with, the line will keep circling around zero.
Your exploit has an exploit too
Let’s go back to the beginning for a moment. The opponent played rock only, you played paper only and took the whole pot. Was YOUR strategy optimal then?
It was, but only against that one specific opponent strategy.
It beat that opponent, but it was just as exploitable. All it took was the opponent catching on and switching to scissors only, and the roles flip. Because every exploit has its own exploit. The counter to rock is paper, the counter to your paper is scissors, and the counter to scissors is rock again. The circle closes. Easy to see here, harder in poker.
Remember this simple mechanism, because at the poker table it happens nonstop: by countering someone’s mistake, you expose yourself to an exploit.
Time for the second-to-last exercise. The opponent below has no fixed mix. He counts your frequencies over the last 10 rounds and counters your most frequent move. Knowing how he plays, can you stay one step ahead of him and win?
The Hint in the game shows how to counter his counter.
How is poker different from this simple game?
In rock, paper, scissors the GTO strategy has one brutal flaw: it does not win. The EV of even mixing against any strategy is zero. It guarantees you a draw with every opponent in the world. And nothing more.
Poker is different. Poker is far more complex.
The GTO strategy in poker wins. Not because it’s clever. Because the people on the other side make mistakes, and their deviations from equilibrium hand you their EV. And there will be plenty of those mistakes, because perfect poker is simply impossible for a human. Too many possible scenarios, too little working memory.
And here we come back to the title. If nobody can play perfectly, then nobody plays GTO. I don’t. You don’t. Solvers compute approximations, and we learn approximations of those approximations. And on those approximations we learn exploits. Because there is no exploit without GTO.
GTO is a reference point. The baseline you have to understand to know when and why you deviate from it.
Did you enjoy this style of article? GTO made simple, plus interactive examples? Let me know by writing to me, or at least with the button at the bottom.
GTO. Game Theory Optimal. Ciężko dzisiaj usłyszeć słowo poker, nie słysząc słowa GTO: "Graj GTO.", "GTO ma odpowiedź", "Solver gra GTO", "GTO nie da się pokonać" etc.
Ale jak poprosisz kogoś o wyjaśnienie co to naprawdę jest, to jakoś trudniej coś sensownego usłyszeć. Może ktoś powtórzy zapamiętaną regułę i tyle.
Pokażę Wam dzisiaj, czym to naprawdę jest, na podstawie prostej, podwórkowej gry: papier, kamień, nożyce. Co ona ma wspólnego z pokerem? Ano całkiem sporo.
Zaczynamy.
Skąd w ogóle to całe GTO
Teoria gier to dział matematyki m.in. o decyzjach, których wynik zależy nie tylko od ciebie, ale też od tego, co zrobią inni. John Nash (btw polecam film "Piękny umysł") udowodnił w 1950 roku, że w takich grach istnieje punkt równowagi: zestaw strategii, przy którym nikomu nie opłaca się nic zmieniać na własną rękę. Dostał za to Nobla. GTO w pokerze to właśnie ta równowaga, przeniesiona do świata kart.
Solver to program, który tę równowagę liczy dla konkretnego rozdania. Dostaje range, stacki i sizingi, a potem gra sam ze sobą miliony razy, poprawiając na zmianę obie strategie, aż żadna strona nie może już nic ugrać. Solvery weszły do pokera nieco ponad dekadę temu. Ale ważna rzecz: solver nie rozumie pokera. Nie wie, co to blef, a co value. Po prostu iteruje do równowagi. Rozumienie, DLACZEGO wynik wygląda tak, a nie inaczej, to nasza robota, nie jego.
I jedno zastrzeżenie: solver liczy przybliżenie równowagi przy założeniach, które sam mu podajesz. Zmienisz range'y albo sizingi lub dystans do równowagi - zmieni się odpowiedź. Solver to nie wyrocznia. To kalkulator. Liczy zadane mu równanie. Wiele osób szuka tam odpowiedzi na pytanie "Czy zagrałem dobrze?". Ale odpowiedź na nie będzie zawsze ograniczona do danych wejściowych, które mu podasz. Masz duży wpływ na odpowiedź, którą dostajesz. I jak tu być obiektywnym?
Co znaczy "optymalnie"
Game Theory Optimal to strategia, której matematycznie nie da się pokonać w długim terminie. Cokolwiek wymyśli przeciwnik, nie zyska nad tobą przewagi. W najgorszym scenariuszu wychodzisz na zero, a jeśli on popełni jakikolwiek błąd - zarabiasz.
Jest w tej definicji bardzo ważne założenie, często pomijane. Szukając strategii GTO zakładamy, że przeciwnik zna naszą strategię. Całą. Widzi ją jak na dłoni i może się dowolnie dostosować. A jego najlepszym wynikiem kontra nasza strategia, jest gra na zero. Tak właśnie działają solvery: grają sobie jeden przeciw drugiemu, aż dojdą do momentu, w którym nikt nie jest w stanie zyskać przewagi.
Miara sukcesu: EV
Zanim zaczniemy grać w grę, umówmy się na sposób pomiaru sukcesu. Wygrana runda to +1, remis 0, przegrana −1. EV, czyli wartość oczekiwana, to średnia z tych wyników ważona tym, co i jak często gra przeciwnik. Mówi ci, ile zarabiasz, jeśli powtórzysz to zagranie wiele wiele razy.
Przykład. Przeciwnik gra sam kamień. Co ty robisz, żeby maksymalnie wykorzystać tendencję przeciwnika? Odpowiedz sobie na to pytanie.
Pokaż odpowiedź
Ty grasz papier. Wygrywasz 100% czasu, a więc:
EV(papier) = 1·(+1) + 0·0 + 0·(−1) = +1
Tyle teorii. W praktyce zobaczysz EV w każdej gierce poniżej. Wszystkie warianty zebrane w jednym miejscu znajdziesz też w narzędziu pawel.life/tools/rps-gto.
Kamień, papier, nożyce
Zasady pewnie znasz. Kamień bije nożyce, nożyce biją papier, papier bije kamień.
Zacznijmy od strategii, które GTO nie są. Mówię ci wprost: rzucam kamień w 100% przypadków. Co robisz? Grasz papier. Każdą rundę. EV(papier) = +1,00, wygrywasz wszystko. Moja strategia jest exploitable do bólu.
Kliknij Podpowiedź w grze, jeśli chcesz gotową kontrę.
Tym razem przeciwnik poprawia się, mimo że nadal uwielbia kamień to rzuca go rzadziej: kamień 80%, papier 10%, nożyce 10%. Jaki jest Twój maksymalny exploit? Odpowiedz, zanim klikniesz.
Pokaż odpowiedź
Nie zmieniasz nic, dalej grasz sam papier. Policzmy:
EV(papier) = 0,8·(+1) + 0,1·0 + 0,1·(−1) = +0,70 na rundę
Leak jest mniejszy, ale nadal miażdżysz przeciwnika. Zauważ na wykresie, że linia skacze mocniej niż w poprzedniej grze. To wariancja: przeciwnik czasem trafi nożyce w twój papier i pojedyncze rundy przegrasz, mimo że zagranie jest zyskowne.
Przeciwnik to widzi i jeszcze raz zmienia swoją strategię, tym razem: kamień 50%, papier 30%, nożyce 20%. Co robisz?
Pokaż odpowiedź
Twoja strategia pozostaje taka sama. Maksymalny exploit to nadal papier 100%, ale zarabiasz już zdecydowanie mniej.
EV(papier) = 0,5·(+1) + 0,3·0 + 0,2·(−1) = +0,30 na rundę
Widzisz schemat. Dopóki mix przeciwnika jest niezbalansowany, ty zawsze znajdziesz na niego kontrstrategię z dodatnim EV. Zauważ, że wariancja ponownie robi się większa, kiedy edge maleje.
Ile wyciągać z błędów przeciwnika?
Powiedzmy, że znam twój leak: lubisz kamień i podświadomie rzucasz go 40% czasu, resztę po równo (30%). Teoretyczny maks dla mnie, to granie samego papieru:
EV(papier) = 0,4·(+1) + 0,3·0 + 0,3·(−1) = +0,10 na rundę
To jest max exploit. Wyciskam z twojego błędu każdy możliwy punkt EV. Ale +0,10 to mały edge: na wykresie możesz zobaczyć, jak przez pierwsze kilkanaście rund wariancja potrafi go ukryć. Małe przewagi realizuje się na dłuższym dystansie, nie w trakcie jednej sesji.
I drugi haczyk. Gram sam papier, więc po kilku rundach widzisz, co się dzieje. I się dostosowujesz.
Przy stole to samo. Zauważyłeś, że ktoś folduje za często na 3-bet, więc 3-betujesz go szerzej. Ale jak zaczniesz to robić co rękę, to jest spora szansa, że się zorientuje. I stracisz przyszłe szanse na ten exploit. Sztuką jest znalezienie exploita. Sztuką też jest używanie go tak, żeby zarabiał jak najwięcej.
Jedyna strategia, której nie pokonasz
Kamień, papier i nożyce po równo, każde 1/3 czasu, wybierane losowo.
Ile tutaj zarabiasz?
Pokaż odpowiedź
EV(papier) = 1/3·(+1) + 1/3·0 + 1/3·(−1) = 0,00
Zero. Same nożyce? To samo zero. Mieszasz po równo jak przeciwnik? Też zero. Nie istnieje adjustment, który cokolwiek tu wygra.
To jest właśnie unexploitable. Żadna wiedza o mixie przeciwnika nic ci nie daje. I działa to w obie strony: każdy mix inny niż 1/3, 1/3, 1/3 da się w tej grze wykorzystać i pokonać. Pograj chwilę i patrz na wykres: cokolwiek wymyślisz, linia będzie się kręcić wokół zera.
Twój exploit też ma exploit
Wróćmy na chwilkę do początku. Przeciwnik grał sam kamień, ty grałeś sam papier i zabierałeś całą pulę. Czy TWOJA strategia była wtedy optymalna?
Była, ale tylko pod tą konkretną strategię przeciwnika.
Wygrywała z przeciwnikiem, ale była tak samo exploitowalna. Wystarczyło, żeby przeciwnik się połapał i przerzucił na same nożyce, i role się odwracają. Bo każdy exploit ma swój exploit. Kontra na kamień to papier, kontra na twój papier to nożyce, a na nożyce znowu kamień. Krąg się zamyka. Tutaj łatwo to dostrzec, w pokerze trudniej.
Warto zapamiętać ten prosty mechanizm, bo przy stole pokerowym dzieje się bez przerwy: kontrując czyjś błąd, sam wystawiasz się na exploit.
Czas na przedostatnią praktykę. Przeciwnik poniżej nie ma sztywnego mixu. Liczy twoje częstotliwości z ostatnich 10 rund i gra kontrę na twój najczęstszy ruch. Czy wiedząc jak on gra, dasz radę być krok przed nim i wygrać?
Podpowiedź w grze pokazuje, jak skontrować jego kontrę.
Czym poker różni się od tej prostej gry?
W kamień, papier i nożyce strategia GTO ma jedną brutalną wadę: nie wygrywa. EV równego miksowania przeciwko jakiejkolwiek strategii to zero. Gwarantuje ci remis z każdym przeciwnikiem świata. I nic ponad to.
W pokerze jest inaczej. Poker jest dużo bardziej skomplikowany.
Strategia GTO w pokerze wygrywa. Nie dlatego, że jest sprytna. Dlatego, że ludzie po drugiej stronie popełniają błędy, a ich odchylenia od równowagi oddają ci EV. A będzie tych błędów mnóstwo, bo perfekcyjna gra w pokera jest zwyczajnie niemożliwa dla człowieka. Za dużo możliwych scenariuszy, za mało pamięci operacyjnej.
I tu wracamy do tytułu. Skoro nikt nie potrafi grać perfekcyjnie, to nikt nie gra GTO. Ja nie gram. Ty nie grasz. Solvery liczą przybliżenia, a my uczymy się przybliżeń tych przybliżeń. I na tych przybliżeniach uczymy się exploitów. Bo nie ma exploitu bez GTO.
GTO to punkt odniesienia. Baseline, który musisz rozumieć, żeby wiedzieć, kiedy i dlaczego od niego odchodzisz.
Podobał Ci się taki styl artykułu? GTO w prostej postaci, a do tego interaktywne przykłady? Daj znać pisząc do mnie, albo przynajmniej przyciskiem na dole.