Version 12 of the popular open-source chess engine Stockfish was
released on Thursday. It is a major update that includes an efficiently
updatable neural network and is significantly stronger than earlier
versions.
According to the official Stockfish blog, version
12 of Stockfish plays significantly stronger than any of its
predecessors: "In a match against Stockfish 11, Stockfish 12 will
typically win at least 10 times more game pairs than it loses."
The rating difference compared to Stockfish 11 is estimated to be about 130 Elo points. This became clear from internal progression tests during development and was recently confirmed In Chess.com's computer chess environment where Stockfish+NNUE and Leela Chess Zero performed quite comparably. Image: Gary Linscott/Github.
The jump in strength is mainly the result of the introduction of
NNUE, which stands for "efficiently updatable neural network." At the
same time, Stockfish remains a CPU-only engine.
Stockfish 12 combines
the brute-force computing approach of traditional Stockfish with the
advanced evaluation capabilities of neural network chess engines such as
AlphaZero and Leela Chess Zero. In
the new version, two evaluation functions are available. Besides the
classical evaluation function, there is now the NNUE evaluation which is
computed with a neural network based on basic inputs. The network is
optimized and trained on the evaluations of millions of positions.
Wotawa, 1937 German Chess Magazine #1979
White to Move
This
study is known to be a tough nut to crack for chess engines. On this
author's computer, Stockfish 11 took 4 minutes and 54 seconds to show
the winning move as its first choice. Stockfish 12 took only 38 seconds.
One early user wrote on Reddit:
"From personal experience, Stockfish 12’s positional understanding is
really unlike anything before it. Even Leela on extremely powerful
hardware can seem blind to Stockfish 12’s search. Pretty much every
major blindspot, like pawn structure weaknesses, seems gone in this new
evaluation."
Stockfish 12 is available for download here with the neural network already embedded and ready to go. Chess.com expects to soon have Stockfish 12 available in our "max analysis" feature. See also:
How to defeat online chess cheating - once and for all
by Marco Saba, August 15, 2020 (Italian translation below the text)
To
understand this document you must be well versed in chess, psychology
and intelligence. And you must know an online chess system, like lichess.org pictured above.
Online chess games are fun and allow you to keep track of your own ELO
score in the various types of game: classic, blitz, fast, or game
variants like the one invented by Fischer and others. The unpleasant
thing is that, since the inscription involves using a name as an alias,
as it happens that someone cheats, it is expelled, but then reincarnated
under a different alias and continues to bother players. In fact, the
cheater uses to play using a chess program that is normally never less
than 3300 ELO points. It is obvious that with such a player there is no
hope and that, since he apparently always has a score of around 1700
(because sooner or later he is discovered and expelled), if you have
2000 and more ELO points a defeat costs you dearly in terms of ELO
score. If you practice a little bit, you will see that your ELO
performance drops a lot when you enter online tournaments where you meet
two or more cheaters. If your usual performance ELO is, say, 2100 ELO,
it makes sense that if you have a drop in performance in the next
tournament to, say 1700 ELO, you will have met some cheaters.
The organizers of the website, when they notice this, will send you a
message that you have played with a cheat and they will reinstate the
points lost, after expelling the cheat from the system (see illustration
above with three system communications about this). But this is not
enough to discourage these cheaters who spend their time hurting others
because they don't know any better. (And, in my opinion, this too is a
side effect of neoliberalism, because they have lost their sense of
collective good and how to play together without cheating. The
imperative Mors Tua Vita Mea, invented by bankers to make the victims
tear each other to pieces, has unfortunately always worked so far as
DIVIDE ET IMPERA).
The aim of the cheaters is certainly not to make their career,
accumulating points, because sooner or later they are expelled anyway.
And it is certainly not to spend the time of their silly life working as
labourers of a program that they use secretly to simulate a skill they
don't have. Their goal is to DESTROY and disrupt the careers of others.
They are repressed sadists who, hidden behind the immunity of
anonymity, prevent others from progressing. They may have suffered
violence as children or be the children of bankers, who knows. Disturbed
people, anyway.
A classic cheater on lichess:
How could one discourage these antisocial cheats?
To this unresolved chess problem, I propose this solution: it is not
enough to reintegrate the victim of the points he lost because of the
undeserved defeat, it must be adequately compensated. Since the penalty
for the cheater is the loss of the score obtained and expulsion from the
system, and it is known that he uses programs that have the score of
3300 ELO points, you should compensate the victim by giving him the
score as if he had WON a game against a 3300 opponent. For example, for a
2000 ELO points player who lost 5 points, you would have to reinstate
the points lost AND ADD THE POINTS AS IF YOU WON !
That is, in the case of a 2000 ELO player cheated by the 3300 ELO program, you would have an increase of + 32 ELO (http://www.ewbilliards.com/EloCalculater/),
so it would go up to 2032, instead of simply being reinstated to 2000.
This system would discourage the only goal that cheaters manage to
achieve: to keep the ELO of honest players artificially low. And it
would avoid that many players decide to abandon lichess.org
as an online gaming platform for good because it's too frustrating to
always lose with a player-program that even if it shows a low score
(from 1300 to 1700) actually has such a high actual score (3300) when the human chess world champion only reaches 2900 !
It is necessary to frustrate sadistic predators and reward honest
players, this is the only way to make sure that everyone plays with
satisfaction seeing the deserved progress they make.
It only
takes one or two cheaters per tournament to lose the desire to play
chess online. It's a shame not to use a bit of psychology to identify
and eliminate cheaters who want to do nothing but harm others, because
they are not capable of doing any good.
A similar system should also be imagined for other situations,
for example the bank credit cartel. But I will discuss this in another
blog dedicated to economics and monetary sovereignty, where for example I
easily deciphered the "miracle" of Norway by simply reading their
banking law:
Come sistemare chi bara nei tornei di scacchi online
I giochi di scacchi online sono divertenti e
permettono di tenere traccia di un proprio punteggio ELO nelle varie
tipologie di gioco: classico, blitz, rapido, oppure varianti del
gioco come quella inventata da Fischer ed altre. La cosa antipatica è
che, poiché l’iscrizione prevede di usare un nome come alias, via
via che accade che qualcuno bara, esso viene espulso, ma poi si
reincarna con un alias differente e continua ad importunare i
giocatori. Infatti, il baro usa giocare utilizzando un programma di
scacchi che di norma non è mai inferiore a 3300 punti ELO. E’
evidente che con un giocatore del genere non c’è speranza e che,
siccome apparentemente ha sempre un punteggio di circa 1700 (perché
prima o poi viene scoperto e viene espulso), se voi avete 2000 e più
punti ELO una sconfitta vi costa cara in termini di punteggio ELO. Se
fate un po’ di pratica, vedretoe che il vostro ELO di performance
cala parecchio quando partecipate a tornei online dove incontrate due
o più cheaters. Se il vostro ELO di performance abitualmente è,
poniamo, 2100 ELO, è logico che se nel torneo successivo avete un
calo di performance a, per esempio 1700 ELO, avrete incontrato
qualche cheaters. Gli organizzatori del sito internet, quando se ne
accorgono, vi inviamo un messaggio in cui vi comunicano che avete
giocato con un baro e vi reintegrano i punti persi, dopo aver espulso
il baro dal sistema (vedi illustrazione con tre comunicazioni del
sistema in merito). Ma questo non è sufficiente per scoraggiare
questi bari deficienti che passano il tempo a danneggiare gli altri
perché non sanno fare altro di meglio. (e, secondo me, anche questo
è un effetto collaterale del neoliberismo, perché si è perso il
senno del bene collettivo e di come giocare assieme senza barare.
L’imperativo Mors Tua Vita Mea, inventato dai banchieri per far
sbranare le vittime tra di loro, ha sempre purtroppo funzionato
finora come DIVIDE ET IMPERA). Lo scopo dei bari non è certo far
carriera loro, accumulando punti, poiché prima o poi vengono
comunque espulsi. E non è certo quello di passare il tempo della
loro vita insulsa lavorando come manovali di un programma che usano
di nascosto per simulare un’abilità che non hanno. Il loro scopo è
DISTRUGGERE la carriera degli altri. Sono dei sadici repressi che,
nascosti dietro l’immunità dell’anonimato, impediscono agli
altri di progredire. Potrebbero aver sofferto violenze da piccoli
oppure essere figli di banchieri, chissà. Persone disturbate,
comunque.
Come
si potrebbe disincentivare questi bari asociali ?
A
questo problema scacchistico, finora non risolto, propongo questa
soluzione: non basta reintegrare la vittima dei punti che ha perso a
causa della sconfitta immeritata, bisogna compensarlo adeguatamente.
Siccome la pena per il baro è la perdita del punteggio ottenuto e
l’espulsione dal sistema, e si sa che usa programmi che hanno il
punteggio di 3300 punti ELO, si dovrebbe indennizzare la vittima
dandogli il punteggio come se avesse VINTO una partita contro un
avversario da 3300. Per esempio, per un giocatore da 2000 punti ELO
che ha perso 5 punti, bisogna reintegrare i punti persi E IN PIÙ
AGGIUNGERE IL PUNTEGGIO COME SE AVESSE VINTO !
Cioè,
nel caso di un giocatore da 2000 truffato dal
programma.3300, si avrebbe un incremento di
+ 32 ELO
(http://www.ewbilliards.com/EloCalculater/)
, così passerebbe a 2032, invece che
essere semplicemente reintegrato a 2000. Questo sistema scoraggerebbe
l’unico obiettivo che riescono ad ottenere i bari: quello di far
rimanere artificialmente basso l’ELO dei giocatori onesti. Ed
eviterebbe che molti giocatori decidano di abbandonare
definitivamente lichess.org come piattaforma di gioco online perché
è troppo frustrante perdere sempre con un giocatore-programma che
anche se mostra un punteggio basso (da 1300 a 1700) ha in realtà un
punteggio effettivo così altro (3300) quando il campione del mondo
umano di scacchi non arriva che a
2900 !
Occorre
frustrare i predatori sadici e premiare i giocatori onesti, questo è
l’unico modo per far sì che tutti giochino con soddisfazione
vedendo effettivamente riconosciuti i meritati progressi che fanno.
Bastano
uno o due bari per torneo per far perdere la voglia di giocare a
scacchi online. E’ un peccato non usare un po’ di psicologia
identificando ed eliminando i bari che non vogliono far altro che
male agli altri, perché essi bene a se stessi non sono capaci di
farlo.
Un
sistema simile dovrebbe essere immaginato anche per altre situazioni,
ad esempio il cartello
del credito bancario. Ma questo lo tratterò in un altro blog
dedicato all’economia e alla sovranità monetaria, dove
ad esempio ho decifrato facilmente il “miracolo” della Norvegia
leggendo semplicemente la loro legge bancaria:
7/31/2020 – Just last week a new chess game launched on
the Steam client — 5D chess. The term 3D stands for third dimension,
while 4D includes “time” plus the three spatial dimensions. 5D chess
boldly claims to go even beyond that as it mixes chess, as we know it,
with a multiverse time-travel function. And this is where my mind starts
to go a bit crazy! | Photos: 5D Chess press kit
Find the right combination!
ChessBase 15 program + new Mega Database 2020 with 8 million games and
more than 80,000 master analyses. Plus ChessBase Magazine (DVD +
magazine) and CB Premium membership for 1 year!
Many fictional films and works of literature tackle the time travel
theme in various forms. Some try to change the past, others attempt to
reshape the future. There are also versions that consider that time
travel simply cannot influence our destiny.
CHESS INFORMANT’S 144 th ADVENTUREJUBILEE Presents 350 pages of the very best in chess:
Matanovic Jubilee – A Tribute to Chess Legend Rogers - Keres Memorial 1985 (Rogers’ Reminiscences) Harikrishna – FIDE Nations Cup Review Gormally – Chess in the 90s (Danny’s Chess Diary) Ivanisevic – My Repertoire Against the Dutch (Ivan’s Short Cuts) Leitao – No Limits for Creativity (Bossa Nova) Perunovic – Chess in Time of Corona (Opening Novelties) Foisor – Magnus Invitational Review (State of Play with Sabina) Prusikin – The Isolated Pawn Couple Moradiabadi – The Ragozin Defence - New Trends Delchev – Tactics Training Marin – Old Wine in New Bottles Griffin – The Birth of Hubner Variation Chess Studies Section – Sergey Didukh
Traditional sections: games, combinations, endings, correspondence chess, endgame blunders, Tournament reviews, the best game from the preceding volume and the most important theoretical novelty from the preceding volume. The periodical that pros use with pleasure is at the same time a must have publication for all serious chess students!
We bring the very best in chess to you for more than 50 years!
La scomparsa di Ennio Morricone è un grande lutto anche per il mondo
degli scacchi. Basti ricordare che Morricone ha musicato l’inno delle Olimpiadi degli Scacchi
di Torino del 2006. Ma era anche un vero appassionato, benché per il
suo lavoro non abbia mai potuto dedicarsi al gioco come forse avrebbe
voluto.
Ennio Morricone imparò a giocare a scacchi da giovane, a 18 anni,
come autodidatta, ed ha sempre coltivato la passione, pur non avendo il
tempo di prendere parte a competizioni agonistiche, salvo alcuni tornei
negli Anni Sessanta, prima del match Fischer-Spassky.
In più occasioni si è però misurato con noti campioni in partite
amichevoli o in simultanea: ricordiamo per esempio una sfida su due
partite semilampo con la celebre Judit Polgar, quando la campionessa
venne in Italia (a Roma 1994 per la precisione) a rappresentare
l’Ungheria in occasione delle cerimonie per l’ingresso ufficiale della
nazione magiara nella comunità europea. Poi una sfida ancora semilampo
con Peter Leko (fu lui a chiedere di giocare con il Maestro Ennio!),
quando Morricone era a Budapest per un concerto; e un’altra con Miso
Cebalo in occasione di un concerto a Zagabria.
Morricone-Cebalo (foto Davor Visnjic PIXSELL)
In simultanea giocò con Mariotti, Zichichi, Tatai (dal quale prese
anche lezioni private), Karpov (con cui a Roma perse per un banale
errore in un finale evidentemente patto), Kasparov e Spassky. Con
quest’ultimo va ricordato il suo pareggio nel novembre 2000, quando il
celebre (ex) campione del mondo fu invitato a Torino per le celebrazioni
dei 90 anni di costituzione della “Scacchistica Torinese”:
simultanea con 26 giocatori, compreso il figlio di Morricone, Andrea; il
Maestro fu l’ultimo a finire, la posizione era evidentemente patta e il
russo fu costretto alla divisione del punto.
Grazie a questo evento, qualche tempo dopo Michele Cordara (Presidente della “Scacchistica Torinese”) gli chiese di musicare l’inno delle Olimpiadi degli Scacchi
di Torino 2006: era preparato ad un rifiuto, invece Morricone non solo
accettò ma scrisse due spartiti, lasciando poi a Cordara il compito di
scegliere il preferito! Per la cronaca l’inno fu suonato con grande
successo all’inaugurazione delle Olimpiadi, ma in ‘playback’ per esplicita richiesta di Morricone stesso.
In più occasioni Ennio Morricone si è dichiarato dispiaciuto per non
essersi potuto dedicare di più agli scacchi. In una intervista a un
settimanale, dopo aver ricevuto l’Oscar alla carriera, alla domanda “Un desiderio che non è riuscito a realizzare?” ha risposto “Diventare un campione di scacchi, più bravo di Kasparov!”
Anche se in alcuni database appaiono partite di Morricone datate
1950, il suo esordio ufficiale in torneo risale al 1964, quando dal 3 al
13 dicembre Alvise Zichichi organizzò la prima manifestazione
scacchistica internazionale di Roma: 77 giocatori complessivamente,
suddivisi in 4 tornei (Principale con 12 giocatori, vinto ex aequo dal
tedesco Lehmann e dall’ungherese Lengyel, poi tornei A, B, C). Tra i
partecipanti del torneo C vi era anche Ennio Morricone, che
nell’occasione fu promosso a Terza categoria sociale.
In seguito arriverà a ‘Seconda Nazionale’. Ma, con buona pace degli scacchisti che badano solo al ‘punteggio Elo’,
Morricone non va considerato tanto come agonista quanto per
l’attenzione che ha saputo richiamare attorno al nostro gioco. Per
esempio quando nel 2000 pattò con Spassky, questo risultato ebbe grande
eco nei giornali e nelle TV.
Così, quando venne a Roma (1994) a rappresentare l’Ungheria in
occasione delle cerimonie per l’ingresso ufficiale della nazione magiara
nella comunità europea, Judit Polgar (allora incinta del primo figlio)
stupì i giornalisti presenti quando disse che prima di ogni partita di
torneo ascoltava un brano di Morricone!
Ricordo che per quell’occasione Ennio mi telefonò – era passata
mezzanotte e devo dire che mi spaventai quando il telefono squillò,
perché temevo qualche brutta notizia – e quando risposi, tutto gioioso
disse “Devo giocare con Judit Polgar! Capece, deve venire a Roma!” al che io un po’ stupidamente risposi “Adesso?”. Ma poi ci chiarimmo e qualche giorno dopo andai a Roma a seguire l’evento nella sede della Accademia d’Ungheria,
facendo anche da arbitro per le due partite semilampo: Morricone, forse
per l’emozione, perse malamente e rapidamente la prima, mentre oppose
un buona resistenza nella seconda. Quando gli proposi di suonare
qualcosa al pianoforte per Judit, il figlio Andrea mi disse “Non illuderti, non lo fa mai quando c’è un po’ di pubblico…”. E invece Ennio si sedette al piano e per un paio di minuti suonò il motivo di ‘Un pianista sull’oceano’ e fu un momento di forte commozione per tutti.
L’emittente pubblica australiana finanziata dai contribuenti, ABC,
sta affrontando il ridicolo dopo aver annunciato che avrebbe ospitato un
dibattito chiedendo se gli scacchi sono razzisti perché il giocatore
con gli scacchi bianchi, tradizionalmente, muove per primo.
Un ex membro dell’Australian Chess Federation ha scoperto che il
dibattito avrebbe avuto luogo quando gli è stato chiesto, da un
produttore radiofonico, di partecipare.
“L’ABC ritiene che gli scacchi siano RAZZISTI, perche` la squadra con
gli scacchi bianchi muove per prima!” ha twittato John Adams,
aggiungendo: “Fidatevi dell’emittente nazionale finanziata dai
contribuenti per applicare quadri ideologici marxisti a qualsiasi cosa
in Australia!”
Gli australiani furiosi si sono rivolti a Twitter per esprimere
indignazione per il fatto che la ABC stava sprecando tempo e denaro per
una domanda così irrilevante.
“La gente vuole che l’emittente nazionale si concentri su questioni
più grandi”, ha affermato uno. “Le persone stanno lottando con
l’economia, con la loro salute, con il lockdown. Non vogliono sprecare i
loro soldi in cazzate.”
Tuttavia, l’esperto di scacchi australiano Kevin Bonham ha dichiarato
che avrebbe preso parte al dibattito, che si e` svolto ieri.
Si scopre che gli scacchi bianchi non hanno assolutamente nulla a che
fare con il “razzismo” e 5 minuti di ricerche affrettate lo avrebbero
confermato.
I pezzi erano tinti di bianco o nero perché erano i colori più
disponibili al momento dell’invenzione del gioco nel nord-ovest
dell’India.
La regola che il bianco muove per primo faceva parte di uno sforzo
per standardizzare gli scacchi a scopi di competizione internazionale
nel 19 ° secolo e non ha letteralmente nulla a che fare con la razza.
Ma viviamo nel 2020, e tutto è potenzialmente razzista.
Il mondo degli eSport viene spesso associato ai videogame, in particolar modo a opere come Fortnite,
Call of Duty, PES e altri titoli di fama mondiale. Avete mai pensato,
però, che gli eSport e il mondo dello streaming sono applicabili anche a
uno dei giochi da tavolo più vecchi al mondo? Parliamo degli scacchi, un gioco che potrebbe sembrare noioso e fuori moda, ma che negli ultimi anni ha visto una nuova ondata di successo.
Un perfetto esempio è Alexandra Botez, attuale detentrice del titolo “Woman FIDE Master”,
prima presidente donna del Club di Scacchi della Stanford University e
una delle migliori giocatrici del Canada. La giovane donna, infatti, ha
iniziato a sfruttare Twitch per giocare in streaming lo scorso
settembre, dedicandovisi sei giorni a settimana. Ora ha 60.000 follower.
Ovviamente non si tratta di una cifra incredibile se paragonata al
successo delle stelle di Twitch, Mixer, Facebook Gaming e YouTube, ma
considerando che parliamo di scacchi, è un risultato notevole. Gli
scacchi, infatti, pur radunando un pubblico inferiore rispetto a quello
dei videogame, hanno visto una crescita del 500% dal 2016 a oggi su Twitch(questo
secondo i dati condivisi dalla piattaforma stessa). Ovviamente questo
successo ha portato denaro in tale mercato, tramite donazioni degli
appassionati ma sopratutto tramite sponsorizzazioni.
Si tratta di un cambiamento non indifferente, sopratutto considerando
che il campionato mondiale di scacchi è stato varie volte annullato per
mancanza di fondi: ora, invece, la nuova natura streaming ed eSport del
gioco permette validi introiti ai giocatori professionisti come Botez.
Parte del merito è da attribuire a Twitch stesso che si è avvicinato
al mondo degli scacchi stringendo una collaborazione con Chess.com, il
più grande sito di scacchi del mondo con circa 33 milioni di membri.
Twitch ha inoltre affermato che il Campionato Mondiale di Scacchi del 2018 ha attirato ben 4.4 milioni di spettatori unici.
Il mondo degli scacchi sta quindi cambiando: ovviamente non siamo
ancora ai livelli dell’eSport dei videogame, ma si tratta di primi
interessanti passi.
Chess player 'won't play for Iran' due to ban on Israeli players
Iran’s top rated chess champion
would be the country's second sports figure in recent months to renounce
his citizenship over pressures on Iranian athletes to forego matches
with Israeli competitors.
DUBAI - Iran’s top rated chess champion has decided not to play
for his country, Iranian news agencies reported on Tuesday, in an
apparent reaction to Tehran’s informal ban on competing against Israeli players.Alireza Firouzja,
the world’s second-highest rated junior player, would be the second
Iranian sports figure in recent months to try to renounce his
citizenship over pressures on Iranian athletes to forego matches with
Israeli competitors.
In October, Iran was banned indefinitely from
international judo by the sport’s world body until it could guarantee
that its athletes would be allowed to face Israelis. The move came after
an Iranian judoka said he was pressured to drop out of bouts to avoid
facing an Israeli athlete.“Firouzja
has made his decision and has told us that he wants to change his
nationality,” the president of Iran’s Chess Federation, Mehrdad
Pahlavanzadeh, told the semi-official news agency Tasnim.“Firouzja
is currently living in France ... and may want to play under the French
or U.S. flag,” Pahlavanzadeh told the news agency ISNA.Firouzja
wanted to take part in an upcoming world championship in Russia even
though Iran had decided not to attend, Pahlavanzadeh said, without
referring to Israel.Firouzja could not be reached for comment.
In April, Iranian media reported that Firouzja had refused to play against an Israeli player in a tournament in Germany.Iranian
political and sports officials have openly called on the country’s
athletes not to compete against Israelis as a sign of opposition to
Iran’s arch-enemy and solidarity with the Palestinians.Iranian Supreme Leader Ayatollah Ali Khamenei has repeatedly praised athletes who have refused to face opponents from Israel.Since its Islamic Revolution in 1979 Iran has refused to recognize Israel.
In a paper published in the journal Science late last year, Google parent company Alphabet’s DeepMind detailed AlphaZero,
an AI system that could teach itself how to master the game of chess, a
Japanese variant of chess called shogi, and the Chinese board game Go.
In each case, it beat a world champion, demonstrating a knack for
learning two-person games with perfect information — that is to say,
games where any decision is informed of all the events that have
previously occurred.
But AlphaZero had the advantage of knowing the rules of games it was
tasked with playing. In pursuit of a performant machine learning model
capable of teaching itself the rules, a team at DeepMind devised MuZero, which combines a tree-based search (where a tree
is a data structure used for locating information from within a set)
with a learned model. MuZero predicts the quantities most relevant to
game planning, such that it achieves industry-leading performance on 57
different Atari games and matches the performance of AlphaZero in Go,
chess, and shogi.
The researchers say MuZero paves the way for learning methods in a
host of real-world domains, particularly those lacking a simulator that
communicates rules or environment dynamics.
“Planning algorithms … have achieved remarkable successes in
artificial intelligence … However, these planning algorithms all rely on
knowledge of the environment’s dynamics, such as the rules of the game
or an accurate simulator,” wrote the scientists in a preprint paper
describing their work. “Model-based … learning aims to address this
issue by first learning a model of the environment’s dynamics, and then
planning with respect to the learned model.”
Model-based reinforcement learning
Fundamentally, MuZero receives observations — i.e., images of a Go
board or Atari screen — and transforms them into a hidden state. This
hidden state is updated iteratively by a process that receives the
previous state and a hypothetical next action, and at every step the
model predicts the policy (e.g., the move to play), value function
(e.g., the predicted winner), and immediate reward (e.g., the points
scored by playing a move).
Above: Evaluation of MuZero throughout training in chess, shogi, Go, and Atari. The y-axis shows Elo rating.
Image Credit: DeepMind
Intuitively, MuZero internally invents game rules or dynamics that lead to accurate planning.
As the DeepMind researchers explain, one form of reinforcement
learning — the technique that’s at the heart of MuZero and AlphaZero, in
which rewards drive an AI agent toward goals — involves models. This
form models a given environment as an intermediate step, using a state
transition model that predicts the next step and a reward model that
anticipates the reward.
Commonly, model-based reinforcement learning focuses on directly
modeling the observation stream at the pixel level, but this level of
granularity is computationally expensive in large-scale environments. In
fact, no prior method has constructed a model that facilitates planning
in visually complex domains such as Atari; the results lag behind
well-tuned model-free methods, even in terms of data efficiency.
Above: Comparison of MuZero against previous agents in Atari.
Image Credit: DeepMind
For MuZero, DeepMind instead pursued an approach focusing on
end-to-end prediction of a value function, where an algorithm is trained
so that the expected sum of rewards matches the expected value with
respect to real-world actions. The system has no semantics of the
environment state but simply outputs policy, value, and reward
predictions, which an algorithm similar to AlphaZero’s search (albeit
generalized to allow for single-agent domains and intermediate rewards)
uses to produce a recommended policy and estimated value. These in turn
are used to inform an action and the final outcomes in played games.
Training and experimentation
The DeepMind team applied MuZero to the classic board games Go,
chess, and shogi as benchmarks for challenging planning problems, and to
all 57 games in the open source Atari Learning Environment as
benchmarks for visually complex reinforcement learning domains. They
trained the system for five hypothetical steps and a million
mini-batches (i.e., small batches of training data) of size 2,048 in
board games and size 1,024 in Atari, which amounted to 800 simulations
per move for each search in Go, chess, and shogi and 50 simulations for
each search in Atari.
With respect to Go, MuZero slightly exceeded the performance of
AlphaZero despite using less overall computation, which the researchers
say is evidence it might have gained a deeper understanding of its
position. As for Atari, MuZero achieved a new state of the art for both
mean and median normalized score across the 57 games, outperforming the
previous state-of-the-art method (R2D2) in 42 out of 57 games and
outperforming the previous best model-based approach in all games.
Above: Evaluations of MuZero on Go (A), all 57 Atari Games (B), and Ms. Pac-Man (C-D).
Image Credit: DeepMind
The researchers next evaluated a version of MuZero — MuZero Reanalyze
— that was optimized for greater sample efficiency, which they applied
to 75 Atari games using 200 million frames of experience per game in
total. They report that it managed a 731% median normalized score
compared to 192%, 231%, and 431% for previous state-of-the-art
model-free approaches IMPALA, Rainbow, and LASER, respectively, while
requiring substantially less training time (12 hours versus Rainbow’s 10
days).
Lastly, in an attempt to better understand the role the model played
in MuZero, the team focused on Go and Ms. Pac-Man. They compared search
in AlphaZero using a perfect model to the performance of search in
MuZero using a learned model, and they found that MuZero matched the
performance of the perfect model even when undertaking larger searches
than those for which it was trained. In fact, with only six simulations
per move — fewer than the number of simulations per move than is enough
to cover all eight possible actions in Ms. Pac-Man — MuZero learned an
effective policy and “improved rapidly.”
“Many of the breakthroughs in artificial intelligence have been based
on either high-performance planning,” wrote the researchers. “In this
paper we have introduced a method that combines the benefits of both
approaches. Our algorithm, MuZero, has both matched the superhuman
performance of high-performance planning algorithms in their favored
domains — logically complex board games such as chess and Go — and
outperformed state-of-the-art model-free [reinforcement learning]
algorithms in their favored domains — visually complex Atari games.”
In 1957, Mikhail Tal played a training game that sparkled of tactics against his trainer, Alexander Koblents.
One of Tal's most famous chess quotes is "You must take your opponent
into a deep, dark forest where 2+2=5 and the path leading out is only
wide enough for one."
POW! BAM! Raymond Keene, aka The Penguin (er, top), finally gets zapped by the Spectator
SPECTATOR editor Fraser Nelson has at last grasped the nettle so many
of his predecessors – including the current prime minister – were too
timorous to touch. He has booted out the mag’s waddling, bow-tied chess
writer Raymond “the Penguin” Keene, so called because of his uncanny
resemblance to a Batman arch-villain.
Keene’s chess column in the
Speccie two weeks ago had a brief postscript revealing that it was his
last, after 42 years – though it omitted to say why he was going. As a
Spectator source puts it, the Penguin’s offence was “keeping on
plagiarising even after he’d been politely asked not to”.
The Eye has been exposing the Penguin’s bubonic plagiarism, in his books
as well as his columns for the Spectator and the Times, for more than
25 years. In June 1993 (Eye 822) we reported that an entire chapter in
his Complete Book of Gambits had been copied almost verbatim, without
acknowledgement, from an article by American international master John
Donaldson. The publisher had to pay Donaldson $3,000 and withdraw the
chapter from the book, but the waddling word-stealer carried on
regardless.
His most frequently pilfered source was Garry Kasparov’s My Great
Predecessors (2003) – a book he looted so shamelessly he didn’t even
bother to change its distinctive punctuation. Here is Kasparov on a 1954
world championship match: “At the most appropriate moment! By driving
the knight from f6 the pawn spreads confusion in the black ranks.” And a
2009 Spectator column by Keene about the same game: “At the most
appropriate moment! By driving the knight from f6 the pawn spreads
confusion in the black ranks.”
Kleptomania
In 2013 Keene published his Little Book of Chess Secrets – which, as Eye
1357 showed, was largely cut-and-pasted from Kasparov’s My Great
Predecessors. A 2013 Spectator column was also lifted from the Kasparov
book (Eye 1343), reprinting its analysis of a 1923 game between Alekhine
and Rubinstein but pretending it was Keene’s own. For good measure, the
Penguin removed quotation marks and attribution in cases where Kasparov
had quoted other greats, such as Bobby Fischer, thus presenting all the
words as if they were his own.
Keene’s editors can’t pretend they were unaware of his kleptomania. In
October 2008 the chess historian Edward Winter sent a letter by
registered post and email to the Spectator’s then editor: “Dear Mr
d’Ancona, May I advise you that over one third of the chess article by
Mr Raymond Keene published on page 64 of The Spectator, 7 June 2008 was
simply copied, word for word, from what I wrote some two years ago…
Thank you very much in advance for informing me of your proposal for
settling this matter.” Matthew d’Ancona never replied.
Grand larcenist
And how did the Penguin react? By, er, plagiarising another article by
Edward Winter for a Spectator column a year later (Eye 1222), this time
about chess coverage in the Guinness Book of Records. When challenged by
online commenters, Keene affected astonishment. “All I can think of
think of is that somewhere Winter’s comments may have been quoted
without authorship or attribution,” he spluttered, “so I regarded them
as being in the public domain.” Which prompted the question: did he
seriously think it OK to pass off someone else’s article as his own so
long as he didn’t know the author’s name?
In 2013 the chess blogger Justin Horton identified no fewer than 137
columns by the Penguin from the previous three years that included
substantial passages stolen from books and articles by other authors.
“How many more daylight robberies can he get away with,” we asked in Eye
1354, “before the editors of the Times and the Spectator call a halt to
his criminal spree?” Six years on the Spectator has at last got the
message, but Times editor John Witherow still seems unconcerned that his
chess writer is a grand larcenist.
Having been caught red-handed so often, Keene does now occasionally
acknowledge sources – though only fleetingly. Thus in the Times on
Monday 7 October, analysing a 1962 game between Petrosian and Korchnoi,
he said his notes were “based on” those in a book by Dutch grandmaster
Jan Timman. But what does he mean by “based on”? Here is Keene’s
assessment of Korchnoi’s fifth move: “Dubious. With reversed colours
this set-up is fine as the king’s bishop has already been fianchettoed,
although the line won’t yield any advantage. Here, however, the missing
tempo makes itself painfully felt…” And here is Timman’s: “Dubious. With
reversed colours this set-up is OK – since the king’s bishop has
already been fianchettoed – although it won’t yield any advantage then.
But the missing tempo makes itself painfully felt…”
For his Times column the next day, he turned to another Petrosian game
from 1962, against Bobby Fischer. No mention of Timman’s book this time,
but anyone familiar with it would have had severe déjà-vu all over
again. Timman, for example, wrote that after 17 moves “White’s position
is anything but healthy and in the text he manages to save his skin in
an endgame that, with accurate play, he will just about be able to
save.” Keene, by contrast, wrote that after 17 moves “White’s position
is anything but healthy. The text is designed to head for an endgame
that, with accurate play, he should just be able to save.”
Can the beleaguered Penguin save his skin at the Times? Tune in next week, same Bat-time, same Bat-channel.