
Technology and AI
If the model already saw the test, the score is mush
If a model has already seen the test, the score is mush. Google DeepMind says it is piloting a double-blind evaluation so neither the model nor the people holding the benchmark get a peek at the other’s secrets.
The first run is Gemini Flash Lite. Partners named on the post include Singapore’s AISI, OpenMined, AVERI, and MLCommons.
The point is a cleaner number: a hidden quiz, not a practice exam the model already crammed.
That is the whole beat. Blind the test, then trust the score.
Source: Google DeepMind