Google DeepMind Seals Gemini Test to Protect AI Benchmarks

Google DeepMind has piloted a double-blind evaluation framework to prevent benchmark contamination during frontier AI testing. Google DeepMind announced on Thursday that it successfully completed a secure evaluation of the Gemini 2.5 Flash Lite model using a cryptographically isolated environment. The project was executed in collaboration with the Singapore AI Safety Institute, OpenMined, AVERI, and … Continue reading Google DeepMind Seals Gemini Test to Protect AI Benchmarks