Google DeepMind Seals Gemini Test to Protect AI Benchmarks

Google DeepMind Seals Gemini Test to Protect AI Benchmarks

Google DeepMind pilots double-blind tests to keep AI benchmarks honest. Image: Google DeepMind

Google DeepMind tested Gemini 2.5 Flash Lite behind a cryptographic wall designed to protect confidential AI benchmarks and proprietary model weights.

Aug 28, 2026

AI benchmarks are supposed to reveal what models can do, but Google DeepMind is now putting the tests behind a cryptographic wall to make sure the models have not seen the answers first.

Google DeepMind said Thursday that it has piloted what it describes as the first double-blind evaluation of a proprietary frontier AI model, using a cryptographically protected environment to keep both the model and evaluation prompts hidden from each side.

The project involved the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. AVERI evaluated Gemini 2.5 Flash Lite using reserved prompts from MLCommons’ AILuminate safety benchmark, covering cyberattacks, chemical and biological hazards, hate speech, self-harm and violent-crime elicitation. Singapore AISI separately tested the model using confidential prompts focused on harmful content in Singapore’s context.

The setup addresses a growing problem in AI testing: benchmark contamination. If a model or its developer has access to test questions before an evaluation, a strong score may reflect familiarity with the benchmark rather than the model’s underlying ability.

Google said traditional protections such as zero-logging policies and contractual restrictions have helped keep evaluation prompts confidential, but cryptographic safeguards can add another layer of protection.

Neither side gets to peek

The system uses Google Cloud’s Confidential Computing technology to place the model and evaluation data inside a protected environment.

The evaluator cannot access Google’s model weights, while Google cannot access the evaluator’s test prompts. The pilot ran on a Google Cloud A3 Confidential VM using Intel TDX host-memory encryption and an NVIDIA H100 Confidential GPU. Hardware encryption and remote attestation were used to keep the benchmark prompts and model weights isolated while verifying the software environment.

The approach is designed to reduce a long-standing trade-off in external AI testing: evaluators previously had to either provide sensitive test material to model developers or ask companies to expose proprietary model weights. Google DeepMind said the approach could be particularly useful for sensitive evaluations involving cybersecurity and government bodies.

More Google coverage

Advertisement

The methodology is public, but the scores are not

The pilot leaves one major question unanswered: how Gemini 2.5 Flash Lite performed. DeepMind’s announcement and technical report describe the evaluation architecture and safety categories but do not publish model scores or a task-by-task results breakdown.

The technical report also acknowledges several limitations. Some proprietary inference code could not be fully inspected or allowlisted, individual Confidential Space builds were not independently reproducible, and Google services were used to sign and verify the attestation report, placing Google in the verification path and increasing the trust required in the model provider.

MLCommons also cautioned that technical secrecy alone is not enough; legal protections and careful benchmark stewardship remain important.

What this could change

The bigger significance of the experiment is not how Gemini scored on one safety benchmark. It is whether AI companies can eventually prove that their benchmark results were earned without allowing evaluators or developers to influence the test.

That distinction could become increasingly important as benchmark scores shape decisions by regulators, researchers and businesses. A secure evaluation process could make independent testing easier without forcing companies to surrender model weights or evaluators to expose valuable test sets.

For IT leaders assessing vendor claims, the method could eventually provide stronger evidence that AI models were tested against independent, previously unseen benchmarks. Until the process becomes reproducible and detailed results are released, buyers should still ask who supplied the benchmark, who evaluated the outputs, what findings were disclosed and which parts of the system required trust in the model provider.

Advertisement

But for double-blind testing to become a meaningful industry standard, the process will need to be independently reproducible, transparent about methodology and capable of scaling across models and benchmarks. Otherwise, the industry could end up with more secure tests without necessarily having more trustworthy results.

Read more: Google’s restricted rollout of Gemini 3.5 Flash Cyber shows why enterprises should examine independent performance evidence and testing controls before adopting specialized AI models.

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. He has written for a wide range of technical and business audiences, from IT professionals and cybersecurity leaders to small business owners, executives, and technology buyers. His work has appeared in publications including: TechRepublic eWEEK Channel Insider Geekflare Enterprise Networking Planet eSecurity Planet CIO Insight Webopedia With a background in computer science, Aminu specializes in translating complex technical subjects into clear, practical, and accessible content. His writing helps readers understand emerging technologies, evaluate business software, strengthen cybersecurity strategies, and make more informed decisions about technology investments. Across his work, Aminu focuses on the real-world impact of technology, connecting technical innovation with business value, operational efficiency, security, and long-term digital transformation.