
Google DeepMind, AVERI, OpenMined and MLCommons published the first double-blind evaluation of a proprietary frontier model on 27 August. The evaluator never saw Gemini's weights, Google never saw the test prompts, and hardware attestation proved it. The enclave added under 5% overhead and the run took 1 minute 11 seconds. Agreeing on what would execute took 28 minutes and 3 seconds.

DeepSeek shipped V4-Flash-0731 on July 31 with the same architecture and size as the April preview and only a new post-training pass. DeepSWE went from 7.3 to 54.4 and the small model now beats DeepSeek's own larger V4-Pro on every agent benchmark published. The weights got a dated Hugging Face repo. The API kept the same floating name.