Tagged: evals
A Model Built to Act Like Users Fooled the Judge 20% of the Time. GPT-6 Astra Managed 0.22%.
AI

A Model Built to Act Like Users Fooled the Judge 20% of the Time. GPT-6 Astra Managed 0.22%.

humans& released Persimmon, a 550B model trained to behave like people rather than help them. In a multi-user Turing test it fooled an LLM judge 19.8% of the time; GPT-6 Astra playing a person managed 0.22%. Frontier models overshare, stay coherent 98% of the time over 80 turns where humans manage 87%, and never change their mind, which means the simulated user in your agent eval is running easy mode.

Cognition's New Model Scores 92.8 and 27.3 on the Same Benchmark. The Difference Is a Version Number.
AI

Cognition's New Model Scores 92.8 and 27.3 on the Same Benchmark. The Difference Is a Version Number.

Cognition's SWE-2 scores 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4. Same model, same benchmark, two versions. Version 4.0 dropped the saturated tasks and the ones with public solutions, then recalibrated the compute and time budget, which makes a score a function of the model, the task-set version, the harness and the resource allowance. Vendors publish the first one.

OpenAI Shipped an Agent Platform You Can't Sign Up For. The Loop Inside It Is Free.
AI

OpenAI Shipped an Agent Platform You Can't Sign Up For. The Loop Inside It Is Free.

OpenAI Presence launched July 22 as a managed layer over its models for enterprise voice and chat agents, with no self-service option and deployments led by OpenAI Forward Deployed Engineers. What it sells is an operating loop, not a model: scope, simulate, review production sessions, approve changes. OpenAI published the same six-stage loop as a free cookbook.

AI

Watching Your Agent Work Is Not the Same as Knowing It Works

Teams instrument their agents before they grade them, 89 percent run observability and only 52 percent run evals. Watching what an agent did is not the same as knowing whether it was any good.

A 30 Minute Eval Harness You Will Actually Run Every Week
AI

A 30 Minute Eval Harness You Will Actually Run Every Week

As open coding models hit similar capability ceilings, the differentiator is internal evals tied to your product. Here is one you will actually run.

All ai evals agents engineering devtools